A Deluge of Complex Repeats: The Solanum Genome.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the headline pipeline outputs 1:1 FROM THE SHIPPED DATA. The GitHub repo (Solanum-Repeats-Metadata @ e3a2ba0) ships the authors' repeat-annotation result data (Solanum.zip, 74 MB), not code; per BRIEF P16 this is a valid reproduction substrate. We downloaded it on a «our HPC» compute node to «infra» and independently re-derived the paper's numbers: (a) merging the shipped per-element repeat coordinates per chromosome (BedTools-style overlap merge) gives 389.76 Mb / 48.08% (potato) and 463.43 Mb / 59.29% (tomato) vs the paper's 395.51 Mb / 48.79% and 470.31 Mb / 60.17% -> ~98.5% of reported, uniformly slightly low; (b) chr12 reproduces as the most repeat-rich chromosome (54.38% potato #1; 63.74% tomato #1 among anchored chromosomes); (c) the de-novo consensus family counts reproduce EXACTLY (1,921 and 1,438); (d) the per-chromosome lengths in the shipped .xls sum EXACTLY to the paper's genome sizes. NOT ATTEMPTED (deliberate 80/20 skip): re-running the authors' own de-novo RepeatModeler+RepeatMasker pipeline on the ~800 Mb Ensembl genomes (days of compute), and the downstream exonization / 3D-structure / TF-binding / miRNA / expression analyses (manual or multi-dataset, out of pipeline scope). FABRICATION ASSESSMENT: no signal -- exact family counts and exact assembly sizes, with coverage reproduced to ~98.5% from the same shipped coordinates; the small uniform downward deficit is consistent with the paper merging additional element sources beyond the shipped coordinate file, not with inflation. Result is provisional and must be human-audited (see AUDIT.md).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 94assessed: 2026-06-16 ⛓ 52ca4d5a8734
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests whether the genomes of Solanum tuberosum and Solanum lycopersicum contain a far greater abundance of complex repetitive elements than previously appreciated, and whether these elements are transcriptionally active and exert functional/regulatory influence over host protein-coding genes.
- ★ ~50–60% of the S. tuberosum and S. lycopersicum genomes are composed of repetitive elements finding
- ★ Complex repetitive elements are associated with >95% of genes in both species, suggesting a major role in gene formation and regulation finding
- ★ Both genomes are predominantly composed of LTR retrotransposons finding
- ★ Two novel repeat families highly similar to LTR/ERV1 and LINE/RTE-BovB are reported for the first time in these species finding
- ★ Many complex repeats are transcriptionally active, as estimated from NGS read data and microarray platforms finding
- ★ Transcription factor binding sites and miRNAs appear to be under the influence of complex repetitive elements, and several genes possess exonized repeats finding
- A de-novo plus homology-based pipeline (RepeatModeler/RepeatMasker/RepeatProteinMasker) was used to identify and annotate known and novel complex repeats genome-wide method
- Exonization of repetitive elements alters orthologous gene/protein sequence and predicted protein secondary and 3D structure mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| de-novo and homology-based repeat identification/annotation (RepeatModeler, RepeatMasker, RepeatProteinMasker, RECON, RepeatScout, TRF) | S. tuberosum (group phureja doubled monoploid clone) and S. lycopersicum (cv. Heinz 1706) genomes | none | repeat family identity, annotation, genomic coordinates and coverage | RepeatModeler/RepeatMasker/Repbase |
| genome-wide distribution / coverage analysis of repeats relative to genes | S. tuberosum and S. lycopersicum genomes (Ensembl Plants) | none | percentage chromosome coverage by repeats; overlap with genes, 5kb upstream, exons vs introns; PCC between repeat and exon coverage | BedTools, in-house PERL, R |
| transcriptional abundance / activity measurement of repeats | S. tuberosum and S. lycopersicum | none | transcriptional abundance of repetitive elements | Next Generation Sequencing read data and Microarray |
| synteny and orthology analysis | S. tuberosum vs S. lycopersicum genomes/proteomes | none | syntenic regions and orthologous gene pairs | Symap v42, BLASTP |
| exonization analysis via global sequence/protein alignment and structure prediction | orthologous gene/protein pairs of S. tuberosum and S. lycopersicum | none | indels/substitutions overlapping repeats and resulting protein secondary/3D structure changes | EMBOSS Stretcher, PsiPred, RaptorX |
| multiple sequence alignment and phylogenetic analysis of repeat families | identified repeat family consensus sequences | none | family relationships / phylogenetic trees | ClustalW, Neighbor-Joining (bootstrap 1000) |
| non-coding RNA / miRNA and regulatory element association analysis | S. tuberosum and S. lycopersicum | none | overlap of repeats with pre-miRNAs, ncRNAs, transcription factor binding sites | miRBase v20, Rfam v11, BLASTN/TBLASTX |
- – Repetitive elements compose ~50–60% of both S. tuberosum and S. lycopersicum genomes 50–60%
- – Complex repetitive elements overlap/associate with >95% of genes in both species >95%
- – LTR retrotransposons are the dominant repeat class in both genomes
- – Two novel repeat families similar to LTR/ERV1 and LINE/RTE-BovB identified 2 families
- – Prior annotations reported 404,861 repetitive elements for S. tuberosum and 719,453 for S. lycopersicum 404,861; 719,453
- – Reported genome sizes: 810.6 Mb (S. tuberosum) and 781.6 Mb (S. lycopersicum) 810.6 Mb; 781.6 Mb
- other 50–60% repetitive content (fraction of S. tuberosum and S. lycopersicum genomes composed of repeats)
- other >95% (percentage of genes associated with complex repetitive elements in both species)
- count 404,861 (annotated repetitive elements in S. tuberosum)
- count 719,453 (annotated repetitive elements in S. lycopersicum)
- other 810.6 Mb (genome size of S. tuberosum)
- other 781.6 Mb (genome size of S. lycopersicum)
- other ~34% (prior repeat content estimate of S. tuberosum BAC sequences (Zhu et al.))
- other ~46% (prior repeat content estimate of S. lycopersicum BAC sequences (Zhu et al.))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational genomics study that identified and characterized repetitive elements genome-wide in Solanum tuberosum and Solanum lycopersicum using de novo and homology-based pipelines (RepeatModeler/RepeatMasker). Enrichment of genes near repetitive elements was assessed with a binomial test, and the relationship between repeat density and exon coverage across chromosomes was quantified using Pearson Correlation Coefficient (PCC) with a t-test for significance. Phylogenetic relationships among repeat families were inferred by Neighbor Joining with 1000 bootstrap replicates. The study is primarily descriptive and comparative, reporting results as percentages and coverage fractions across the two genomes.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Binomial test | Testing significant enrichment of genes overlapping with repetitive elements (null: no significant enrichment near repeats) | Count of genes overlapping repetitive elements; exact n not stated | not stated |
| Pearson Correlation Coefficient (PCC) with t-test for p-value | Correlation between per-chromosome percentage coverage of repetitive elements and percentage coverage of coding regions (exons) | Number of chromosomes per species; not explicitly stated | not stated |
| Neighbor Joining with 1000 bootstrap replicates | Phylogenetic trees for selected repeat family consensus sequences aligned with ClustalW | Number of sequences per alignment; not stated | na |
-
A binomial test was used to assess whether genes are enriched near repetitive elements↳ Could also: A permutation or randomization test (e.g., shuffling gene or repeat coordinates across the genome and resampling the overlap count) could also be used — Permutation tests make fewer parametric assumptions about the null distribution and directly account for the non-uniform, structured nature of genomic coordinate data (chromosome lengths, centromeric gaps), which can influence the binomial null
-
Pearson Correlation Coefficient (PCC) was used to relate per-chromosome repeat coverage to exon coverage↳ Could also: Spearman rank correlation could also be used for this chromosome-level analysis — Genomic coverage fractions across chromosomes may not follow a bivariate normal distribution, and with a small n equal to the number of chromosomes, Spearman rank correlation is a commonly chosen non-parametric alternative that does not assume linearity or normality
-
Multiple per-chromosome PCC tests and multiple binomial tests across repeat categories were conducted without a stated multiple-testing correction↳ Could also: A false discovery rate (FDR) correction such as Benjamini-Hochberg, or a family-wise correction such as Bonferroni, could also be applied across the family of chromosome-level or category-level tests — When multiple hypothesis tests are conducted simultaneously, applying a correction limits the expected proportion of false positives; reporting which correction (if any) was used is standard practice for transparency and reproducibility
-
Neighbor Joining (NJ) was used to construct phylogenetic trees for repeat families with 1000 bootstrap replicates↳ Could also: Maximum likelihood (ML) or Bayesian inference (e.g., RAxML, IQ-TREE, MrBayes) phylogenetics could also be applied to these alignments — ML and Bayesian methods incorporate explicit substitution models and are generally considered to produce more statistically rigorous topologies and branch-length estimates than distance-based NJ, particularly for divergent repeat sequences; they also provide posterior probabilities or model-corrected bootstrap support
-
Repeat annotation confidence was resolved by a hierarchy of tools (RepeatModeler, RepeatMasker, RepeatProteinMasker) with manual inspection as a fallback↳ Could also: A formal probabilistic annotation framework (e.g., REPET pipeline or EDTA) could also integrate evidence from multiple sources into a single scored annotation — Probabilistic pipelines can provide quantitative confidence scores for each annotation decision rather than a deterministic rule-based hierarchy, which may aid downstream interpretation of borderline classifications
-
Transcriptional activity of repeats was assessed using NGS read data and microarray platforms (described in abstract), with results reported descriptively↳ Could also: A formal differential expression or abundance testing framework (e.g., DESeq2, edgeR, or limma-voom for RNA-seq; limma for microarray) could also be applied to quantify and test repeat-element transcriptional abundance — Formal count-based or linear-model frameworks provide statistical tests and FDR-controlled p-values for transcript abundance differences, enabling more precise statements about which repeat families show evidence of active transcription above background
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Complex repetitive elements overlap or associate with >95% of genes in both S. tuberosum and S. lycopersicum.other solanum-spp 2015×1papers★ This paper is the founder (earliest)
-
LTR retrotransposons are the dominant repeat class in both S. tuberosum and S. lycopersicum genomes.other solanum-spp 2015×1papers★ This paper is the founder (earliest)
-
Two novel repeat families with similarity to LTR/ERV1 and LINE/RTE-BovB were identified in Solanum genomes.other solanum-spp 2015×1papers★ This paper is the founder (earliest)
-
Repetitive elements constitute 50–60% of both S. tuberosum and S. lycopersicum genomes.other solanum-spp 2015×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-26241045
Paper: Mehra M, Gangwar I, Shankar R (2015). A Deluge of Complex Repeats: The Solanum Genome. PLoS ONE 10(8):e0133962. PMID 26241045 / PMC4524691 / DOI 10.1371/journal.pone.0133962.
Repo: https://github.com/mrigayamehrajha/Solanum-Repeats-Metadata (commit e3a2ba0a546207250c6e1859d7ff29c70d0331c2, pushed 2015-06-03). Contains README (90 B) + Solanum.zip (74 MB) = "Complete repeat information for the genomes of Solanum tuberosum and Solanum lycopersicum". This repo ships the authors' RESULT DATA (the repeat annotation), not their analysis code. Per BRIEF rule P16 a third-party tool on the paper's data is equally valid; here, more directly, we can check whether the paper's headline numbers are derivable from the shipped annotation — a 1:1 reproduction-from-data + fabrication check.
Pipeline used by the paper (Methods)
De-novo repeat discovery + annotation:
- RepeatModeler (RECON + RepeatScout + TRF) → de-novo consensus repeat library (reported: 1,921 families for S. tuberosum, 1,438 for S. lycopersicum).
- RepeatMasker annotates the genome against the de-novo library (+ Repbase).
- RepeatProteinMasker for protein-domain repeats.
- BedTools merge to collapse overlapping repeat coordinates → non-redundant repeat bp.
- Genomes from Ensembl Plants: S. tuberosum group phureja DM (PGSC_DM_v4.03), S. lycopersicum cv. Heinz 1706 (ITAG v2.3 / SL2.40).
IN SCOPE (pipeline-derived, attempted)
| # | Result | Reported | Pipeline | Reproduction route |
|---|---|---|---|---|
| C1 | Genome-wide repeat bp/% — S. tuberosum | 395,513,917 bp / 810,654,046 bp = 48.79% (~49%) (Table 1, Abstract) | RepeatModeler→RepeatMasker→BedTools merge | Parse shipped annotation, sum non-redundant repeat bp, compare |
| C2 | Genome-wide repeat bp/% — S. lycopersicum | 470,312,762 bp / 781,666,411 bp = 60.17% (~60%) (Table 1, Abstract) | same | same |
| C3 | Per-chromosome repeat coverage (Table 1); chr12 most repeat-rich (~55% potato, ~64.6% tomato) | Table 1 | same | Aggregate shipped annotation per chromosome |
| C4 | RepeatModeler consensus family counts (1,921 / 1,438) | Results | RepeatModeler | Check if shipped library/annotation exposes family counts |
OUT OF SCOPE (not attempted — why)
- De-novo RepeatModeler re-run on full ~800 Mb genomes: days of compute; the hard last 20%. We reproduce from the shipped annotation instead and document this as the deliberately-skipped expensive step.
- Exonization / orthologous gene-pair indels, secondary-structure changes (27,923 pairs; 61 structural): multi-step downstream + manual structure work.
- 3D structure threading (RaptorX/PsiPred/LigPlot+), TF binding-site gain percentages, miRNA-repeat association: downstream/manual, out of scope.
- Expression (RPKM via SeqMap/R-seq), small-RNA (Bowtie): separate datasets (SRP*/E-MTAB*/GSE*), not the core repeat result; out of scope.
Note: the manifest lists data: geo:GSE22300, but that GEO series is only a
S. lycopersicum microarray cited for expression context — not the input to
the repeat pipeline. The actual repeat-pipeline inputs are the Ensembl Plants
genome assemblies; the reproducible result data is the repo's Solanum.zip.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean reproduction from the authors' own deposited result data (Solanum.zip): de-novo family counts (1,921/1,438) and assembly sizes match exactly, chr12 is confirmed as the densest chromosome, and genome-wide repeat coverage reproduces to ~98.5%. The only deviation is a small, uniform ~1.5% downward deficit in total repeat bp (-5.75 Mb potato, -6.88 Mb tomato), most plausibly because the paper merged additional element sources beyond the shipped coordinate file — an input/preprocessing difference on the data side, not a computation error. Severity is negligible and the central conclusion holds fully; no fabrication signal.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.