Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A Deluge of Complex Repeats: The Solanum Genome.

PLoS One · 2015
L1 94/100 PQI 94
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
94/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 87% of all assessed papers rank 133 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the headline pipeline outputs 1:1 FROM THE SHIPPED DATA. The GitHub repo (Solanum-Repeats-Metadata @ e3a2ba0) ships the authors' repeat-annotation result data (Solanum.zip, 74 MB), not code; per BRIEF P16 this is a valid reproduction substrate. We downloaded it on a «our HPC» compute node to «infra» and independently re-derived the paper's numbers: (a) merging the shipped per-element repeat coordinates per chromosome (BedTools-style overlap merge) gives 389.76 Mb / 48.08% (potato) and 463.43 Mb / 59.29% (tomato) vs the paper's 395.51 Mb / 48.79% and 470.31 Mb / 60.17% -> ~98.5% of reported, uniformly slightly low; (b) chr12 reproduces as the most repeat-rich chromosome (54.38% potato #1; 63.74% tomato #1 among anchored chromosomes); (c) the de-novo consensus family counts reproduce EXACTLY (1,921 and 1,438); (d) the per-chromosome lengths in the shipped .xls sum EXACTLY to the paper's genome sizes. NOT ATTEMPTED (deliberate 80/20 skip): re-running the authors' own de-novo RepeatModeler+RepeatMasker pipeline on the ~800 Mb Ensembl genomes (days of compute), and the downstream exonization / 3D-structure / TF-binding / miRNA / expression analyses (manual or multi-dataset, out of pipeline scope). FABRICATION ASSESSMENT: no signal -- exact family counts and exact assembly sizes, with coverage reproduced to ~98.5% from the same shipped coordinates; the small uniform downward deficit is consistent with the paper merging additional element sources beyond the shipped coordinate file, not with inflation. Result is provisional and must be human-audited (see AUDIT.md).

💻 Code ↗ 🗄 Data: GSE22300

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 94
    assessed: 2026-06-16 ⛓ 52ca4d5a8734
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether the genomes of Solanum tuberosum and Solanum lycopersicum contain a far greater abundance of complex repetitive elements than previously appreciated, and whether these elements are transcriptionally active and exert functional/regulatory influence over host protein-coding genes.

Core claims
  • ~50–60% of the S. tuberosum and S. lycopersicum genomes are composed of repetitive elements finding
  • Complex repetitive elements are associated with >95% of genes in both species, suggesting a major role in gene formation and regulation finding
  • Both genomes are predominantly composed of LTR retrotransposons finding
  • Two novel repeat families highly similar to LTR/ERV1 and LINE/RTE-BovB are reported for the first time in these species finding
  • Many complex repeats are transcriptionally active, as estimated from NGS read data and microarray platforms finding
  • Transcription factor binding sites and miRNAs appear to be under the influence of complex repetitive elements, and several genes possess exonized repeats finding
  • A de-novo plus homology-based pipeline (RepeatModeler/RepeatMasker/RepeatProteinMasker) was used to identify and annotate known and novel complex repeats genome-wide method
  • Exonization of repetitive elements alters orthologous gene/protein sequence and predicted protein secondary and 3D structure mechanism
Experimental setups
Assay System Perturbation Readout Platform
de-novo and homology-based repeat identification/annotation (RepeatModeler, RepeatMasker, RepeatProteinMasker, RECON, RepeatScout, TRF) S. tuberosum (group phureja doubled monoploid clone) and S. lycopersicum (cv. Heinz 1706) genomes none repeat family identity, annotation, genomic coordinates and coverage RepeatModeler/RepeatMasker/Repbase
genome-wide distribution / coverage analysis of repeats relative to genes S. tuberosum and S. lycopersicum genomes (Ensembl Plants) none percentage chromosome coverage by repeats; overlap with genes, 5kb upstream, exons vs introns; PCC between repeat and exon coverage BedTools, in-house PERL, R
transcriptional abundance / activity measurement of repeats S. tuberosum and S. lycopersicum none transcriptional abundance of repetitive elements Next Generation Sequencing read data and Microarray
synteny and orthology analysis S. tuberosum vs S. lycopersicum genomes/proteomes none syntenic regions and orthologous gene pairs Symap v42, BLASTP
exonization analysis via global sequence/protein alignment and structure prediction orthologous gene/protein pairs of S. tuberosum and S. lycopersicum none indels/substitutions overlapping repeats and resulting protein secondary/3D structure changes EMBOSS Stretcher, PsiPred, RaptorX
multiple sequence alignment and phylogenetic analysis of repeat families identified repeat family consensus sequences none family relationships / phylogenetic trees ClustalW, Neighbor-Joining (bootstrap 1000)
non-coding RNA / miRNA and regulatory element association analysis S. tuberosum and S. lycopersicum none overlap of repeats with pre-miRNAs, ncRNAs, transcription factor binding sites miRBase v20, Rfam v11, BLASTN/TBLASTX
Key results
  • Repetitive elements compose ~50–60% of both S. tuberosum and S. lycopersicum genomes 50–60%
  • Complex repetitive elements overlap/associate with >95% of genes in both species >95%
  • LTR retrotransposons are the dominant repeat class in both genomes
  • Two novel repeat families similar to LTR/ERV1 and LINE/RTE-BovB identified 2 families
  • Prior annotations reported 404,861 repetitive elements for S. tuberosum and 719,453 for S. lycopersicum 404,861; 719,453
  • Reported genome sizes: 810.6 Mb (S. tuberosum) and 781.6 Mb (S. lycopersicum) 810.6 Mb; 781.6 Mb
Key statistics
  • other 50–60% repetitive content (fraction of S. tuberosum and S. lycopersicum genomes composed of repeats)
  • other >95% (percentage of genes associated with complex repetitive elements in both species)
  • count 404,861 (annotated repetitive elements in S. tuberosum)
  • count 719,453 (annotated repetitive elements in S. lycopersicum)
  • other 810.6 Mb (genome size of S. tuberosum)
  • other 781.6 Mb (genome size of S. lycopersicum)
  • other ~34% (prior repeat content estimate of S. tuberosum BAC sequences (Zhu et al.))
  • other ~46% (prior repeat content estimate of S. lycopersicum BAC sequences (Zhu et al.))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational genomics study that identified and characterized repetitive elements genome-wide in Solanum tuberosum and Solanum lycopersicum using de novo and homology-based pipelines (RepeatModeler/RepeatMasker). Enrichment of genes near repetitive elements was assessed with a binomial test, and the relationship between repeat density and exon coverage across chromosomes was quantified using Pearson Correlation Coefficient (PCC) with a t-test for significance. Phylogenetic relationships among repeat families were inferred by Neighbor Joining with 1000 bootstrap replicates. The study is primarily descriptive and comparative, reporting results as percentages and coverage fractions across the two genomes.

Replicationunclear Sample sizeWhole reference genome assemblies of two species used (S. tuberosum group phureja DM clone; S. lycopersicum cv. Heinz 1706); no biological or technical replication described; sample size/power framing not applicable to this computational design GroupsS. tuberosum vs. S. lycopersicum genomes; chromosomes compared for repeat density vs. gene/exon density Pairingna Randomization/blindingnot stated Dispersionnone Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Binomial test Testing significant enrichment of genes overlapping with repetitive elements (null: no significant enrichment near repeats) Count of genes overlapping repetitive elements; exact n not stated not stated
Pearson Correlation Coefficient (PCC) with t-test for p-value Correlation between per-chromosome percentage coverage of repetitive elements and percentage coverage of coding regions (exons) Number of chromosomes per species; not explicitly stated not stated
Neighbor Joining with 1000 bootstrap replicates Phylogenetic trees for selected repeat family consensus sequences aligned with ClustalW Number of sequences per alignment; not stated na
Approaches that could also have been used
  • A binomial test was used to assess whether genes are enriched near repetitive elements
    Could also: A permutation or randomization test (e.g., shuffling gene or repeat coordinates across the genome and resampling the overlap count) could also be used — Permutation tests make fewer parametric assumptions about the null distribution and directly account for the non-uniform, structured nature of genomic coordinate data (chromosome lengths, centromeric gaps), which can influence the binomial null
  • Pearson Correlation Coefficient (PCC) was used to relate per-chromosome repeat coverage to exon coverage
    Could also: Spearman rank correlation could also be used for this chromosome-level analysis — Genomic coverage fractions across chromosomes may not follow a bivariate normal distribution, and with a small n equal to the number of chromosomes, Spearman rank correlation is a commonly chosen non-parametric alternative that does not assume linearity or normality
  • Multiple per-chromosome PCC tests and multiple binomial tests across repeat categories were conducted without a stated multiple-testing correction
    Could also: A false discovery rate (FDR) correction such as Benjamini-Hochberg, or a family-wise correction such as Bonferroni, could also be applied across the family of chromosome-level or category-level tests — When multiple hypothesis tests are conducted simultaneously, applying a correction limits the expected proportion of false positives; reporting which correction (if any) was used is standard practice for transparency and reproducibility
  • Neighbor Joining (NJ) was used to construct phylogenetic trees for repeat families with 1000 bootstrap replicates
    Could also: Maximum likelihood (ML) or Bayesian inference (e.g., RAxML, IQ-TREE, MrBayes) phylogenetics could also be applied to these alignments — ML and Bayesian methods incorporate explicit substitution models and are generally considered to produce more statistically rigorous topologies and branch-length estimates than distance-based NJ, particularly for divergent repeat sequences; they also provide posterior probabilities or model-corrected bootstrap support
  • Repeat annotation confidence was resolved by a hierarchy of tools (RepeatModeler, RepeatMasker, RepeatProteinMasker) with manual inspection as a fallback
    Could also: A formal probabilistic annotation framework (e.g., REPET pipeline or EDTA) could also integrate evidence from multiple sources into a single scored annotation — Probabilistic pipelines can provide quantitative confidence scores for each annotation decision rather than a deterministic rule-based hierarchy, which may aid downstream interpretation of borderline classifications
  • Transcriptional activity of repeats was assessed using NGS read data and microarray platforms (described in abstract), with results reported descriptively
    Could also: A formal differential expression or abundance testing framework (e.g., DESeq2, edgeR, or limma-voom for RNA-seq; limma for microarray) could also be applied to quantify and test repeat-element transcriptional abundance — Formal count-based or linear-model frameworks provide statistical tests and FDR-controlled p-values for transcript abundance differences, enabling more precise statements about which repeat families show evidence of active transcription above background
Software: R · RepeatModeler · RepeatMasker · RepeatProteinMasker · BedTools · Symap 42 · ClustalW · NCBI BLAST (BLASTP, BLASTN, TBLASTX) · EMBOSS Stretcher · PsiPred · RaptorX

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
26
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

E-MTAB-629 ArrayExpress in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
E-MTAB-634 ArrayExpress in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE18110 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE22300 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRP033230 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-26241045

Paper: Mehra M, Gangwar I, Shankar R (2015). A Deluge of Complex Repeats: The Solanum Genome. PLoS ONE 10(8):e0133962. PMID 26241045 / PMC4524691 / DOI 10.1371/journal.pone.0133962.

Repo: https://github.com/mrigayamehrajha/Solanum-Repeats-Metadata (commit e3a2ba0a546207250c6e1859d7ff29c70d0331c2, pushed 2015-06-03). Contains README (90 B) + Solanum.zip (74 MB) = "Complete repeat information for the genomes of Solanum tuberosum and Solanum lycopersicum". This repo ships the authors' RESULT DATA (the repeat annotation), not their analysis code. Per BRIEF rule P16 a third-party tool on the paper's data is equally valid; here, more directly, we can check whether the paper's headline numbers are derivable from the shipped annotation — a 1:1 reproduction-from-data + fabrication check.

Pipeline used by the paper (Methods)

De-novo repeat discovery + annotation:

  • RepeatModeler (RECON + RepeatScout + TRF) → de-novo consensus repeat library (reported: 1,921 families for S. tuberosum, 1,438 for S. lycopersicum).
  • RepeatMasker annotates the genome against the de-novo library (+ Repbase).
  • RepeatProteinMasker for protein-domain repeats.
  • BedTools merge to collapse overlapping repeat coordinates → non-redundant repeat bp.
  • Genomes from Ensembl Plants: S. tuberosum group phureja DM (PGSC_DM_v4.03), S. lycopersicum cv. Heinz 1706 (ITAG v2.3 / SL2.40).

IN SCOPE (pipeline-derived, attempted)

# Result Reported Pipeline Reproduction route
C1 Genome-wide repeat bp/% — S. tuberosum 395,513,917 bp / 810,654,046 bp = 48.79% (~49%) (Table 1, Abstract) RepeatModeler→RepeatMasker→BedTools merge Parse shipped annotation, sum non-redundant repeat bp, compare
C2 Genome-wide repeat bp/% — S. lycopersicum 470,312,762 bp / 781,666,411 bp = 60.17% (~60%) (Table 1, Abstract) same same
C3 Per-chromosome repeat coverage (Table 1); chr12 most repeat-rich (~55% potato, ~64.6% tomato) Table 1 same Aggregate shipped annotation per chromosome
C4 RepeatModeler consensus family counts (1,921 / 1,438) Results RepeatModeler Check if shipped library/annotation exposes family counts

OUT OF SCOPE (not attempted — why)

  • De-novo RepeatModeler re-run on full ~800 Mb genomes: days of compute; the hard last 20%. We reproduce from the shipped annotation instead and document this as the deliberately-skipped expensive step.
  • Exonization / orthologous gene-pair indels, secondary-structure changes (27,923 pairs; 61 structural): multi-step downstream + manual structure work.
  • 3D structure threading (RaptorX/PsiPred/LigPlot+), TF binding-site gain percentages, miRNA-repeat association: downstream/manual, out of scope.
  • Expression (RPKM via SeqMap/R-seq), small-RNA (Bowtie): separate datasets (SRP*/E-MTAB*/GSE*), not the core repeat result; out of scope.

Note: the manifest lists data: geo:GSE22300, but that GEO series is only a S. lycopersicum microarray cited for expression context — not the input to the repeat pipeline. The actual repeat-pipeline inputs are the Ensembl Plants genome assemblies; the reproducible result data is the repo's Solanum.zip.

Figures / tables: Table
C1
Reported
395,513,917 bp = 48.79% (S. tuberosum genome-wide repeat content)
Reproduced
389,763,736 bp = 48.08%
within tolerance
C2
Reported
470,312,762 bp = 60.17% (S. lycopersicum genome-wide repeat content)
Reproduced
463,433,170 bp = 59.29%
within tolerance
C3a
Reported
chr12 most repeat-rich, S. tuberosum (~55%)
Reproduced
chr12 = 54.38%, rank #1 of 13
exact
C3b
Reported
chr12 most repeat-rich, S. lycopersicum (~64.6%)
Reproduced
chr12 = 63.74%, densest of anchored chr1-12 (chr0=scaffolds excluded)
within tolerance
C4a
Reported
1,921 RepeatModeler consensus families (S. tuberosum)
Reproduced
1,921
exact
C4b
Reported
1,438 RepeatModeler consensus families (S. lycopersicum)
Reproduced
1,438
exact
C5
Reported
Assembly sizes 810,654,046 / 781,666,411 bp
Reproduced
810,654,046 / 781,666,411 bp (exact)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 94/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

This is a clean reproduction from the authors' own deposited result data (Solanum.zip): de-novo family counts (1,921/1,438) and assembly sizes match exactly, chr12 is confirmed as the densest chromosome, and genome-wide repeat coverage reproduces to ~98.5%. The only deviation is a small, uniform ~1.5% downward deficit in total repeat bp (-5.75 Mb potato, -6.88 Mb tomato), most plausibly because the paper merged additional element sources beyond the shipped coordinate file — an input/preprocessing difference on the data side, not a computation error. Severity is negligible and the central conclusion holds fully; no fabrication signal.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

174.4 k
tokens (I/O) · 12.3 M incl. cache
19 min
runtime · 0 CPU-h
0.2 GB
peak RAM
8 (1 failed)
HPC jobs
hummel
machine