Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Insights into the Evolution of the New World Diploid Cottons (Gossypium, Subgenus Houzingenia) Based on Genome Sequencing.

Genome Biol Evol · 2019
L1 87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and reproduced 1:1. This P16-style reproduction re-ran the authors' OWN post-processing scripts (jackKnife.R, chisq.R) on their OWN shipped intermediate inputs on «our HPC» (SLURM 2176464, R 4.3.3). Claim 1 (the paper's headline introgression result, Z=-3.64 for G. davidsonii->G. aridum Colima admixture via ANGSD ABBA-BABA block jackknife) reproduced EXACTLY: our output is byte-identical to the repo's shipped aridum.txt (same SHA256) and rounds to the paper's -3.64. Claim 2 (Table 4 introgressed-SNP chi-square) reproduced within Monte-Carlo tolerance: all deterministic floor p-values match to the last digit and every significance call is identical; borderline/non-significant rows differ only in Monte-Carlo noise (simulate.p.value=TRUE, unknown seed). NOT attempted (heavy ~20%, needs full raw SRA PRJNA488266 / days of compute): genome assembly+stats (Tables 1-2), BUSCO (Table 3), k-mer/GenomeScope genome sizes, read mapping + SNP/indel calling, dN/dS, TE clustering/dating, RAxML phylogenies. The SNP COUNTS feeding Table 4 are upstream variant-calling output and were used as shipped inputs (graded partial), not regenerated. No fabrication indicators: every checked value is fully derivable from the shipped data + scripts.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 87
    assessed: 2026-06-14 ⛓ 12df57f24901
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

How did the New World diploid D-genome cottons (Gossypium subgenus Houzingenia) diversify, and what are the pace and patterns of molecular evolution of their genes and repetitive sequences, including phylogenetic relationships, divergence timing, introgression, and genome-size dynamics?

Core claims
  • Subgenus Houzingenia originated via transoceanic dispersal from Africa ~6.6 Ma, with most biodiversity arising from rapid mid-Pleistocene (0.5–2.0 Ma) diversification plus multiple long-distance dispersals. finding
  • Comparison of cpDNA versus nuclear data reveals several clear cases of interspecific introgression in the subgenus. finding
  • Repetitive DNA comprises roughly half of the ~880 Mb genome, but most transposable element families are relatively old and stable among species. finding
  • A 2-fold bias toward small deletions over insertions counteracts TE-driven genome-size growth, explaining the small genomes restricted to this subgenus. mechanism
  • Pairwise synonymous mutation rates average ~1% per Myr, with nonsynonymous changes about seven times less frequent. finding
  • The rate of indel occurrence is much lower than nucleotide substitution, averaging ~17 nucleotide substitutions per indel event. finding
  • Whole-genome resequencing produced new genome and plastome assemblies for most included species, providing community resources. resource
  • Phylogenomic + ancestral-state reconstruction workflow (assembly, mapping, ML phylogeny, divergence dating, indel/SNP and repeat characterization) for closely related plant species. method
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome resequencing (Illumina, 2×100 bp and 2×150 bp) Gossypium subgenus Houzingenia D-genome species (leaves), plus outgroup G. longicalyx none genomic sequence reads / coverage Illumina HiSeq2500 (HiSeq PE Rapid v2 / SBS Kit v2); BGI, Novogene, and USDA-ARS GBRU libraries (Nextera, Accel-NGS 2S PCR-Free)
De novo genome assembly and annotation D-genome Gossypium species assemblies none assembly size, E-size, % genome covered, gene models, BUSCO recovery ABySS v2.0.1, Chromosomer v0.1.3, MAKER v2.31.6, BUSCO v2.0
Reference mapping and per-gene assembly for phylogenetics Houzingenia species mapped to G. raimondii reference none gene alignments for ML phylogeny BWA v0.7.10, samtools, BamBam v1.3, RaxML
Divergence time estimation / ancestral state reconstruction Houzingenia species tree none node ages (Ma), ancestral genome sizes R chronos/ape, phytools fastAnc
Indel and SNP characterization and effect analysis Houzingenia genomes vs G. raimondii reference none >1.1 million indels polarized, SNP introgression proportions, indel effects on genes GATK, ANGSD, SnpEff, SnpSift
Gene family / copy number evolution analysis orthologous gene clusters across species none gene family sizes and ancestral states OrthoFinder, Count (birth-and-death model)
Repetitive sequence (graph-based) characterization 1% genome-size-equivalent read subsamples per accession none repeat cluster abundance, genome occupation (Mb) per repeat type RepeatExplorer, RepeatMasker, custom cotton-enriched repeat library
Repeat heterogeneity / relative age estimation repeat clusters from D-genome accessions none pairwise percent identity divergence profiles classifying clusters as young/old all-vs-all BLASTn, regression models with BIC
Key results
  • Subgenus likely originated from African transoceanic dispersal ~6.6 Ma
  • Rapid mid-Pleistocene diversification accounts for nearly all biodiversity 0.5–2.0 Ma
  • Bias toward small deletions over insertions detected from polarized indels 2-fold
  • Synonymous substitution rate ~1% per Myr
  • Nonsynonymous changes much rarer than synonymous ~7-fold less frequent
  • Substitution events per indel event ~17 substitutions per indel
  • Repetitive DNA fraction of genome ~50% of ~880 Mb
  • Assemblies cover most of each genome 67–85% covered; 585–775 Mbp (avg 643 Mbp)
Key statistics
  • other average of 33× genomic coverage (average per-sample sequencing coverage after filtering)
  • count over 1.1 million indels (indels detected and phylogenetically polarized)
  • fold_change 2-fold bias toward deletions over small insertions (indel deletion vs insertion bias)
  • other ~1% per Myr (average pairwise synonymous mutation rate)
  • fold_change ~7 times less frequent (nonsynonymous vs synonymous change frequency)
  • count genome size 841–934 Mb (~1.11-fold range) (genome sizes across D-genome cottons (Hendrix and Stewart 2005))
  • count 169.4 M raw reads/accession; 136.9 M after filtering (range 67.2–260.2 M) (sequencing read counts per accession)
  • count 20,522–45,244 gene models per accession (annotation output across assemblies)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a phylogenomic and molecular-evolution study based on whole-genome resequencing of New World diploid cottons. The overall approach is model-based and comparative rather than hypothesis-test-driven: phylogeny was inferred by maximum likelihood (RAxML, GTRGAMMA) with rapid bootstrapping, divergence times were estimated via penalized/maximum likelihood, gene-family size and ancestral states were reconstructed with likelihood-based birth-death and continuous-trait models, repeat composition was summarized with Principal Coordinate Analysis, and relative repeat age was classified using regression models selected by Bayesian Information Criterion. Results are reported largely as estimated quantities (coverage, rates, counts, percentages, divergence dates) rather than as p-values or formal significance tests.

Replicationbiological Sample sizesequencing coverage described (avg 33x; raw 22–65x; avg 169.4 M reads/accession, post-filter avg 136.9 M); at least one accession per D-genome species; no formal power analysis described Groups13–14 Houzingenia species/accessions, rooted with G. longicalyx outgroup Pairingna Randomization/blindingna Dispersionunclear Exact p-valuesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Maximum likelihood phylogenetic inference (RAxML, GTRGAMMA model) concatenated nuclear gene alignment phylogeny of subgenus Houzingenia genes filtered by coverage/length with >=1 accession per species (exact count not stated) stated
Rapid bootstrapping with consensus tree generation node support on the ML phylogeny (10,000 alternative runs on distinct starting trees) na
Penalized likelihood / maximum likelihood divergence-time estimation (chronos, ape) species divergence time / temporal framework stated
Likelihood-based phylogenetic birth-and-death model (Count) gene-family copy-number evolution and ancestral state reconstruction 1,000 resampling-with-replacement permutations stated
Continuous-trait ancestral state reconstruction (fastAnc, phytools; fitContinuous, Geiger) genome size and per-cluster repeat abundance across the tree not stated
Principal Coordinate Analysis (PCoA) multivariate repeat-cluster abundance patterns per genome (log-normalized or genome-size-normalized) na
Approaches that could also have been used
  • Node support was assessed with rapid bootstrapping in a maximum likelihood framework.
    Could also: Bayesian inference (e.g., MrBayes or RevBayes) yielding posterior probabilities, and/or coalescent-based species-tree methods (e.g., ASTRAL) on per-gene trees could also be used. — A coalescent species-tree approach would also accommodate gene-tree discordance from incomplete lineage sorting and introgression, which is relevant given the reported introgression cases, while Bayesian posteriors offer a complementary measure of clade support.
  • Genes were concatenated into a single supermatrix for the ML phylogeny.
    Could also: Partitioned models with per-locus or per-codon-position rate parameters could also be applied to the same alignment. — Allowing substitution parameters to vary among loci/codon positions would also capture among-gene rate heterogeneity and is a widely used complement to a single GTRGAMMA model across all sites.
  • Divergence times were estimated with penalized/maximum likelihood (chronos) using calibrated node ages.
    Could also: A Bayesian relaxed-clock approach (e.g., BEAST or MCMCtree) could also be used. — A Bayesian dating analysis would also propagate calibration and rate uncertainty into credibility intervals around each node age, providing an explicit measure of dating uncertainty.
  • Repeat-cluster abundance patterns were summarized and visualized with Principal Coordinate Analysis.
    Could also: Formal multivariate tests such as PERMANOVA (adonis) or model-based differential-abundance frameworks could also accompany the ordination. — These would also provide a quantitative, testable assessment of how much repeat composition differs among groups, complementing the visual ordination.
  • Relative repeat age was classified by selecting regression models with the Bayesian Information Criterion.
    Could also: Reporting alongside AIC/AICc, or model-averaging across candidate models, could also be used. — Comparing information criteria or averaging would also convey how strongly the data favor the selected trend versus close alternatives.
  • Estimated quantities (rates, counts, divergence dates) were generally reported as point values.
    Could also: Accompanying these with confidence/credible intervals or bootstrap-derived ranges could also be reported. — Reporting interval estimates alongside point estimates would also convey the precision of each estimate, which is often preferred when summarizing evolutionary rates and dates.
Software: RAxML (maximum likelihood phylogenetics) · R 3.4.4 · ape (chronos, tree visualization) · phytools (fastAnc) · Geiger (fitContinuous) · Count (gene-family birth-death modeling) · GATK (Genome Analysis ToolKit; SNP/indel characterization) · ANGSD (introgression) · SnpEff / SnpSift (indel effects) · OrthoFinder (orthology) · RepeatExplorer / RepeatMasker (repeat profiling)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
57
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

MH477706 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
MH477724 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
PRJNA488266 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-30476109

Paper: Grover et al. 2019, "Insights into the Evolution of the New World Diploid Cottons (Gossypium, Subgenus Houzingenia) Based on Genome Sequencing." GBE 11(1):53-71. PMID 30476109 · PMCID PMC6320677 · DOI 10.1093/gbe/evy256 Code: https://github.com/IGBB/D_Cottons_USDA (master, commit f9675740faf21a361d6f33a89ce3665565908e98, pushed 2019-07-18) Data: SRA PRJNA488266 (raw Illumina reads, 13 D-genome cotton species)

Pipeline map (which reported results come from a bioinformatic pipeline)

Result Pipeline In scope? Why
Genome assemblies (ABySS) + stats (Table 1/2) trim→ABySS v2.0.1→E-size select OUT (heavy 20%) needs all raw SRA reads + many-kmer assemblies; days of compute, not low-hanging
BUSCO (Table 3) MAKER annot → BUSCO v2 OUT depends on assemblies above
k-mer genome sizes (GenomeScope) jellyfish→GenomeScope per accession OUT (heavy) needs raw reads per accession to build k-mer histograms
Read mapping + SNP/indel calling BWA→samtools→variant calls OUT (heavy) needs raw reads + reference; upstream of the stats below
ABBA-BABA D-statistic / Z (introgression) ANGSD abbababa → block-jackknife jackKnife.R IN shipped input (aridum.abbababa) + script + expected output (aridum.txt); deterministic; seconds
Introgressed-SNP chi-square (Table 4) chisq.R on shared.derived.snps IN (post-proc) shipped input + script + expected p-values (Table4 xlsx); Monte-Carlo (within-tol). NOTE: the SNP counts are upstream variant-calling output (out of scope to regenerate); only the chi-square step is reproduced
CNV evolution rates count.bootstrap.R (CAFE-style) OUT (optional) candidate, not attempted (80/20)
dN/dS, indel, TE/clustering, phylogeny (RAxML) various R + RAxML OUT heavy / depend on upstream alignments

What we attempt (IN scope, low-hanging, fully shipped)

  1. D-statistic introgression test — re-run the authors' jackKnife.R on the shipped aridum.abbababa (3 individuals: D4-12.Garid1=aridum Colima, D4-185=aridum Jalisco, D3-D-27.Gdavi1=davidsonii H3) and check the resulting Patterson's D + block-jackknife Z against the shipped aridum.txt AND the paper's reported Z = -3.64.
  2. Introgression chi-square (Table 4) — re-run the authors' chisq.R on the shipped shared.derived.snps and compare Monte-Carlo p-values to the manuscript Table 4 xlsx.

All compute on «our HPC» («infra»). Repo cloned + R env built inside the compute job («infra»).

D-stat-Z-aridumColima
Reported
Z = -3.64 (ABBA-BABA introgression of a G. davidsonii-like species into G. aridum Colima)
Reproduced
Z = -3.640015 ; D = -0.005534797 (byte-identical to shipped aridum.txt, SHA256 fb7b037a...)
exact
D-stat-davi-vs-arid
Reported
D = 0.600078 (Z=150.29) and D = 0.603608 (Z=154.10) (repo aridum.txt rows 2-3)
Reproduced
D = 0.600078 (Z=150.2918) and D = 0.603608 (Z=154.096); byte-identical
exact
introg-snp-overall-p
Reported
Table 4 overall p = 0.0004997501 (=1/2001)
Reproduced
p = 0.0004997501 (=1/2001)
exact
introg-snp-genic-p
Reported
Table 4 genic p = 0.0014992504 (=3/2001)
Reproduced
p = 0.0004997501 (=1/2001); same significance call (p<<0.01), differs only by Monte-Carlo noise
within tolerance
introg-snp-chisq-vector
Reported
Table 4 full p-value vector (16 rows)
Reproduced
all 8 floor rows match exactly at 1/2001; every significance call matches; non-sig rows agree within 2nd-3rd sig-fig Monte-Carlo noise
within tolerance
introg-snp-counts
Reported
Table 4 shared-derived SNP counts (overall 188472/182563; genic 7843/7419)
Reproduced
shipped intermediate input verified == Table 4; NOT regenerated from raw SRA (upstream variant calling out of scope)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

Re-ran the authors' own scripts (jackKnife.R, chisq.R) on their shipped intermediate inputs. The paper's headline introgression result Z=-3.64 reproduces byte-identical (Z=-3.640015, D=-0.005534797 — same SHA256 as the shipped aridum.txt), the supporting davidsonii-vs-aridum D-stats are byte-identical, and Table 4's chi-square reproduces with all 8 deterministic floor p-values exact (1/2001) and every significance call identical. The only differences are unseeded Monte-Carlo noise on borderline/non-significant rows (genic 1/2001 vs 3/2001, same p<<0.01 conclusion). Upstream variant-calling (the SNP counts) was out of scope and used as shipped inputs, verified == Table 4. No fabrication indicators — a clean 1:1 reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

105.4 k
tokens (I/O) · 6.5 M incl. cache
13 min
runtime · 0.02 CPU-h
8.3 GB
peak RAM
1
HPC jobs
hummel
machine