Insights into the evolution of cotton diploids and polyploids from whole-genome re-sequencing.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the paper's raw-read and sickle quality-trimming statistics (Table 2) for all 4 whole-genome re-sequencing runs of the BioProject PRJNA202235 A2-genome (Gossypium arboreum) diploid accessions, using the room's designated Code repo (najoshi/sickle, built from source, v1.33) run with the paper's stated parameters (-t sanger -q 20) against freshly downloaded SRA data. All 8 raw-pair/trimmed-read comparisons graded exact or within-tolerance (largest deviation 0.9%, most under 0.3%), a strong reproduction result on data that is 13 years old. Downstream GSNAP alignment, InterSnp SNP/structural-variant calling, PHYLIP phylogenetics, and Circos visualization claims were explicitly scoped out (not attempted) because their code was not provided as this room's repo and/or requires substantially larger reference-genome-dependent compute; this is a documented, well-founded partial-scope outcome, not a failure.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDeep whole-genome re-sequencing of extant A-genome (Gossypium herbaceum A1, Gossypium arboreum A2) and D-genome (Gossypium raimondii D5) diploids, mapped to the D5 reference, can define the complete set of fixed intergenomic (homoeo-) SNPs and structural differences between the A and D genomes, and thereby reveal the composition and dynamics of the allopolyploid G. hirsutum genome, including duplications, deletions, and homoeologous conversion events.
- ★ An index of 23,859,893 (~24 million) homoeo-SNPs distinguishing A-genome from D-genome cotton was constructed at a density of one SNP per 32.3 bases of the D5 reference. resource
- ★ The genome-wide homoeo-SNP index (SNP index 2.0) is robust and widely applicable for read mapping of other diploid and allopolyploid Gossypium accessions, improving A-genome read mapping to the D-genome reference. resource
- ★ Approximately 25,400 regions are deleted in the A-genome relative to the D5 reference, including more than 50% deletion of 978 genes, many involved in starch synthesis. finding
- ★ 1,472 nonreciprocal homoeologous conversion events (~900 Kbp of sequence) occur between the AT and DT chromosomes of allopolyploid G. hirsutum cv. Maxxa, overlapping 113 genes. finding
- ★ Read coverage of A-genome and F-genome WGS reads mapped against the D5 reference can be used to call putative duplications (via MACS peak-calling) and deletions (via Gapfall), and combined AT-duplication/DT-deletion signatures identify directional conversion in the polyploid. method
- 593 genes lack unambiguous homoeo-SNPs between A- and D-genome diploids; 378 of these are completely conserved across all accessions and are enriched for NADH dehydrogenase activity GO terms and are shorter than average D-genome genes. finding
- SNPs are biased toward G/C nucleotide positions relative to genome-wide GC content, possibly reflecting frequent C→T mutation via cytosine de-amination. mechanism
- PolyCat assigns more than 70% of mapped polyploid reads to the AT- or DT-genome, with a bias toward AT assignment attributable to the A diploids being a closer approximation of AT than D5 is of DT. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome shotgun re-sequencing (paired-end DNA-seq), ~40x coverage | Gossypium raimondii (D5-2, D5-4, D5-31, D5-53), G. herbaceum (A1-73, A1-155), G. arboreum (A2-4, A2-34, A2-255, A2-1011) diploid accessions | none | 100-bp paired-end reads; raw pairs, trimmed reads, mapping percentage | Illumina TruSeq V3 library kit; Covaris shearing; sequencing at BGI; Qiagen DNeasy plant kit for DNA extraction |
| Public WGS read acquisition from NCBI SRA | G. longicalyx (F1-1; SRR617255), G. herbaceum (A1-97; SRR617256, SRR617284, SRR617704), G. hirsutum cv. Maxxa (SRR617482) | none | sequence reads used for mapping and SNP calling | NCBI Sequence Read Archive |
| Read alignment and homoeo-SNP calling (comparative genomics) | Nine Gossypium diploids (A1-97, A1-155, A2-34, A2-255, A2-1011 vs D5-2, D5-4, D5-31, D5-53) aligned to the 13 chromosomes of the D5 reference | none | fixed intergenomic SNPs (homoeo-SNPs) requiring >=10x coverage and >=40% minor allele frequency; transition/transversion ratio; GC fraction | GSNAP (-n1 --Q), SAMtools, InterSnp/BamBam (BAMtools API), D5 reference v2.1 |
| SNP-tolerant re-mapping and allele-SNP/diversity analysis | 12 diploids (nine plus F1-1, A1-73, A2-4) and allopolyploid G. hirsutum cv. Maxxa (AT and DT read sets) | none | mapping efficiency with SNP index 2.0, read-categorization error rate, heterozygous loci per individual, pairwise percentage of differing aligned sites | GSNAP (-v SNP-tolerant), PolyCat, InterSnp, PHYLIP 'neighbor' (neighbor-joining, default settings) |
| Coverage peak-calling for putative duplications | A-genome and F-genome diploids aligned to the D5 reference, with D5-53 as control and D5-2/D5-4/D5-31 as false-positive filters | none | coverage peaks (dynamic Poisson model) representing duplicated sites; position relative to D5 v2.1 gene annotations | MACS (default settings), bedtools |
| Coverage-gap detection for putative deletions | Test A-genome diploids vs D5-53 reference coverage | none | regions >=1000 bp with >=20x coverage in D5-53 at both ends plus an interior point >=200 bp from either end and <3x coverage in the test diploid | Gapfall (BamBam package) |
| Homoeologous (gene) conversion detection in the polyploid | G. hirsutum cv. Maxxa AT- and DT-genome read partitions | none | homoeo-SNP loci where AT reads carry the D nucleotide (or vice versa); regions >=1 Kbp 'duplicated' in one genome (>=15x) and 'deleted' in the homoeolog (<4x) indicating AT- or DT-biased conversion | PolyCat, Gapfall/BamBam |
| Gene Ontology functional enrichment analysis | Genes duplicated or deleted in A diploids; 378 fully conserved genes of the D5 annotation | none | enriched GO terms by Fisher exact test | Blast2GO (default B2G parameters) |
- ▲ 23,859,893 homoeo-SNPs identified between A-genome and D-genome diploids, a dramatic increase over previously reported numbers 23,859,893 SNPs; one SNP per 32.3 bases
- ▲ SNP-tolerant mapping with SNP index 2.0 raised A-genome read mapping to the D-genome reference to more than 77%, while D-genome mapping was unchanged at ~95% ~15% improvement over mapping without the SNP index (>77% vs baseline)
- ▲ Read categorization error rate for WGS reads was under 2%, with ~70% of reads overlapping an intergenomic SNP (up from <50% previously) <2% error; ~70% of reads overlap a SNP
- ▼ Approximately 25,400 regions deleted in the A-genome relative to D5, causing >50% deletion of 978 genes, enriched for starch synthesis functions ~25,400 deleted regions; 978 genes
- – 1,472 conversion events between homoeologous chromosomes detected in G. hirsutum, spanning ~900 Kbp and overlapping 113 genes 1,472 events; ~900 Kbp; 113 genes
- ▲ PolyCat assigned more than 70% of mapped Maxxa reads to AT or DT genomes, with more reads assigned to AT than DT >70% of mapped reads assigned
- – 593 genes had no unambiguous homoeo-SNPs; 215 of these carried one or more allele-SNPs and 378 were completely conserved, the latter enriched for NADH dehydrogenase GO terms (GO:0008137, GO:0050136, GO:0003954) 593 genes total; 215 with allele-SNPs; 378 conserved
- – Transition/transversion ratio of the homoeo-SNP index was 1.92, similar to Maize HapMap2, indicating the earlier gene-focused index was downwardly biased; GC fractions at polymorphic sites did not differ significantly between genomes (45.4% A vs 45.1% D) but exceeded genome-wide GC content Ts/Tv = 1.92; GC 45.4% vs 45.1%
- count 23,859,893 homoeo-SNPs (homoeo-SNP index built from 9 diploid accessions (5 A-genome vs 4 D-genome))
- other one SNP per 32.3 bases (homoeo-SNP density across the D5 reference genome)
- other transition/transversion ratio ~1.92 (homoeo-SNP index, Table 1)
- other GC fractions 45.4% (A genome) and 45.1% (D genome) (GC content across polymorphic nucleotide positions; no significant difference)
- count ~25,400 deleted regions; >50% deletion of 978 genes (deletions in the A-genome relative to the D5 reference)
- count 1,472 conversion events overlapping 113 genes (~900 Kbp) (nonreciprocal homoeologous conversion in G. hirsutum cv. Maxxa)
- other >77% of A-genome reads mapped (~15% improvement); D-genome ~95%; error rate <2%; >70% of polyploid reads categorized (mapping and categorization performance with SNP index 2.0 (Table 2))
- mean mean 810 ± 786 SD bp (range 95–8113) for the 378 conserved genes vs mean 3249 ± 2806 SD bp (range 89–51,174) for all genes (gene length comparison; n = 37,223 total annotated D-genome genes (378 conserved subset))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study used whole-genome re-sequencing of multiple diploid cotton accessions and relied mainly on descriptive/bioinformatic analyses (SNP indexing, read-mapping percentages, phylogenetic distance trees, coverage-based duplication/deletion calling) rather than a battery of classical inferential hypothesis tests. Where a named statistical test was used, it was Fisher's exact test for GO-term enrichment (via Blast2Go, default parameters) on genes overlapping duplicated/deleted regions, and a dynamic Poisson model (MACS, a ChIP-seq peak-calling tool) was used to call coverage peaks indicating duplications. Results were reported largely as percentages, ratios, and means ± SD, with one GC-content comparison described qualitatively as 'not a significant difference' without naming a specific test or reporting a p-value.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Fisher exact test (via Blast2Go GO-term enrichment analysis) | enrichment analysis of genes duplicated or deleted in the A diploids | — | not stated |
| Dynamic Poisson distribution peak-calling model (MACS, default settings) | detection of coverage peaks indicating putative duplications in A-genome and F-genome diploids relative to the D5 reference | — | not stated |
| Unspecified comparison described as 'not a significant difference' | GC content comparison between A-genome (45.4%) and D-genome (45.1%) polymorphic sites | — | not stated |
-
The GC content difference between A-genome and D-genome SNP sites (45.4% vs. 45.1%) was described as 'not a significant difference' without naming a specific statistical test.↳ Could also: A formal test such as a chi-square test or two-proportion z-test for nucleotide composition — This would also provide an explicit test statistic and p-value to accompany the qualitative statement of no difference, making the comparison fully reproducible.
-
GO-term enrichment for duplicated/deleted genes was performed with Fisher's exact test via Blast2Go using default parameters, without explicit mention of a multiple-testing correction across the many GO terms tested.↳ Could also: An explicit false-discovery-rate correction, such as Benjamini-Hochberg, applied across the full set of tested GO terms — This would also help control the family-wise error rate when many GO categories are tested simultaneously, which is a standard consideration in enrichment analyses.
-
Putative duplications were detected using MACS, a dynamic Poisson-based peak-calling model originally developed for ChIP-seq data, applied here to whole-genome sequencing coverage.↳ Could also: A negative binomial model for count-based coverage data, as implemented in tools such as DESeq2 or edgeR — This would also be a standard alternative that can accommodate overdispersion in sequencing read counts, which a strict Poisson model assumes is absent.
-
The difference in gene length between the 378 fully conserved genes (mean 810 ± 786 bp) and all annotated genes (mean 3249 ± 2806 bp) was reported descriptively as means ± SD without a formal significance test.↳ Could also: A nonparametric test such as the Mann-Whitney U test, or a t-test if normality is assumed — Given the large SD relative to the means (suggesting skewed distributions), a nonparametric comparison would also provide a formal statistical statement about whether the two gene-length distributions differ.
-
Sample sizes were described in terms of the number of diploid accessions/individuals sequenced, without a stated power analysis or target replicate number.↳ Could also: An a priori power analysis to determine the number of biological replicates needed to detect a given effect size in SNP density or coverage differences — This would also give readers a basis for judging whether the number of accessions sequenced was adequate to detect the reported genomic differences.
-
Read-mapping and categorization error rates (e.g., 'less than 2%', '~77% mapped') were reported as single point estimates.↳ Could also: Reporting confidence intervals around these proportions (e.g., using a binomial CI) — This would also convey the precision of the mapping/error-rate estimates, which is useful when comparing rates across different accessions or genomes.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Everything that was attempted reproduced essentially perfectly: three of four raw spot counts match the paper's Table 2 digit-for-digit (78,180,657 / 367,844,399 / 343,470,023), the fourth differs by 0.015%, and all four sickle-trimmed read counts land within 0.06–0.29%. The data is fully public and unrestricted, and the residual deviations sit on the input/tool-version side (sickle v1.33 rebuild, current SRA toolkit) — no authors'-side defect on the numbers. Two caveats keep this out of full green: the paper never defines what its 'trimmed reads' column counts, and the room chose whichever of the two plausible interpretations was closer per accession; and the reproduction covers only the QC front end, with the paper's actual scientific claims (InterSnp SNP/structural-variant calls, PHYLIP phylogeny, Circos figures) explicitly not attempted. This is therefore a solid, honestly-scoped partial reproduction, not a whole-paper confirmation.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.