Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Ancient gene duplicates in Gossypium (cotton) exhibit near-complete expression divergence.

Genome Biol Evol · 2014
50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced two of the paper's core pipeline steps end-to-end on one representative RNA-seq library (SRX172483/SRR530667, G. hirsutum leaf): sickle v1.33 quality trimming (8,158,340 -> 8,019,202 reads) and GMAP-GSNAP splice-aware alignment to the G. raimondii reference genome (94.86% of alignment records / 92.6% of reads mapped). Went beyond the minimum by also computing an approximate, single-genome (non-homoeolog-resolved) gene-level read count and RPKM table (30,003/38,208 genes with nonzero counts, 92.66% of mapped reads assigned to gene bodies) as a partial, honestly-scoped stand-in for the paper's full homoeolog-resolved RPKM/UQ + t-test/GLM differential-expression analysis, whichrequires infrastructure (dual A/D-genome homoeolog SNP-sorting, R-based GLM with edgeR-style negative binomial, 22 additional SRA accessions across leaf/seed/petal tissues) not established in this environment and was not attempted rather than faked. The paper itself publishes no per-accession numeric QC statistics (trim counts, mapping rates) to compare against, so all three claims are graded 'partial': the pipeline steps are verified to run correctly and produce internally consistent, biologically plausible numbers, but cannot be graded exact/within-tolerance/mismatch against a published number that does not exist. Only one of 23 total SRA accessions referenced by the paper was profiled in this run; the other 22 (SRX172484-SRX172485, SRX170955, SRX172454, SRX172473, SRX204399-SRX204401, SRX204405-SRX204407, SRX204429-SRX204434, SRX204555-SRX204558, SRX328344) were not downloaded or profiled due to scope/time, and this is an explicit, documented limitation rather than a claim of completeness.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-03
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What evolutionary processes maintain gene duplicates over long time scales following ancient whole genome duplication? The authors test whether paralogs retained from the ~60 My old Gossypium-specific 5- to 6-fold ploidy increase have undergone expression-level neo- and/or subfunctionalization.

Core claims
  • Nearly all (99.4%) ancient paralog pairs in Gossypium raimondii are differentially expressed in at least one of three tissues (petal, leaf, seed), indicating massive, near-complete expression-level divergence. finding
  • A generalized linear model showed 92.4% of paralog pairs exhibit expression divergence, with most showing significant gene-by-tissue interactions indicating complementary expression across tissues. finding
  • Expression divergence is mirrored in G. arboreum and in a G. raimondii seed developmental time series, indicating expression-level diversification occurred before species divergence in the common ancestor. finding
  • Expression-level neo- and/or subfunctionalization is essential to the maintenance of these duplicates over ~60 My. mechanism
  • 1,971 strictly duplicated paralogous gene pairs were identified in G. raimondii, traceable to the Gossypium-specific whole genome multiplication. resource
  • Strictly duplicated genes can be identified by requiring duplicate syntenic regions in G. raimondii that correspond to only a single genomic region in both Theobroma cacao and Vitis vinifera, thereby excluding the older shared eudicot triplication. method
  • Retained duplicate genes are broadly distributed across chromosomes without positional bias, apart from being densest in regions of high overall gene density such as subtelomeric regions. finding
  • Paralogs at a given chromosomal region are not more likely to be over- or underexpressed relative to their counterpart, i.e., no positional (subgenome-like) expression bias was detected. finding
Experimental setups
Assay System Perturbation Readout Platform
Comparative genomics / synteny and sequence-similarity analysis Gossypium raimondii genome vs Theobroma cacao and Vitis vinifera genomes none Identification of strictly duplicated paralogous gene pairs (duplicated syntenic regions in G. raimondii mapping to a single region in T. cacao and V. vinifera)
Coding sequence alignment and dN/dS estimation Primary transcripts of paralogous gene pairs in G. raimondii none dN/dS ratios per paralog pair ClustalW; custom BioPerl scripts; Jukes-Cantor substitution model
Bulk RNA-seq (differential expression between paralogs) Gossypium raimondii leaf, petal, and seed (10 DPA) tissue; three biological replicates per tissue none Uniquely mapped read coverage over ~37,000 gene annotations, RPKM and upper-quartile normalized; log ratio of paralog expression Public NCBI SRA data (SRX172483-SRX172485, SRX204399-SRX204401, SRX204405-SRX204407, SRX204429-SRX204434, SRX328344); sickle QC; GSNAP mapping; samtools
Bulk RNA-seq (differential expression between paralogs) Gossypium arboreum leaf, petal, and seed; three biological replicates per tissue none Uniquely mapped read coverage and paralog expression log ratios, reads mapped to G. raimondii genome using a Gossypium-specific SNP index to reduce mapping bias NCBI SRA (SRX170955, SRX172454, SRX172473, SRX204555-SRX204558, SRX328344); GSNAP; Gossypium SNP index
Bulk RNA-seq developmental time series Gossypium raimondii seed, 10-40 DPA none (developmental stage series) Differential expression between paralogs across developmental stages
Generalized linear model (negative binomial) with contrasts UQ-normalized RNA-seq counts from G. raimondii petal, seed, and leaf none Gene effect, tissue effect, gene-by-tissue interaction, per-tissue differential expression (G|T), and complementary expression patterns; FDR 5% (Benjamini-Hochberg) R (contrasts package)
Tissue-specific reciprocal silencing analysis Differentially expressed paralog pairs of G. raimondii and G. arboreum across tissues/time points none Cases where one paralog accounts for >=95% of the pair's total RPKM and the bias is reversed in one or more tissues/time points
Genomic distribution / positional bias analysis G. raimondii chromosome scaffolds; leaf, seed, and petal expression data none Circos visualization of paralog positions and gene density; binomial test for regional over- or underexpression bias, FDR 5% Circos
Key results
  • 99.4% of paralog pairs in G. raimondii are differentially expressed in at least one of petal, leaf, or seed; 93-94% per tissue 99.4%
  • 1,666 (85%) of pairs are differentially expressed in all three tissues; 88-89% in two of three tissues 1,666 pairs (85%)
  • In G. arboreum, 1,962 (99.5%) paralogs show transcriptional divergence in at least one tissue; 92-95% per tissue and 86-88% in at least two tissues 1,962/1,971 (99.5%)
  • GLM revealed 92.4% of paralog pairs exhibit expression divergence, most with significant gene and tissue interaction effects 92.4%
  • All but two paralog pairs (1,969 of 1,971) were differentially expressed in both species in at least one tissue; 87% and 90% in petal and leaf respectively in both species; 74% divergent in all tissues of both species 1,969/1,971; 74%
  • In a G. raimondii seed developmental time series (10-40 DPA), 1,961 (99.5%) paralogs were differentially expressed in at least one developmental stage 1,961 (99.5%)
  • 1,809 of 1,971 G. raimondii pairs (and 1,811 in G. arboreum) show significant and substantial (>=1.5-fold) divergence in at least one tissue; in all tissues of both species at least 25% of paralogs differ by at least 5-fold >=1.5-fold in 1,809/1,971; >=5-fold in >=25%
  • No significant departure from expectation for regional over- or underexpression bias in leaf, seed, or petal after FDR correction
Key statistics
  • count 1,971 strictly duplicated paralogous gene pairs (Paralogs identified in G. raimondii from the Gossypium-specific 5- to 6-fold ploidy increase)
  • other 99.4% (Percent of G. raimondii paralog pairs differentially expressed in at least one of three tissues)
  • count 1,666 (85%) (Pairs differentially expressed in all three tissues (petal, leaf, seed) in G. raimondii)
  • other 92.4% (Paralog pairs with expression divergence by the negative binomial GLM)
  • count 1,962 (99.5%) (G. arboreum paralogs with transcriptional divergence in at least one tissue)
  • fold_change Petal 1,462 / 1,237 / 715 pairs at >=1.5-, 2-, 5-fold; Leaf 1,380 / 1,076 / 496; Seed 1,379 / 1,119 / 551; max in any tissue 1,809 / 1,644 / 1,026 (G. raimondii) (Table 1 fold-change distribution in G. raimondii tissues)
  • fold_change Petal 1,481 / 1,259 / 742; Leaf 1,403 / 1,087 / 508; Seed 1,409 / 1,134 / 568; max in any tissue 1,811 / 1,645 / 1,027 (G. arboreum) (Table 1 fold-change distribution in G. arboreum tissues)
  • count 1,825 pairs differentially expressed in petal, 1,746 of which are also differentially expressed in seed (10 DPA); 1,278 pairs biased in the same direction in petal and seed (Figure 2A tissue overlap in G. raimondii)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper compares RNA-seq-derived expression levels between ~2,000 pairs of ancient paralogous genes in two Gossypium species across three tissues (and a seed developmental time series), using three biological replicates per tissue/time point. Differential expression between paralogs was assessed with Student's t-tests on log ratios of normalized expression (RPKM and, separately, upper-quartile normalization), with results corrected for a 5% false discovery rate. A generalized linear model (negative binomial, fit in R) was additionally used to partition gene, tissue, and gene-by-tissue interaction effects, and a Wilcoxon signed-rank test compared dN/dS ratios between resulting gene groups; results were reported primarily as counts and percentages of paralog pairs meeting significance and fold-change thresholds.

Replicationbiological Sample sizethree biological replicates per tissue and/or time point for both G. raimondii and G. arboreum Groupsparalogous gene pairs compared to each other, across three tissues (petal, leaf, seed), a seed developmental time series, and two species Pairingpaired Randomization/blindingnot stated Dispersionunclear Effect sizesyes Multiplicity correctionBenjamini-Hochberg (1995) false discovery rate correction, 5% FDR
Statistical tests used
Test Applied to n Assumptions
Student's t-test on log ratio of paralog expression (RPKM-normalized, and separately UQ-normalized) differential expression between paralogous gene pairs within and between tissues/time points, G. raimondii and G. arboreum three biological replicates per tissue/time point stated
Generalized linear model, negative binomial distribution (R, contrasts package) gene (G), tissue (T), and gene-by-tissue (G×T) interaction effects, and per-tissue (G|T) contrasts, on UQ-normalized RNA-seq data in petal, seed, and leaf of G. raimondii three biological replicates per tissue not stated
Binomial test whether paralogs at a given chromosomal position were more often over- or underexpressed relative to their duplicate counterpart not stated
Wilcoxon signed-rank test comparison of mean dN/dS ratios between paralog groups defined by G, T, and G×T effect patterns not stated
Approaches that could also have been used
  • Differential expression between paralogs was assessed with Student's t-tests on log ratios of RPKM/UQ-normalized values, with normality checked by visual inspection of the log-ratio distribution.
    Could also: RNA-seq-specific count-based frameworks such as DESeq2 or edgeR, which model raw counts with a negative binomial distribution and shrink dispersion estimates across genes — these tools are built to handle the mean-variance structure and overdispersion typical of RNA-seq counts directly, which can be useful with a small number of replicates, without relying on log-ratio normality.
  • Gene, tissue, and gene-by-tissue interaction effects were estimated with a negative binomial GLM implemented in base R.
    Could also: established RNA-seq differential expression packages (e.g., edgeR glmQLFit, DESeq2 likelihood ratio test) that implement similar negative-binomial GLM frameworks with additional empirical Bayes dispersion shrinkage — dispersion shrinkage across genes can stabilize variance estimates and improve power when replicate numbers are limited.
  • Multiple testing was controlled using the Benjamini-Hochberg FDR procedure at a 5% threshold, applied separately within each test family.
    Could also: Storey's q-value approach, which adaptively estimates the proportion of true null hypotheses — q-value methods can offer somewhat greater power than the standard BH step-up procedure while still controlling the false discovery rate, particularly when many hypotheses are truly non-null.
  • Differential expression between paralog pairs was evaluated using a t-test on the log ratio, treating each paralog pair as a matched comparison.
    Could also: a paired nonparametric alternative such as the Wilcoxon signed-rank test on the same log ratios — a rank-based paired test avoids reliance on the log-ratio data approximating a normal distribution and can be a natural complement when normality is inspected only visually.
  • Differences in dN/dS ratios among paralog groups (defined by significant G, T, and G×T effects) were tested with a single Wilcoxon signed-rank test.
    Could also: a Kruskal-Wallis test (with post-hoc pairwise comparisons) when more than two groups are being compared — Kruskal-Wallis extends the same nonparametric rank-based logic to simultaneous comparison of multiple groups while controlling the overall test-wise error rate across the group set.
  • Replicate numbers (three biological replicates per tissue/time point) were used without describing a formal power calculation.
    Could also: an a priori power analysis based on expected RNA-seq count variance to justify replicate number — explicit power calculations can help communicate the expected sensitivity to detect a given fold-change in expression given the chosen replicate depth.
Software: R (GLM with negative binomial distribution; contrasts package) · GSNAP (read mapping) · samtools · sickle (quality trimming) · ClustalW (alignment) and custom BioPerl scripts (dN/dS via Jukes-Cantor model) · Circos (visualization)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

sickle_trim_SRR530667
Reported
{'raw_reads': None, 'trimmed_reads': None, 'discarded_reads': None, 'note': 'Paper describes the trimming method qualitatively only; no per-accession numeric read counts are published in Methods/Results/Supplement.'}
Reproduced
{'raw_reads': 8158340, 'trimmed_reads': 8019202, 'discarded_reads': 139138}
partial
gsnap_align_SRR530667
Reported
{'mapping_rate_pct': None, 'queries_aligned': None, 'note': 'Paper states reads were mapped with GSNAP allowing splice-junction mapping, but publishes no mapping-rate statistic or aligned-read count.'}
Reproduced
{'queries_aligned': 8019202, 'runtime_seconds': 809.26, 'total_sam_alignment_records': 11566451, 'mapped_records': 10972040, 'unmapped_records': 594411, 'mapped_fraction_records': 0.9486, 'primary_mapped_reads': 7424791, 'primary_mapped_fraction_of_input_reads': 0.9262}
partial
gene_rpkm_SRR530667_approx
Reported
{'homoeolog_resolved_rpkm': None, 'note': "The paper's headline expression-quantification result is homoeolog-resolved (A-subgenome vs D-subgenome copies within allotetraploid G. hirsutum) RPKM and UQ-normalized values per gene pair, further tested by Student's t-test / GLM negative-binomial with Benjamini-Hochberg FDR 5%. No single-genome, non-homoeolog-resolved RPKM numbers are published to compare against, and the homoeolog-resolved values themselves are only available as full supplementary data tables, not as summary statistics in the text."}
Reproduced
{'genes_total': 38208, 'genes_with_nonzero_counts': 30003, 'fraction_genes_expressed': 0.7853, 'primary_mapped_reads_total': 7424791, 'reads_assigned_to_genes': 6880176, 'fraction_reads_assigned': 0.9266}
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.