Polyploidy and the petal transcriptome of Gossypium.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
This is a PARTIAL reproduction. In scope: sickle quality trimming (najoshi/sickle, the paper's cited repo) and GSNAP alignment to the G. raimondii D5 reference, applied to all 6 publicly available SRA runs. The trimming step's OUTPUT matches the paper's Table 1 exactly for all 6 samples (but the raw->trimmed transformation itself is untestable since raw reads were never deposited -- only post-trim reads are public). GSNAP alignment succeeded for all 6 samples with plausible high mapping rates (83-97% primary mapped), but these are systematically 10-13 percentage points higher than the paper's Table 2 'Mapped %' (73-85%) for every sample -- most likely because Table 2's metric is actually the fraction of reads PolyCat could categorize via homeoSNP overlap (A+D+X+N reads), not the raw GSNAP alignment rate, and PolyCat itself was not run in this pass (third-party tool, separate implementation effort beyond the cited sickle repo). A gene-coverage stretch check (80% of genes have >=1 mapped read, per paper text) reproduced the same qualitative finding (the vast majority of annotated genes are covered) but not the exact figure: 99.63% vs the paper's stated 80%, plausibly for the same reason (no unique-mapping filter applied, and reads pooled across all 6 heterogeneous genotypes onto a single D5 reference rather than per-homeolog assessment as in the original pipeline). NOT attempted: PolyCat categorization itself, RPKM/UPC/EdgeR expression quantification, Blast2GO/KEGG functional analysis -- all out of scope as separate downstream tools/analyses beyond the cited repo and this pass's pipeline-reproduction focus. Dataset profiling of SRP028270 uncovered two significant, independently-discovered data-quality issues: (1) SRA metadata mislabels 5 of 6 runs (all labeled 'D5' when only 1 actually is D5, established via exact read-count cross-match to paper Table 1), and (2) only 6 of the 18 biological-replicate-level runs described in Methods are publicly deposited (1 of 3 replicates per genotype). Both materially affect anyone reusing this dataset without independently cross-checking against the publication.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-31
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusHow are duplicated homoeologous genes (A_T and D_T copies) deployed in the petal transcriptome of allopolyploid Gossypium, and does homoeolog expression bias and expression level dominance arise at genomic merger or evolve after polyploidization? The study uses RNA-seq against the G. raimondii reference to quantify homoeolog-specific expression in diploids, a diploid F1 hybrid, and natural tetraploids.
- ★ Most homoeologous gene pairs in polyploid cotton petals are expressed at equal levels, indicating a surprising level of expression homeostasis; only ~20% of expressed genes show significant genome bias. finding
- ★ The direction of homoeolog expression bias is conserved between natural polyploids and a recently formed diploid F1 hybrid, suggesting cis-acting modifications that arose in the diploid progenitors prior to polyploidization. mechanism
- ★ A small conserved set of 478 homoeologous gene pairs (~4% of commonly expressed genes) is biased in every polyploid accession and shows a significantly larger degree of bias than other biased pairs. finding
- ★ Comparisons between diploids and tetraploids suggest different regulatory mechanisms of gene expression, described as three phases in the evolution of cotton genomes contributing to expression in the polyploid nucleus. mechanism
- ★ RNA-seq reads assigned to homoeologs with PolyCat using a 24 M homoeo-SNP index against the G. raimondii genome provides a more accurate assessment of transcriptome composition than microarrays or EST assemblies, which suffer from probe specificity, ascertainment bias, and cross-hybridization. method
- A gene-expression-based phylogeny of the accessions recovers branching patterns matching accepted genetic relationships, with two main branches corresponding to the A_T- and D_T-genomes. finding
- Petal tissue expresses roughly 45-50% of the genome, lower than fiber tissue (75-90%), consistent with greater canalization of petal development. finding
- The petal transcriptome resource comprises 11,469 commonly expressed genes with GO/KEGG annotation (4,565 enzyme-coding genes, 654 enzymes, 93 pathways), deposited at NCBI SRA (SRP028270 / SRX328344). resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA-seq (50 bp single-end), 3 biological replicates per accession | Petal tissue at full expansion from Gossypium arboreum (A2, AKA8401), G. raimondii (D5, GN33), G. hirsutum cv. Acala Maxxa (AD1), G. hirsutum cv. Tx2094 (AD1), G. tomentosum (AD3, WT936), and a sterile diploid A2 x D5 F1 hybrid | none (natural genotype/ploidy comparison: diploids vs. diploid F1 hybrid vs. natural tetraploids) | Transcript abundance as raw read counts converted to RPKM | Illumina TruSeq RNA library prep kit; Illumina HiSeq v.2 chemistry; Covaris sonication (200-400 bp); Agilent Bioanalyzer and Ribogreen (Invitrogen) for RNA QC |
| Read mapping and homoeolog (A_T vs D_T) read categorization | RNA-seq reads from all six Gossypium accessions mapped to the 13 pseudo-molecules of the G. raimondii D-genome reference (v. 2.2.1) | none | Counts of A-, D-, X- (chimeric) and N- (uncategorized) reads and percent mapped per accession | GSNAP aligner; PolyCat with an index of 24 M homoeo-SNPs; sickle trimming at phred threshold 20 |
| Universal Probability of expression Codes (UPC) expression calling | Petal transcriptomes of each Gossypium accession | none | Probability that a gene is actively expressed; number/percent of expressed genes and 'commonly expressed' genes shared across accessions | UPC mixed-model method |
| Functional annotation (BLASTX, GO assignment, KEGG/Enzyme Code pathway mapping) | 11,469 commonly expressed petal genes of Gossypium | none | GO term distribution across CC/BP/MP ontologies; enzyme codes, enzyme counts, KEGG pathway membership and transcript abundance per pathway | Blast2GO with RefSeq BLAST hits; BLASTX run on the Fulton Supercomputer at BYU; KEGG |
| Differential expression analysis (exact test, negative binomial) | Petal RNA-seq of diploid F1 hybrid, G. tomentosum, G. hirsutum Tx2094, G. hirsutum Maxxa | none | Genes differentially expressed between accessions (nested interaction design with factors 'accession' (4 levels) and 'genome' (2 levels)) and between the two genomes within an accession (single factor, 8 levels); FDR < 0.05 | EdgeR |
| Homoeolog expression bias analysis | Polyploid and diploid F1-hybrid petal transcriptomes of Gossypium | none | Number and percentage of genes with significant A_T/D_T bias, direction of bias, A_T:D_T bias ratio, chromosomal distribution of biased genes | — |
| Expression-based phylogeny (neighbor joining) | Commonly expressed genes across all six Gossypium accessions (A_T and D_T partitions) | none | Distance matrix of sum of squared differences of expression levels between accessions; second tree from differences between homoeologous expression levels | PHYLIP (neighbor joining) |
| Expression level dominance analysis | Each polyploid Gossypium accession compared against diploid parents A2 (G. arboreum) and D5 (G. raimondii) | none | Per-gene categorization from RPKM comparisons of A2 and D5 to total polyploid expression; matrix counts per dominance category (UPC-nonexpressed genes excluded) | — |
- – Approximately 20% of genes expressed in petals showed significant bias toward the A_T- or D_T-genome; the large majority of homoeolog pairs were unbiased (percent biased: 13.3% diploid F1-hybrid, 27.0% G. tomentosum, 20.6% Maxxa, 19.8% Tx2094) ~20% biased; bias ratios A_T/D_T of 0.99, 1.04, 0.98, 1.01 across accessions
- ▲ 478 homoeologous gene pairs (~4% of commonly expressed genes) were biased in every polyploid petal transcriptome, and their degree of bias was significantly greater than that of the remaining biased pairs 1.3-1.6 fold greater on average, p < 0.001
- – Of the 478 conserved biased pairs, 246 were consistently biased in the A_T direction and 202 in the D_T direction, distributed across all thirteen D-genome chromosomes 246 A_T vs 202 D_T of 478 pairs
- – Direction of expression bias among the 770 gene pairs biased in the natural polyploids was conserved for all but 1-2 pairs in any pairwise comparison, indicating conserved cis-regulation; 19 gene pairs in the diploid F1-hybrid had a contrary direction of bias relative to the natural polyploids (plus 6 unique in G. tomentosum, 2 in Maxxa, 1 in Tx2094) all but 1-2 of 770 pairs conserved
- ▼ Bias differed between the diploid F1-hybrid and natural polyploids for 1,009 genes (339 biased only in the F1-hybrid, 770 biased only in the natural polyploids); the F1-hybrid had the smallest proportion of expressed genes with detectable bias 339 + 770 = 1,009 genes
- – Differential expression between accessions: 692 genes between the two G. hirsutum accessions, 1,394 between G. hirsutum accessions and G. tomentosum, and 2,671 between the diploid F1-hybrid and the natural polyploids; the F1-hybrid's total expression more closely resembled the diploid species 692 / 1,394 / 2,671 genes (FDR < 0.05)
- ▼ Of 37,223 annotated reference genes, 11,469 were commonly expressed in petals of all polyploid accessions (~80% of the expressed genes in each accession); ~45-50% of the genome is expressed in petals versus 75-90% in fiber 45-50% (petal) vs 75-90% (fiber)
- – Read categorization: ~25% of mapped reads were assigned to each of the A_T and D_T genomes, with 78.1% overall mapping and 88.9 M reads uncategorized (N-reads) due to limited A_T/D_T coding-sequence divergence; 80% of annotated reference genes had at least one mapped petal read 78.1% mapped; 46.8 M A-reads, 46.9 M D-reads, 4.1 M X-reads, 88.9 M N-reads of 188.2 M total
- count 11,469 commonly expressed genes of 37,223 annotated genes in the reference D-genome (Genes expressed in petals of all polyploid accessions by UPC; ~80% of each accession's expressed genes)
- fold_change 1.3 - 1.6 fold greater degree of bias on average (478 conserved biased homoeolog pairs vs. remaining biased pairs)
- pvalue p < 0.001 (Greater degree of bias for the 478 conserved biased pairs (e.g. vs the remaining 814 genes in the diploid F1-hybrid))
- pvalue p < 0.005 (Correlation between total genes per chromosome and biased genes per chromosome, arguing against a chromosomal-location effect)
- count 692; 1,394; 2,671 differentially expressed genes (Maxxa vs Tx2094; G. hirsutum accessions vs G. tomentosum; diploid F1-hybrid vs natural polyploids (EdgeR, FDR < 0.05))
- count 2,060 biased genes in the diploid F1-hybrid, of which 1,292 were expressed in all accessions; 3,706 in G. tomentosum; 3,146 in Maxxa; 2,686 in Tx2094 (Genes with homoeologous expression bias per accession)
- other 4,565 genes assigned an enzyme code, corresponding to 654 enzymes across 93 enzymatic pathways (Blast2GO/KEGG annotation of commonly expressed petal genes)
- other GO assignment: cellular component 88%, biological process 17%, molecular process 9%; cytoplasm 28% and cytoplasmic part 27% (CC); cellular protein metabolic process 31% (BP); kinase activity 41% (MP) (Gene ontology distribution of commonly expressed petal genes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used RNA-seq (three biological replicates per accession) to quantify homoeolog (A-genome/D-genome) gene expression in diploid and tetraploid Gossypium petals, with reads categorized to genome-of-origin via homoeo-SNPs. Differential expression between accessions and between genomes was tested with edgeR's exact test for the negative binomial distribution within a factorial model design (accession x genome), using an FDR (q-value) < 0.05 threshold to call genes differentially expressed. Gene expression detection/presence calls used a mixed-model approach (UPC), and a small number of additional comparisons (e.g., degree-of-bias differences, chromosome-level correlation) were reported with threshold p-values. Results were reported mainly as gene counts/percentages, fold-differences, and p-value or FDR thresholds rather than exact p-values or interval estimates.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| edgeR exact test for the negative binomial distribution (nested interaction model: accession x genome) | Differential expression between the four accessions (diploid F1-hybrid, G. tomentosum, G. hirsutum Maxxa, G. hirsutum Tx2094) | 3 biological replicates per accession | not stated |
| edgeR exact test for the negative binomial distribution (single-factor, 8-level design) | Differential expression between A-genome (A_T) and D-genome (D_T) homoeologs within each accession | 3 biological replicates per accession | not stated |
| UPC (Universal Probability of expression Codes), mixed-model approach | Determining which genes were actively/detectably expressed in petal tissue per accession | 3 biological replicates per accession | not stated |
| Unspecified statistical test comparing degree of expression bias between gene sets (reported as p < 0.001) | Comparison of the 478 'conserved' biased gene pairs vs. remaining biased gene pairs (e.g., 814 in the diploid F1-hybrid) | 478 vs. remaining biased gene pairs (varies by accession) | not stated |
| Correlation test (type unspecified, reported as p < 0.005) | Relationship between total genes per chromosome and number of biased genes per chromosome | 13 D-genome chromosomes | not stated |
-
Differential expression was assessed with edgeR's exact test for the negative binomial distribution, applied within a factorial (accession x genome) model design.↳ Could also: edgeR's generalized linear model with quasi-likelihood F-test (glmQLFit/glmQLFTest), or DESeq2's Wald/likelihood-ratio test — These GLM-based approaches can directly model multi-factor designs (e.g., accession and genome as combined factors) with a single fitted model and often give more conservative, robust error control at low replicate numbers, which some researchers prefer for multi-factor RNA-seq designs.
-
Genes were called differentially expressed using edgeR's FDR (q-value) < 0.05 threshold, applied separately per contrast.↳ Could also: A single joint multiple-testing correction (e.g., Benjamini-Hochberg) applied across the combined set of all contrasts performed in the study — Applying one correction across the full family of tests performed (rather than per-contrast) is another standard way to control the overall false discovery rate when several separate comparisons are made in the same study.
-
Active gene expression was called using UPC, a mixed-model probability-of-expression approach.↳ Could also: A simple count/abundance threshold (e.g., requiring a minimum CPM or RPKM in at least a set number of replicates) — Threshold-based expression calls are a widely used, simpler alternative that some researchers prefer for its transparency and ease of reproducibility, though UPC's probabilistic approach can better account for background noise.
-
The difference in degree of expression bias between the 478 'conserved' biased gene pairs and the remaining biased gene pairs was reported as a fold-difference with p < 0.001, without specifying the test used.↳ Could also: A nonparametric Mann-Whitney U (Wilcoxon rank-sum) test — Because bias-ratio/fold-change data are often right-skewed, a rank-based test is a standard alternative that does not assume normality and can also be used to compare two distributions of ratio values.
-
The relationship between total genes per chromosome and number of biased genes per chromosome was assessed with a correlation test (p < 0.005), based on 13 chromosomes.↳ Could also: A permutation-based (randomization) test of the observed chromosomal distribution of biased genes against a null expectation — With only 13 data points (chromosomes), a permutation approach is a standard alternative that does not rely on asymptotic assumptions of a parametric correlation test and can be well suited to small-n count data.
-
Expression-level and bias results were summarized mainly as point estimates (percentages, fold-differences, ratios) without an explicitly reported dispersion measure.↳ Could also: Reporting standard deviation, standard error, or a 95% confidence interval alongside the point estimates (e.g., for average fold-difference in bias) — Adding a dispersion or interval estimate is a standard complement to point estimates and conveys the precision/spread of the estimate in addition to its central value.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
What deviates: Table 1's six 'Trimmed Reads' values reproduce exactly (e.g. D5 = 39,974,015 to the read), but every Table 2 'Mapped %' is 10-13 points above the published figure (D5 84.70% -> 97.19%, A2 73.30% -> 83.29%, F1 78.80% -> 89.61%, Maxxa 77.70% -> 88.69%, AD3 77.60% -> 88.55%, Tx2094 76.60% -> 88.55%), and the '80% of genes covered' statement reproduces as 99.63% (37,085/37,223). Whose side: the offset is uniform in sign and size across all six samples, which points squarely at our own scope decision — we ran GSNAP alone and skipped PolyCat's homeoSNP categorization, the step that actually generates Table 2 — not at an authors' error. Authors' side is nonetheless implicated on data availability: only 6 of 18 described libraries are deposited, raw pre-trim reads are missing entirely, and 5 of 6 SRA runs are mislabeled as 'D5' under one sample accession, a genuine reuse hazard we discovered independently. Severity: moderate and explainable — no significance flip, no direction reversal, rank ordering across genotypes preserved, and no fabrication signal; the paper's actual central claims about homeolog expression under polyploidy were simply never put to the test in this pass, so they remain untested rather than unconfirmed.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.