RNA-Seq transcriptome profiling of upland cotton (Gossypium hirsutum L.) root tissue under water-deficit stress.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Different-but-partial: the raw sequencing data (single pooled SRA run, 109,596,793 reads) exactly matches the paper's stated combined read total, and the described pipeline steps (Sickle Q20 trimming, GSNAP -N1 genomic alignment) were executed faithfully and completed cleanly (81.1% mapping rate; 55,009/78,371 annotated transcripts received >=1 read). However the paper's HEADLINE result -- 1,530 genes differentially expressed between water-deficit and well-watered treatments -- cannot be reproduced from public data: NCBI SRA/BioProject PRJNA210770 deposited only ONE pooled run representing all six barcoded libraries combined, with no demultiplexing key, per-sample split files, or supplementary DE table ever made public. This is a genuine data-availability gap in the original deposit, not a pipeline failure on our end. Consequently PolyCat subgenome localization (Table 3), the DE heatmap/PCA (Figures 1-4), and RT-qPCR concordance (Figure 5) were not attempted (unreproducible or wet-lab/out of scope). The transcript-count comparison (55,009 vs paper's 33,930) is graded partial/order-of-magnitude because the annotation version used here (current NCBI, 78,371 transcript models) differs from the paper's 2013-vintage Phytozome v2.1 annotation.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan next-generation Illumina RNA-seq analysis, anchored to the newly published diploid Gossypium raimondii (D5) whole genome sequence, be used to measure global gene expression profiles in field-grown tetraploid upland cotton root tissue under water-deficit stress in order to identify candidate genes for molecular cotton breeding? The study also asks whether differentially expressed transcripts can be putatively assigned to the AT or DT subgenomes of allotetraploid G. hirsutum.
- ★ A total of 1,530 transcripts were differentially expressed between well-watered and water-deficit stressed field-grown upland cotton root tissues (913 up-regulated, 617 down-regulated). finding
- ★ This is the first application of the newly published diploid D5 G. raimondii genome sequence to RNA-seq transcriptome analysis of tetraploid AD1 upland cotton. method
- ★ Putative sequence-based genome localization using the PolyCat pipeline detected A-genome (AT) specific gene expression under water-deficit stress, with up-regulated genes predominately containing AT-specific reads. finding
- ★ Down-regulated transcripts were more evenly distributed between AT and DT subgenome assignment than up-regulated transcripts. finding
- ★ RNA-seq recovered transcripts previously identified by cDNA-AFLP in the same experimental system (water uptake, heat stress and carbohydrate metabolism genes), confirming accuracy of the technique for future cotton genomics studies. finding
- Differentially expressed genes were distributed across all 13 chromosomes of the diploid progenitor G. raimondii genome. finding
- Several genes and major biochemical pathways were up-regulated in root tissue under water-deficit stress. mechanism
- ★ The dataset is a resource for identifying candidate genes benefiting applied plant breeding programs (NCBI SRA Accession PRJNA210770; also to be made available through CottonGen). resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (mRNA-seq, single-end 50 bp) | Gossypium hirsutum cultivar 'Siokra L-23' root tissue, field-grown at North Carolina State University Sandhills Research Station, Jackson Springs, NC; 3 individual plants per treatment (6 barcoded libraries) | water-deficit stress (naturally rain-fed) vs. well-watered irrigation, field conditions | mapped read counts per transcript; differentially expressed transcripts between treatments | Illumina RNA TruSeq kit; single lane of 50 bp Illumina HiSeq 2000; library QC on Agilent Bioanalyzer 2100 with qPCR quantification |
| leaf water potential measurement (physiological phenotyping) | uppermost fully expanded leaves of field-grown G. hirsutum 'Siokra L-23'; 3 plants per treatment | water-deficit vs. well-watered | xylem/leaf water potential in MPa (stressed defined as -2.0 MPa or greater; well-watered -1.9 MPa or lower) | pressure bomb, Model 600, PMS Instrument Company, Albany, OR |
| RT-qPCR validation of RNA-seq | root tissue from additional G. hirsutum 'Siokra L-23' plants grown in the same plots and experimental conditions | water-deficit vs. well-watered | Relative Expression Ratios (RER) by ΔCt method for ten differentially expressed target genes, normalized to internal reference transcript Gorai.012G141300 (validated with RefFinder); TUA11 (Gorai.010G125700) used as no-RT DNA contamination control | Maxima SYBR Green/ROX qPCR Master Mix (2X, Fermentas/Thermo); Bio-Rad iCycler Real Time PCR Detection System, two-step amplification plus melt; RNA via Sigma-Aldrich Spectrum Plant Total RNA kit with On-Column DNase I; NanoDrop; Bioanalyzer 2100; cDNA via Invitrogen SuperScript III First-Strand Synthesis SuperMix |
| read trimming and SNP-tolerant reference genome alignment | reads from six G. hirsutum root libraries mapped to G. raimondii 2.1 whole genome reference (33,930 transcripts) | none (computational) | trimmed/mapped read counts; novel splice site identification | Sickle (quality cutoff 20); GSNAP with '-N 1' and a SNP index from deep coverage of G. arboreum and G. raimondii |
| differential expression and data quality analysis | mapped read count matrix from six cotton root libraries | water-deficit vs. well-watered contrast | significantly differentially expressed transcripts at 5% FDR; Euclidean distance heatmap and principal component analysis of samples | DESeq version 1.9.12 in R |
| subgenome read categorization / genome localization | NGS reads of allotetraploid G. hirsutum (AD1) assigned to progenitor diploid genomes G. arboreum (A2) and G. raimondii (D5) | none (computational) | counts and percentages of differentially expressed transcripts whose reads map predominantly or exclusively to AT vs. DT genomes | PolyCat annotation/read mapping pipeline |
| functional annotation, GO enrichment and pathway mapping | 1,530 significant transcripts plus splice variants (2,942 sequences) from G. raimondii v2.1 / Phytozome | none (computational) | number of successfully annotated sequences, added and confirmed annotations, enriched GO terms, KEGG pathway assignments | Blast2GO (with ANNEX and validation step), InterProScan, AgriGO, Phytozome KEGG Orthology IDs with KEGG 'Search and Color' Pathway tool against reference pathway (KO) |
| amplicon cloning and Sanger sequencing for primer specificity confirmation | purified RT-qPCR PCR amplicons from G. hirsutum root cDNA/DNA | none | insert presence/orientation (T3/T7 PCR) and sequence identity of amplicons by BLAST homology; four colonies bi-directionally sequenced per amplicon | Invitrogen TOPO Zero Blunt or TA Cloning Systems with OneShot Top10 competent cells; Promega Wizard SV Gel and PCR Clean Up System; Invitrogen Qubit dsDNA HS Assay Kit; Geneious version 6.1 |
- – 1,530 genes were differentially expressed between water-deficit and well-watered root samples at FDR 0.05 1530 genes (913 up, 617 down)
- ▲ Up-regulated transcripts predominately contained AT genome specific reads in both water-deficit and well-watered comparisons 407 of 913 transcripts (44.6%)
- ▼ Down-regulated transcripts were more evenly split between AT and DT subgenome read assignment AT 225 (36.5%) vs. DT 217 (35.4%)
- – Very few differentially expressed transcripts were exclusively subgenome-specific up-regulated: 2 (0.2%) AT-only, 3 (0.3%) DT-only; down-regulated: 5 (0.8%) AT-only, 5 DT-only
- – 101 differentially expressed transcripts could not be assigned to a specific genome within tetraploid cotton 101 transcripts
- – Approximately 109.6 million 50 bp reads from six libraries were trimmed and mapped to 33,930 G. raimondii transcripts 109.6 million reads; 33,930 transcripts
- – Of 2,942 significant transcripts and splice variants, 2,416 were successfully annotated; 112 additional annotations were added and 1,821 confirmed after InterProScan/ANNEX enhancement 2416/2942 annotated; 102 exceeded >8000 bp size limit; 74 had no BLAST homology
- – RNA-seq detected water-deficit responsive transcripts overlapping a prior cDNA-AFLP study, including aquaporin PIP1;3, Heat Shock Protein 26, and mannose-6-phosphate isomerase 304 DE genes in cDNA-AFLP vs. 1530 in RNA-seq
- count 1530 differentially expressed genes (913 up-regulated, 617 down-regulated) (DESeq analysis, water-deficit vs. well-watered root libraries, FDR 0.05)
- other FDR = 5% (0.05) (significance threshold for differential expression in DESeq 1.9.12)
- count ~109.6 million 50 bp reads (total trimmed reads across all six RNA-seq libraries)
- count 33,930 transcripts (G. raimondii 2.1 reference transcripts to which reads were mapped)
- mean well-watered -1.60, -1.35, -1.45 MPa; water-deficit -2.20, -2.70, -2.85 MPa (leaf water potential of the three plants per treatment used for RNA-seq (Table 1))
- count 407 (44.6%) AT-majority up-regulated transcripts; 225 (36.5%) AT and 217 (35.4%) DT down-regulated transcripts (PolyCat putative subgenome localization of differentially expressed transcripts (Table 3))
- count 2416 of 2942 sequences annotated; 112 annotations added; 1821 annotations confirmed (Blast2GO plus InterProScan/ANNEX functional annotation of significant transcripts and splice variants)
- other >90% of transcripts had 0–1000 mapped reads; 50% had fewer than 100 reads; 7% had more than 1000 mapped reads (distribution of mapped read depth across identified transcripts)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used an RNA-seq design comparing well-watered and water-deficit cotton root tissue, with three biological replicates (individual plants) per treatment sequenced as six barcoded Illumina libraries. Differential expression between treatments was tested using the DESeq package (version 1.9.12) in R, applying a 5% false discovery rate to call 1,530 significantly up- or down-regulated transcripts. Functional characterization used Blast2GO/InterProScan annotation and AgriGO gene ontology enrichment, and ten of the differentially expressed genes were validated by RT-qPCR using the ΔCt relative expression ratio method with a reference gene chosen via RefFinder. Results were reported mainly as counts of transcripts passing the FDR threshold, genome-of-origin proportions from PolyCat, and RT-qPCR expression ratios, rather than as individual p-values or fold-change magnitudes in the main text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq (v1.9.12) count-based differential expression test with 5% FDR threshold | Differential expression between water-deficit and well-watered root transcripts | 3 biological replicates (individual plants) per treatment, 6 libraries total | not stated |
| Gene ontology enrichment analysis (AgriGO) | Functional/GO enrichment of differentially expressed transcripts | — | not stated |
| ΔCt relative expression ratio method | RT-qPCR validation of 10 target genes plus reference gene TUA11 | duplicate reactions from two independent cDNA synthesis reactions per sample | not stated |
-
Differential expression was assessed using DESeq (version 1.9.12) with raw mapped read counts and a single 5% FDR threshold for calling significance.↳ Could also: A more recent count-based DE tool such as DESeq2 or edgeR, which incorporate shrinkage estimation of dispersion and effect size — Shrinkage-based dispersion and log2 fold-change estimates can improve stability of significance calls with small sample sizes and let readers evaluate effect size alongside FDR.
-
Three biological replicates per treatment were used for RNA-seq, without a stated power or sample-size justification.↳ Could also: Reporting a formal power calculation or a minimum detectable fold-change alongside the replicate count — This would give readers additional context on the sensitivity of the design to detect smaller expression differences.
-
Differential expression results were summarized mainly as counts of transcripts passing the FDR threshold (1,530 total), without per-gene fold-change magnitudes or exact adjusted p-values presented in the main text.↳ Could also: A volcano plot or supplementary table listing log2 fold-change and adjusted p-value for each transcript — Pairing effect size with statistical significance for each transcript gives readers a fuller view of both the magnitude and confidence of expression changes.
-
RT-qPCR validation used the ΔCt method normalized to a single reference gene (TUA11) selected via RefFinder.↳ Could also: The ΔΔCt method with multiple validated reference genes, plus a formal concordance statistic (e.g., Pearson or Spearman correlation) between RNA-seq and RT-qPCR fold-changes — Using multiple reference genes can reduce sensitivity to expression variability in any single reference, and a correlation statistic quantifies platform agreement numerically rather than qualitatively.
-
GO term enrichment was performed with AgriGO without the underlying statistical test or multiple-testing correction being described in the text.↳ Could also: Explicitly reporting the enrichment test (e.g., Fisher's exact test) and a correction method (e.g., Bonferroni or FDR) applied across GO terms — Stating the specific test and correction clarifies how significance was determined across the many GO categories evaluated simultaneously.
-
Leaf water potential values used to classify treatments (Table 1) are presented as individual plant measurements without a formal statistical comparison between groups.↳ Could also: A two-sample t-test or non-parametric Mann-Whitney U test comparing water potential between well-watered and water-deficit plants — A formal test would provide a quantitative measure of separation between the two water-status groups used to define the experimental treatments.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
What deviates: Only one quantitative discrepancy was actually measurable — 55,009 transcripts with >=1 mapped read versus the paper's 33,930 — and it is cleanly explained by annotation vintage (current NCBI 2021 set, 78,371 transcript models, vs the paper's 2013 Phytozome G. raimondii v2.1). The read-volume claim reproduced exactly (109,596,793 = reported ~109.6M) and Sickle Q20 / GSNAP -N1 ran faithfully (98.78% reads kept, 81.11% mapped). Whose side: The blocking defect is on the authors'/deposit side — PRJNA210770 contains a single pooled run representing all six barcoded libraries with no demultiplexing key, per-sample FASTQs, count matrix or supplementary DE table, so the headline 1,530-gene DESeq result and the dependent PolyCat Table 3 are structurally uncomputable from public data. Severity: Moderate. Nothing in the paper was contradicted and there is no fabrication signal; the central conclusion is simply unverifiable, which warrants yellow on q5/q7/q8 and red on q4 (cause lies with incomplete public deposition), not a red criticality driven by a demonstrated discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.