Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Transcriptomic and physiological analysis of atractylodes chinensis in response to drought stress reveals the putative genes related to sesquiterpenoid biosynth

BMC Plant Biol · 2024
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Re-run of a previously requeued room (the prior «infra» workdir had been reclaimed by the janitor). De novo plant RNA-seq drought study (Atractylodes chinensis rhizome, CK/D3/D9 x3); the metadata 'Code' link is the third-party SeqPrep tool, applied with Sickle + Trinity to the paper's own public PRJNA930596 data (valid per P16). RESULT = partial, described well enough to attempt, no fabrication signal. Claim A (raw reads/group) is an EXACT zero-compute match to ENA spot counts. Claim B (clean bases per group + Q30) reproduces WITHIN-TOL via the paper's named SeqPrep+Sickle defaults run on «our HPC»: 6.261/6.279/6.309 Gb vs 6.28/6.30/6.33 (all <0.3% low) and Q30 94.83% vs 94.59% -- the deposited reads + named tools regenerate Table 1's clean-data columns. Claim C (Trinity assembly stats) could NOT be completed: the exact pipeline (Trinity 2.15.1 defaults) ran correctly and finished Phase 1 (Inchworm/Chrysalis -> 682,761 cluster commands), but Phase 2 (Butterfly) runs at ~0.47 cmd/s (~396h projected) and cannot finish within the cluster's 12h/job wall -- a compute-budget blocker on our side, and the unigene count is non-deterministic regardless. NOT attempted (all cascade off C): DESeq2 DEG counts, functional-annotation counts, sesquiterpenoid candidate genes; physiology + qRT-PCR are wet-lab (out of scope). Dataset PRJNA930596 profiled: 9/9 PE runs present, all 18 fastq md5-verified against ENA, full parse clean -> grade A, delivers-promised yes.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 67
    assessed: 2026-06-15 ⛓ f010ff352a60
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests how drought stress affects the physiology and transcriptome of Atractylodes chinensis seedlings over time, aiming to identify candidate genes involved in sesquiterpenoid and triterpenoid biosynthesis that respond to drought.

Core claims
  • Drought stress significantly increases MDA, proline, soluble sugar, and crude protein content and antioxidative enzyme (SOD, POD, CAT) activity in A. chinensis seedlings finding
  • Transcriptomic analysis assembled 215,665 unigenes as a resource for A. chinensis drought response resource
  • 29,449 DEGs were identified between control and D3, and 14,538 DEGs between control and D9 finding
  • Terpenoid backbone biosynthesis had the highest number of unigenes among terpenoid and polyketide metabolism pathways under drought stress finding
  • 22 unigenes encoding enzymes in the terpenoid backbone biosynthetic pathway and 15 unigenes encoding enzymes in the sesquiterpenoid/triterpenoid biosynthetic pathway were identified as candidate genes under drought stress finding
  • Integration of physiological and transcriptomic analyses reveals the molecular mechanism underlying A. chinensis drought stress response mechanism
  • qRT-PCR was used to verify the expression of 15 selected DEGs involved in sesquiterpenoid and triterpenoid biosynthetic pathways method
  • RNA-seq transcriptome data were deposited in NCBI SRA under accession PRJNA930596 resource
Experimental setups
Assay System Perturbation Readout Platform
Physiological index measurement (MDA, proline, soluble sugar, crude protein content; SOD, POD, CAT enzyme activity) A. chinensis seedlings (whole plant/root) drought stress (0, 3, 9 days) biochemical content and antioxidant enzyme activity colorimetric/enzymatic assay kits (Nanjing Jiangcheng Bioengineering Institute)
Relative water content measurement (soil and plant) A. chinensis seedlings / seedling nutrient media (soil) drought stress (0, 3, 9 days) soil and plant relative water content (%)
Bulk RNA-seq A. chinensis seedlings/rhizome drought stress (0, 3, 9 days) gene expression (TPM), differentially expressed genes Illumina NovaSeq 6000
De novo transcriptome assembly and functional annotation A. chinensis drought stress unigenes annotated against NR, COG, KEGG, Pfam, Swiss-Prot, GO databases Trinity, BLASTX, BLAST2GO
GO and KEGG enrichment analysis A. chinensis DEGs drought stress enriched GO terms and KEGG pathways Goatools, KOBAS
Quantitative real-time PCR (qRT-PCR) A. chinensis seedlings drought stress (0, 3, 9 days) relative expression of 15 selected DEGs (2^-ΔCt method) ABI 7500 Fast Real-Time System
Key results
  • Soil relative water content decreased from 74.40% (control) to 35.54% (D3) and 30.09% (D9)
  • MDA content increased significantly with prolonged drought stress and was more than two times higher in D9 than in control more than 2-fold
  • Soluble sugar content was significantly increased in D3 and D9 compared with control
  • Proline content and SOD activity increased significantly with prolonged drought stress
  • Crude protein content and CAT/POD activity were significantly increased in D9 but showed no significant difference at D3
  • 29,449 DEGs identified between control and D3; 14,538 DEGs identified between control and D9
  • 215,665 unigenes assembled with average length 759.09 bp and N50 of 1140 bp
  • 22 unigenes encoding enzymes identified in terpenoid backbone biosynthetic pathway; 15 unigenes encoding enzymes identified in sesquiterpenoid/triterpenoid biosynthetic pathways
Key statistics
  • other 74.40% to 35.54% (soil relative water content decrease from control to D3)
  • other 30.09% (soil relative water content at D9)
  • fold_change more than two times higher (MDA content in D9 vs control)
  • count 29,449 (DEGs between control and D3)
  • count 14,538 (DEGs between control and D9)
  • count 215,665 (total assembled unigenes)
  • other N50 = 1,140 bp; average length = 759.09 bp (unigene assembly quality metrics)
  • pvalue P < 0.05 (significance threshold for physiological index differences among control, D3, D9)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

A. chinensis seedlings were subjected to three drought-stress durations (0, 3, 9 days; n=3 biological replicates each) and evaluated with eight physiological indices (each replicate pooled from 5 plants) plus de-novo transcriptomics (RNA-seq, 3 replicates per condition). Physiological comparisons used Student's t-test and/or one-way ANOVA (the methods section states t-test; figure legends indicate one-way ANOVA with letters denoting significance). Differential expression between control and each drought treatment was identified with DESeq2 (Padj<0.05, |log2FC|≥1), and GO/KEGG enrichment significance was assessed with Bonferroni correction. Results are reported as means ± SE, with significance thresholds of P<0.05 and P<0.01.

Replicationbiological Sample sizeThree biological replicates per treatment for transcriptomics and physiological assays; each replicate pooled from one plastic plate (~50 seedlings) or five whole plants; five biological replicates used for plant RWC specifically GroupsControl (0 days drought) vs. D3 (3 days) vs. D9 (9 days) Pairingunpaired Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR (DESeq2 default, Padj<0.05) for DEGs; Bonferroni correction for GO/KEGG enrichment; no correction stated for the multiple physiological t-tests
Statistical tests used
Test Applied to n Assumptions
Student's t-test (two-group, implied two-tailed) Physiological indices (MDA, soluble sugar, crude protein, proline, SOD, POD, CAT, soil/plant RWC) — control vs. D3 and control vs. D9 n=3 biological replicates per treatment not stated
One-way ANOVA (with post-hoc letters) All physiological index figures (Fig. 2A–I) — stated in figure legends as basis for significance letters across control, D3, and D9 n=3 biological replicates per treatment not stated
DESeq2 Wald test with BH-adjusted p-value Differential expression: control vs. D3 (29,449 DEGs) and control vs. D9 (14,538 DEGs) n=3 biological replicates per condition not stated
Bonferroni-corrected enrichment test (hypergeometric, via Goatools / KOBAS) GO term and KEGG pathway enrichment of DEGs vs. whole transcriptome background not stated
2^(-ΔCt) relative quantification qRT-PCR validation of 15 selected DEGs in sesquiterpenoid/triterpenoid pathways; UBQ2 as internal reference n=3 biological replicates per condition not stated
Approaches that could also have been used
  • Eight physiological indices were each compared across three groups using t-tests (and/or one-way ANOVA per figure legends), without a stated correction for running eight simultaneous comparisons
    Could also: A one-way ANOVA followed by a Tukey HSD or Dunnett's post-hoc test applied jointly across all indices, combined with a Bonferroni or BH-FDR correction across the eight outcomes, would also be a standard approach — Correcting for the family of comparisons across multiple physiological indices limits the probability of at least one false positive across the set; Dunnett's test is specifically designed for comparing multiple treatment groups against a single control
  • Dispersion around means is reported as SEM throughout (mean ± SE, n=3)
    Could also: SD or a 95% bootstrap confidence interval would also summarize variability — With n=3 per group, SEM is very small relative to the underlying biological spread; SD conveys the actual observed variation in the sample, and a CI makes the uncertainty in the mean estimate explicit — both are often preferred for small-n biological data
  • qRT-PCR relative expression was calculated using the 2^(-ΔCt) method, normalizing to UBQ2 but not to a reference condition
    Could also: The 2^(-ΔΔCt) (Livak) method would also express expression as fold-change relative to the control condition — 2^(-ΔΔCt) directly yields a ratio to the calibrator condition (control), making it straightforward to compare direction and magnitude of change across genes and treatments on a common scale; it is also the conventional form expected by many readers for validation qPCR
  • DESeq2 was used for differential expression analysis with n=3 replicates per condition
    Could also: edgeR (exact negative-binomial test) or limma-voom would also be standard for RNA-seq DE analysis at this sample size — edgeR and limma-voom are widely validated for small-n RNA-seq experiments and are frequently used as comparators or alternatives to DESeq2; reporting concordance across at least two methods is a common practice in transcriptomics studies to strengthen confidence in the DEG list
  • GO/KEGG enrichment used Bonferroni correction, which is the most conservative standard approach
    Could also: Benjamini-Hochberg FDR at a threshold of 0.05 or 0.10 would also be a standard correction for enrichment analyses — Bonferroni controls the family-wise error rate under an assumption of independent tests; for GO/KEGG terms that are often correlated (parent–child relationships), BH-FDR is commonly preferred as it maintains a less conservative false-discovery rate while accounting for the large number of correlated terms
  • Transcriptome assembly was performed de novo with Trinity because no reference genome is available for A. chinensis
    Could also: If a closely related Atractylodes or Asteraceae reference genome becomes available, reference-guided assembly (e.g., HISAT2 + StringTie) would also be applicable — Reference-guided approaches reduce assembly artifacts and improve isoform resolution; reporting assembly quality metrics (e.g., BUSCO completeness score) alongside N50 and average length would also further characterize assembly completeness for readers
Software: SPSS 26 · DESeq2 (R/Bioconductor) · Trinity · RSEM · Goatools · KOBAS

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
17
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

5.3.3.2 Brenda in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
5.4.99.39 Brenda in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
5.4.99.41 Brenda in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
5.4.99.7 Brenda in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
PRJNA930596 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Reproduction scope — pmid-38317086

Paper: Ma et al. 2024, Transcriptomic and physiological analysis of Atractylodes chinensis in response to drought stress reveals the putative genes related to sesquiterpenoid biosynthesis. BMC Plant Biol. DOI 10.1186/s12870-024-04780-8.

Design: de novo RNA-seq of rhizome, 9 libraries = 3 groups × 3 biol. reps: control (CK), 3-day drought (D3), 9-day drought (D9). Illumina NovaSeq 6000, 2×150 bp PE. Data: SRA PRJNA930596 (9 runs SRR23330181–SRR23330189, public).

The "Code" link in the metadata (github.com/jstjohn/SeqPrep) is a third-party read-preprocessing tool, not the authors' own analysis code. Per project rule P16, running that tool (and the rest of the described pipeline) on the paper's own public data is a valid reproduction.

Run → group map (from ENA, library_name)

group runs
CK (control) SRR23330189 (CK-1), SRR23330188 (CK-2), SRR23330187 (CK-3)
D3 (3-day) SRR23330186 (D3-1), SRR23330185 (D3-2), SRR23330184 (D3-3)
D9 (9-day) SRR23330183 (D9-1), SRR23330182 (D9-2), SRR23330181 (D9-3)

IN SCOPE (pipeline-derived) — attempted, in 80/20 order

# reported result location pipeline cost
A Raw reads per group: 43.3 / 44.0 / 44.2 M Table 1 (none — SRA metadata) free
B Clean bases per group: 6.28 / 6.30 / 6.33 Gb; Q30 = 94.59% Table 1 SeqPrep + Sickle (named code), default params light–moderate
C Assembly: 215,665 unigenes, avg 759.09 bp, N50 1140 bp, GC 45.24% Table 1 / Results Trinity (de novo) heavy (stretch)

IN SCOPE but NOT attempted in first pass (hard 20%, cascading + non-deterministic)

  • DEG counts via DESeq2: CK-vs-D3 29,449 (18,932 up / 10,517 down); CK-vs-D9 83,238 (6,908 up / 76,330 down). Downstream of C + quantification; depends on a near-identical assembly, so not gradable until C reproduces well.
  • Functional annotation: 130,195 unigenes annotated; 52,574 with GO. Needs the exact 2020-vintage NR/KEGG/Swissprot/Pfam DBs (versioned in paper) — DB drift + downstream of C. Out of first-pass scope.
  • Sesquiterpenoid biosynthesis candidate genes (Tables 2–3, e.g. TRINITY_DN833… DS log2FC 6.93/5.66). Downstream of C + annotation; not attempted.

OUT OF SCOPE (wet-lab / manual / not a pipeline)

  • Physiological measurements (relative water content, MDA, proline, antioxidant enzyme activity, chlorophyll) — wet-lab assays.
  • qRT-PCR validation of 15 DEGs (Fig 9) — wet-lab.
  • Sample collection, RNA extraction, library prep, sequencing.

Reproducibility caveats (recorded up front, not as excuses)

  • Trinity version + parameters are unspecified ("default"); de novo assembly is inherently non-deterministic (k-mer graph, thread order). An exact 215,665-unigene match is not expected; agreement is judged on order-of-magnitude + summary-stat proximity (N50, mean length, GC).
  • "Default parameters" for SeqPrep/Sickle leaves the adapter and quality cutoffs implicit; we use each tool's documented defaults and record them in environment.lock.
Figures / tables: Table
A_raw_reads
Reported
CK 43.3M / D3 44.0M / D9 44.2M raw reads per group (Table 1)
Reproduced
CK 43.33M / D3 44.05M / D9 44.17M (ENA read_count spots x2 mates, avg per sample)
exact
B_clean_data
Reported
Clean bases 6.28/6.30/6.33 Gb per group; Q30 94.59% (Table 1)
Reproduced
CK 6.261 / D3 6.279 / D9 6.309 Gb clean bases; Q30 94.83% (SeqPrep+Sickle defaults, «job» COMPLETED). All groups within 0.3% of reported; Q30 +0.24pp.
within tolerance
C_assembly
Reported
Trinity: 215,665 unigenes, avg 759.09 bp, N50 1140 bp, GC 45.24% (Table 1)
Reproduced
NOT COMPLETED (computational blocker). Trinity 2.15.1 defaults ran on the 9 reproduced clean libs; Phase 1 (Inchworm/Chrysalis) COMPLETED -> 682,761 cluster-assembly commands; Phase 2 (Butterfly, full 192-CPU node) ran at ~0.47 cmd/s = 0.79% in 3h09m, projecting ~396h, infeasible within the cluster's 12h/job MaxWall. No Trinity.fasta produced. Non-deterministic regardless.
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Claim A (Table 1 per-group raw reads) reproduces exactly from ENA spot counts (43.33/44.05/44.17M vs 43.3/44.0/44.2M), confirming the public SRA data is the paper's raw input — no fabrication signal. Claims B (clean bases + Q30 94.59%) and C (Trinity 215,665 unigenes etc.) were not completed on our side: B was mid-run on «our HPC» at finalize, C was staged but never launched, and all downstream DEG/annotation/sesquiterpenoid claims cascade off the un-run assembly. The only deviation seen is rounding; the gap is our incompleteness, not an authors' defect, so derivability and core-claim support are graded yellow (partly demonstrated) rather than red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

347.3 k
tokens (I/O) · 31.9 M incl. cache
215 min
runtime · 660.76 CPU-h
123.9 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine