Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Regulatory and evolutionary adaptation of yeast to acute lethal ethanol stress.

PLoS One · 2020
L1 68/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
68/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 32% of all assessed papers rank 765 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED-WELL-ENOUGH: YES (for the downstream pipeline). 1:1 reproduction. The paper ships PROCESSED data in GEO (GSE151784 TAG counts; GSE151785 kallisto counts+TPM) and the methods (DESeq2 BH<0.05; Pearson correlation) are fully specified, so the reported numbers are re-derivable directly from the shipped data without re-running alignment. PRIMARY CLEAR DATA POINTS both reproduced within tolerance: (1) deletion-library DESeq2 screen 186 enriched / 714 depleted vs reported 192 / 735 (total 900 vs 927, within ~3%); (2) headline RNA-seq Pearson correlation r=0.84 (log2 TPM) vs reported 0.88. Named-gene Table-2 scores reproduce with the same sign and ~80-100% magnitude (VPS69 -13.0 vs -12.4 near-exact). The repo (github yang-jamie/SCRIPTS, commit 7b4104a) is a loose pile of one-off scripts with no README/env/license/orchestration; we reproduced by running the described tool (DESeq2 1.50.2, R 4.5.3) on the paper's own data (valid per P16). KEY AUDITOR NOTE (not fabrication): the paper's enrichment score = DESeq2 condition_pre_vs_post LFC = log2(pre/post) used un-negated exactly as in the authors' YKO_DEseq.R; the counts/signs reproduce ONLY under this convention (naive log2(post/pre) inverts every sign). NOT ATTEMPTED (optional ~20%): raw-FASTQ realignment (bowtie2 for TAGs, kallisto for RNA-seq) from GSE15178*_RAW.tar; iPAGE/GO enrichment figures (no single pinnable number); all wet-lab/phenotype results (out of scope). Magnitude gaps are DESeq2-version-level, not method-level. No fabrication detected.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 68
    assessed: 2026-06-15 ⛓ c33d9c4c78b6
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The genetic and regulatory factors underlying yeast (S. cerevisiae) adaptation and survival to stresses that cross the lethality threshold—specifically acute lethal ethanol exposure—have not been systematically studied; this paper tests how yeast cells survive and evolutionarily adapt to acute lethal ethanol stress.

Core claims
  • Yeast cells activate a rapid transcriptional reprogramming process following acute lethal ethanol stress that is likely adaptive for post-stress survival. finding
  • The early transcriptional response to acute lethal ethanol stress is dominated by global downregulation of gene expression (~5x more downregulated than upregulated genes). finding
  • Fitness profiling of the pooled yeast deletion library identifies non-essential genes whose deletion increases (e.g., ribosomal, mitochondrial, TOR pathway) or decreases (e.g., vacuolar functions) survival under lethal ethanol stress. finding
  • Repeated cycles of lethal ethanol exposure can experimentally evolve yeast strains with an order-of-magnitude higher ethanol tolerance/survival without compromising bulk growth rate. finding
  • Hyper-ethanol-tolerant evolved strains reprogram their pre-stress gene expression states to match the likely adaptive post-stress response of the wild-type strain. mechanism
  • A gene's transcriptional response is concordant with its fitness contribution: negative fitness scores correspond to positive expression changes and positive fitness scores to negative expression changes. finding
  • Meiosis/sporulation-associated genes (condensed chromosome, spore wall assembly) are modulated in haploid yeast via UME6/IME1 regulation, suggesting novel adaptive value under acute lethal ethanol stress. mechanism
  • A short 2-minute lethal ethanol exposure paradigm minimizes transcriptional response during stress, isolating pre-stress state and post-stress recovery as dominant survival contributors. method
Experimental setups
Assay System Perturbation Readout Platform
Colony-forming-unit survival assay (stress-survival curve) S. cerevisiae haploid strain BY4741 2-minute acute ethanol exposure (19%-26%) fraction survival of CFUs
Bulk RNA-seq time-course transcriptional profiling S. cerevisiae haploid strain BY4741 2-minute 20% (threshold lethal) ethanol exposure genome-wide gene expression / differentially expressed genes across post-stress time points
Pooled deletion-library fitness/survival profiling (barcode abundance) S. cerevisiae pooled haploid deletion library (~4,500 non-essential gene deletions, 20-nt barcodes) 2-minute 24.5% ethanol exposure (1% survival concentration) fitness/survival scores from post- vs pre-stress barcode abundance
Experimental evolution S. cerevisiae populations repeated cycles of lethal ethanol exposure evolved ethanol tolerance/survival and bulk growth rate
Key results
  • Ethanol exposure above 20% causes lethality, with survival dropping exponentially within the lethal range. 10^-5 at 26% ethanol
  • Peak transcriptional response occurs early (15-minute time point) with the most genes showing a two-fold change.
  • At early time points (15 and 30 min) ~1500 genes show two-fold decrease vs ~300 with two-fold increase (~5:1). ~5-fold more downregulated (~1500 vs ~300)
  • 735 gene deletions significantly diminished survival and 192 significantly improved survival. 735 down, 192 up
  • Negative fitness scores have significant positive expression change and positive fitness scores have significant negative expression change. p<2.2e-16 and p=1.9e-13
  • Evolved strains achieved an order of magnitude improvement in survival without compromising growth rate. ~10-fold (order of magnitude)
  • Top enriched deletions include AAH1 (8.26), LTV1 (8.05), FTR1 (7.72), RPL13A (7.72); top depleted include VPS69 (-12.4), SIW14 (-6.92), NGG1 (-6.81). scores 8.26 to -12.4
  • Spore wall assembly and condensed chromosome genes initially decrease then increase by 60 min, tracking UME6 upregulation and IME1 downregulation early.
Key statistics
  • count 10^-5 survival at 26% ethanol (fraction survival of CFUs at highest ethanol concentration)
  • count ~1500 downregulated vs ~300 upregulated genes (two-fold) (early post-stress time points (15 and 30 min))
  • pvalue p < 2.2 x 10^-16 (negative fitness scores have significant positive expression change)
  • pvalue p = 1.9 x 10^-13 (positive fitness scores have significant negative expression change)
  • count 735 gene deletions diminished survival; 192 improved survival (significant fitness effects from deletion library)
  • count 4,500 deletions; 20-nucleotide barcodes (pooled haploid yeast deletion library coverage)
  • count 24.5% ethanol = 1% wild-type survival (concentration used for deletion-library fitness profiling)
  • other 900 ESR genes (300 induced, 600 repressed) (environmental stress response reference from prior literature)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combines RNA-seq time-course expression profiling, pooled deletion-library fitness profiling (performed in triplicate), and experimental evolution to characterize yeast responses to acute lethal ethanol stress. Differential expression was assessed by a two-fold change threshold relative to a pre-stress reference, patterns were explored with k-means and hierarchical clustering plus enrichment tools (iPAGE GO analysis at p<0.001, FIRE motif discovery), and fitness scores were used to call gene deletions that significantly increased or decreased survival. Group differences in expression change between positive- and negative-fitness gene sets were reported with exact p-values on boxplots, and survival curves were shown with standard-error bars.

Replicationmixed Sample sizefitness profiling stated as performed in triplicate; no formal sample-size or power calculation described Groupspost-stress vs pre-stress time points; post-stress vs pre-stress deletion abundance; positive vs negative fitness-score gene sets; evolved vs parental strain Pairingunclear Randomization/blindingnot stated DispersionSEM Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
two-fold expression change threshold (differential expression call) Fig 1C, number of differentially expressed genes per post-stress time point vs pre-stress not stated
GO/functional enrichment via iPAGE (significance threshold p<0.001) Fig 1D/1E clusters and post-stress vs pre-stress comparisons; S1 Table; S3 Fig not stated
cis-regulatory motif enrichment via FIRE Fig 1D cluster 2 (URS1 motif discovery); S2 Fig not stated
significance test of expression-change difference between positive- vs negative-fitness gene groups (test type not stated) Fig 2D boxplot (p<2.2x10^-16 and p=1.9x10^-13) not stated
fitness-score significance call for gene-deletion survival effects (method referenced to Materials and Methods) Fig 2B/2C; 192 gene deletions improving and 735 diminishing survival triplicate fitness profiling not stated
Approaches that could also have been used
  • Differentially expressed genes were identified using a two-fold expression-change cutoff relative to the pre-stress time point.
    Could also: A model-based differential-expression framework such as DESeq2 or edgeR (Wald or likelihood-ratio test with Benjamini-Hochberg FDR) could also have been used. — Such models combine effect size with a statistical significance/uncertainty estimate and provide formal multiple-testing control across thousands of genes, which complements a fixed fold-change threshold.
  • Survival curves were summarized with standard-error bars.
    Could also: Standard deviation or a 95% confidence interval could also be displayed. — SD conveys the spread of the underlying replicates and CIs convey the precision of the estimate; both are often preferred, especially with small n, to make the meaning of the error bars explicit.
  • The number of replicates is described as triplicate for fitness profiling, without a stated power or sample-size rationale.
    Could also: A brief a priori power consideration or a statement of the basis for the replicate number could also be included. — Documenting the basis for n helps readers gauge the sensitivity of the design and supports reproducibility.
  • Group differences in Fig 2D were reported with p-values but the specific test is not named in the available text.
    Could also: Explicitly naming the test (e.g., Mann-Whitney U / Wilcoxon rank-sum, or a t-test) and reporting an accompanying effect size could also be done. — Naming the test and giving an effect size (e.g., rank-biserial correlation or difference in medians) clarifies the assumptions and the practical magnitude alongside the very small p-values.
  • Clustering used k-means with 10 clusters and complementary hierarchical clustering.
    Could also: A cluster-number selection criterion (e.g., gap statistic, silhouette analysis) or model-based clustering could also be reported. — Reporting how the number of clusters was chosen adds transparency about the robustness of the partition and how sensitive the patterns are to that choice.
  • Multiple genome-wide comparisons (per-time-point expression and per-gene fitness) were performed.
    Could also: A single explicitly stated family-wise or FDR correction scheme spanning these comparisons could also be described. — Stating the multiplicity scope and method makes the control of false positives across the many tests transparent to the reader.
Software: iPAGE (pathway/GO enrichment) · FIRE (cis-regulatory motif discovery) · k-means clustering algorithm · RNA-seq analysis pipeline (unspecified)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
46
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33170850

Paper: Yang J, Tavazoie S. Regulatory and evolutionary adaptation of yeast to acute lethal ethanol stress. PLoS One 2020. PMID 33170850 · PMC7654773 · DOI 10.1371/journal.pone.0239528.

Code: https://github.com/yang-jamie/SCRIPTS (commit 7b4104a, 2020-08-22). A loose collection of one-off scripts (Perl/R/C/shell) — no README, no environment file, no orchestrated pipeline, no license. The scripts document the methods (parameters, tools) but are not a runnable end-to-end workflow. Per BRIEF rule 2 this does not disqualify the paper: we reproduce the pipeline-derived downstream results by running the described standard tools on the paper's own processed data shipped in GEO.

Data: GEO GSE151786 (SuperSeries) = two SubSeries:

  • GSE151784 — pooled yeast deletion-library (YKO) barcode/fitness screen. 6 samples: pre_rep1-3 (GSM4590920-922), post_rep1-3 (GSM4590923-925). Each ships *_counts.txt = abundance of each 20-nt TAG (processed).
  • GSE151785 — RNA-seq, WT vs evolved strain (JY304) timecourse. 10 samples (GSM4590926-935). Each ships *_abundance.txt = kallisto est-counts + TPM per transcript (processed).

In scope (pipeline-derived, attempted)

# Result Pipeline (as described) Input data Reported value
C1 # significantly enriched gene deletions (improved survival) DESeq2 on TAG counts, paired ~pair+condition, BH padj<0.05 GSE151784 6 count files 192
C2 # significantly depleted gene deletions (diminished survival) same DESeq2 run GSE151784 735
C3 Top enrichment scores (= shrunken log2FC) for named deletions DESeq2 lfcShrink(apeglm) GSE151784 AAH1=8.26, LTV1=8.05, FTR1=7.72; VPS69=−12.4
C4 Pearson correlation, evolved (JY304) pre-stress vs WT 15-min post-stress (headline, Fig 4C-E) cor.test() on per-gene TPM GSE151785 GSM4590932 vs GSM4590928 r = 0.88
C5 Evolved pre-stress vs WT pre-stress: #genes up/down >2-fold (Table 3) TPM fold-change threshold GSE151785 GSM4590932 vs GSM4590927 266 up / 52 down

Out of scope (not attempted) — and why

  • Raw alignment (bowtie2 for TAGs; kallisto for RNA-seq from FASTQ): the paper ships the processed outputs of these steps, so re-deriving them from the GSE15178*_RAW.tar FASTQs is the optional hard ~20% (BRIEF rule 3). We start from the provided processed counts/TPM — a faithful reproduction of the downstream statistical pipeline that produced the reported numbers.
  • Wet-lab / phenotypic results (survival assays, the >10-fold survival of the evolved strain, experimental evolution): not computational — out of scope.
  • iPAGE / GO-category enrichment figures (Suppl. Fig 3): qualitative pathway statements without a single pinnable number; not attempted.
  • C5 is a secondary check; C1/C2 (YKO DESeq2 counts) and C4 (r=0.88 correlation) are the primary CLEAR data points.

Primary CLEAR data points

C1+C2 (192/735 DESeq2 hits) and C4 (Pearson r=0.88). Both are computed directly from the provided processed data with the exact described method.

Figures / tables: TableFig 4C
C1
Reported
192 enriched gene deletions (DESeq2 BH<0.05)
Reproduced
186
within tolerance
C2
Reported
735 depleted gene deletions (DESeq2 BH<0.05)
Reproduced
714
within tolerance
Ctot
Reported
927 total significant deletions
Reproduced
900
within tolerance
C3d
Reported
VPS69 enrichment score -12.4 (Table 2)
Reproduced
-13.0
within tolerance
C3a
Reported
AAH1 enrichment score 8.26 (Table 2)
Reproduced
6.61
partial
C3b
Reported
LTV1 enrichment score 8.05 (Table 2)
Reproduced
6.51
partial
C3c
Reported
FTR1 enrichment score 7.72 (Table 2)
Reproduced
6.55
partial
C4
Reported
Pearson r 0.88 evolved-pre vs WT-15min (Fig 4C-E)
Reproduced
0.84 (log2 TPM) / 0.96 (raw TPM)
within tolerance
C5a
Reported
266 genes up >2-fold (Table 3)
Reproduced
192
partial
C5b
Reported
52 genes down >2-fold (Table 3)
Reproduced
80
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 68/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

Reproduced directly from the authors' shipped GEO processed data: the headline deletion-screen counts (186/714 vs 192/735, within ~3%) and the central evolved-pre vs WT-15min Pearson correlation (r=0.84–0.96 vs 0.88) both confirm the paper's core claims, and named Table-2 scores reproduce with correct sign (VPS69 −13.0 vs −12.4 near-exact). Remaining gaps are on our/technical side: ~80% Table-2 magnitudes from a newer DESeq2 version, and the C5 transcript-vs-gene aggregation choice (266→192). The single non-obvious step — the log2(pre/post) sign convention — is fully derivable from the authors' committed code, so no fabrication. Overall a solid, explainable reproduction, hence yellow rather than a clean 1:1.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

183 k
tokens (I/O) · 14.9 M incl. cache
20 min
runtime · 0.01 CPU-h
0.8 GB
peak RAM
3
HPC jobs
hummel
machine