What's Genetic Variation Got to Do with It? Starvation-Induced Self-Fertilization Enhances Survival in Paramecium.
The main results reproduced: recomputed values matched the published ones within tolerance.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Clean 1:1 reproduction of all four in-scope pipeline-derived claims, independently re-run end-to-end on «our HPC» compute nodes (the prior «infra» workdir had been reclaimed; this run rebuilt env+data+index+results from scratch). C1 read depth reproduced within-tol (counted read pairs == ENA read_count exactly on the 3-sample subset; mean 29.70M pairs/sample; 27/27 runs >25M reads). C2 trimming exact (trim-galore 0.4.5 + cutadapt 1.18 default, 99.96% pairs retained, Illumina adapter auto-detected). C3 alignment within-tol (STAR 2.7.11a uniquely-mapped mean 91.30% vs reported ~90%; minor STAR 2.7-vs-2.5 drift, noted). C4 within-tol: full 27-sample featureCounts->CPM>1 gave 31,848 genes unstranded vs reported 31,830 (+0.06%); 31,681 reverse-stranded (-0.47%). The paper describes its computational core well enough to reproduce, and the public data (PRJEB33070, 27/27 runs) delivers exactly what is promised. NOT attempted: C5 DE trend (under-specified trend test + ambiguous 8-of-9-day selection) and C6 HSP70 GO enrichment (online DAVID/PANTHER) - both honestly out of scope, not dropped. Robustness note: long single-node jobs repeatedly died at ~1-2h with clean logs (not a wall-time limit); splitting alignment into a 27-task SLURM array of short jobs resolved it. All grades provisional pending human audit.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 89assessed: 2026-06-20 ⛓ 6b810f84d3b2
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether self-fertilization (autogamy) in Paramecium tetraurelia can enhance survival (stress resistance) independent of new genetic variation, based on the hypothesis that the molecular machinery underlying sex overlaps with the cellular stress response.
- ★ Self-fertilization enhances heat shock resistance in P. tetraurelia in the absence of new genetic variation finding
- ★ Cells that have recently achieved competence for self-fertilization also display increased heat shock resistance finding
- ★ Enhanced stress resistance is negatively coupled with cell proliferation, possibly shaped by withdrawal of growth factor(s) finding
- ★ Self-fertilization, onset of sexual maturation, stress response, and cell proliferation are interlinked in P. tetraurelia finding
- ★ Heat shock proteins with meiotic/fertilization functions (HSP_MF) may mechanistically link sexual reproduction to the stress response mechanism
- A time-course transcriptomics study across the clonal life cycle post-autogamy provides insight into molecular mechanisms underlying the survival advantage method
- ★ Age at autogamy maturity relates to cell proliferation rate, assessed via meta-analysis across strains and diets finding
- ★ Dietary restriction (starvation) enhances survival under prolonged lethal heat stress compared with nutrient-rich fed cells finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Heat shock tolerance test | Paramecium tetraurelia, strain d12 | heat shock (43°C, 90s) | binary cell survival (alive/dead) after 24h recovery | Eppendorf Mastercycler pro S thermocycler |
| Autogamy reactivity test (DAPI staining) | Paramecium tetraurelia, strain d12 | starvation/nutritional deprivation | fraction of autogamous cells (macronuclear fragmentation) | ZEISS Axioskop 2 epifluorescence microscope |
| Bulk RNA-seq (time-course transcriptomics) | Paramecium tetraurelia, strain d12 | none (clonal aging post-autogamy, 9 time points) | transcript abundance/differential gene expression | Illumina HS4000 |
| Cell proliferation and sexual maturation meta-analysis | Paramecium tetraurelia, strains d12, d4-2, 51 | diet (Alcaligenes faecalis vs Enterobacter aerogenes) | cell proliferation rate and age at autogamy maturity | — |
| Dietary restriction and survival-to-stress assay | Paramecium tetraurelia, postautogamous cells | starvation vs nutrient-rich diet, then 40°C heat exposure | cell survival/density until total cell death | — |
- ▲ Cell survival after heat shock increases in association with rising autogamy competence across the clonal life cycle
- – Rate of cell proliferation between days did not differ statistically significantly (Kruskal-Wallis test) P=0.10
- – Average cell proliferation rate across clonally aging mass cultures ~4 fissions/day
- – On average, RNA-seq reads aligned to the P. tetraurelia reference genome 90%
- – Genes retained for differential expression analysis (CPM>1 in ≥1 of 27 samples) 31,830 genes
- – Average RNA Integrity Number (RIN) of extracted total RNA 8.3
- pvalue P=0.10 (Kruskal-Wallis test for differences in daily cell proliferation rate across the time course)
- count 31,830 (genes retained for differential expression analysis (CPM>1 in ≥1 of 27 samples))
- mean 8.3 (average RNA Integrity Number (RIN) of extracted total RNA)
- other 90% (average percentage of reads aligned to the reference genome per sample)
- count 27 (RNA samples sequenced (9 time points × 3 replicates))
- pvalue <0.05 (adjusted P value threshold (Benjamini-Hochberg) for calling differentially expressed genes)
- count ~10^5 (log-phase cells concentrated per RNA extraction sample)
- other ~4 fissions/day (average cell proliferation rate across clonally aging mass cultures)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combined a nine-day time-course transcriptomics experiment in Paramecium tetraurelia (3 biological replicates per time point) with phenotypic assays (heat-shock tolerance, autogamy competence, dietary-restriction survival) to link self-fertilization to stress resistance. Differential gene expression was evaluated with two independent pipelines (DESeq2 and EdgeR), retaining only genes flagged as significant by both after Benjamini–Hochberg FDR correction. Phenotypic comparisons used the Kruskal–Wallis test and linear regression; dispersion was reported as standard error of the mean.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Kruskal–Wallis test | Comparison of daily cell-proliferation rates across days 1–8 post-autogamy (fig. 2A) | 3 biological replicates per day, 9 days | not stated |
| DESeq2 Wald test (negative binomial GLM) | Pairwise differential gene expression analysis across time points | 27 RNA-seq samples (9 time points × 3 biological replicates) | not stated |
| EdgeR likelihood ratio / quasi-likelihood test (negative binomial GLM) | Pairwise differential gene expression analysis across time points (parallel to DESeq2) | 27 RNA-seq samples (9 time points × 3 biological replicates) | not stated |
| Linear regression | Autogamy competence proportion regressed on clonal age (DPA) to extrapolate age at 90% autogamy maturity | 3 strains × ≥3 biological replicates per bacterium condition | not stated |
| Principal component analysis (PCA) | Visualization of between-sample transcriptomic distances (fig. 2B) | 27 samples | na |
-
Dispersion in phenotypic data (proliferation rate, survival) was summarized with standard error of the mean (SEM)↳ Could also: Standard deviation (SD) or 95% confidence intervals could also describe the spread of the data — With small biological replicate numbers (n=3 per time point), SD directly conveys the variability among replicates, while 95% CIs make estimation uncertainty explicit and are increasingly preferred in biological reporting guidelines
-
Proliferation-rate differences across nine daily time points were evaluated with a Kruskal–Wallis test↳ Could also: A linear mixed-effects model or repeated-measures ANOVA could also have been applied, treating the time-course structure explicitly — The three mass cultures were tracked longitudinally across nine days; a mixed-effects or repeated-measures approach would account for within-culture correlation across days and allow estimation of a time trend rather than only an omnibus difference test
-
The relationship between clonal age and autogamy competence (proportion of cells capable of autogamy) was modeled with linear regression↳ Could also: Logistic regression or a beta regression could also model a bounded proportion outcome as a function of age — Autogamy competence is a proportion (bounded 0–1); logistic or beta regression naturally respects these bounds and can model the sigmoidal accumulation of competence with age without extrapolating beyond the feasible range
-
Differential gene expression was identified by taking the intersection of DESeq2 and EdgeR results (consensus filtering)↳ Could also: A single pipeline with a more stringent FDR threshold, or a unified framework such as limma-voom, could also have been used — Consensus filtering increases specificity at the cost of sensitivity; limma-voom or a single well-validated pipeline with a tighter FDR threshold (e.g., 0.01) is a widely used alternative that provides a principled statistical justification rather than relying on the overlap of two correlated procedures
-
Survival of fed vs. starved cells under lethal heat was assessed by measuring cell density over time, with the statistical comparison not explicitly named in the available text↳ Could also: A formal survival or time-to-event analysis (e.g., Kaplan–Meier curves with a log-rank test, or a parametric survival model) could also characterize the time-to-death profiles — When cell density declines over a heat-exposure period, survival analysis handles censoring and captures the full temporal shape of the mortality curve; it also provides a hazard ratio as a natural effect-size metric for the dietary-group comparison
-
Heat-shock survival (binary alive/dead outcome per well) was compared across time points from day 0 to day 8↳ Could also: A generalized linear mixed model (binomial GLMM) could also have been applied to the binary per-cell outcome, with day as a fixed effect and biological replicate as a random effect — A GLMM on the raw binary data uses all individual-cell information rather than aggregating to proportions first, appropriately accounts for the nested structure (cells within replicates within days), and allows simultaneous estimation of the time-course effect with confidence intervals
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32163147
Paper: Thind, Vitali, Guarracino, Catania (2020) Genome Biol Evol 12(4). "What's Genetic Variation Got to Do with It? Starvation-Induced Self-Fertilization Enhances Survival in Paramecium." DOI 10.1093/gbe/evaa052.
Linked code artifact (registry): https://github.com/FelixKrueger/TrimGalore — the third-party read-QC/adapter-trimming tool (Trim Galore, a wrapper around Cutadapt + FastQC) the paper used as its read-cleaning step. Per brief rule P16, applying this third-party tool to the paper's own data is a valid reproduction.
Data: ENA PRJEB33070 — 27 paired-end strand-specific RNA-seq runs of
Paramecium tetraurelia strain d12 (sample design "9 days × 3 replicates").
All 27 runs public, FASTQ on ENA FTP. Confirmed via ENA portal: 27 runs, all
RNA-Seq / PAIRED, total 801,990,715 read pairs, mean 29.70 M pairs/sample,
range 23.36 M – 39.97 M, total ~127 GB compressed FASTQ.
Exact reported claims (verified against PMC7239694 full text)
- C1 "On an average, more than 25 million reads per samples were produced from the sequencing." (Methods/Results)
- C2 "Low-quality reads were removed using default options of trim galore version 0.4.5"; "Quality control … using the FastQC (version 0.11.57)".
- C3 "On average, for each sample 90% of the reads aligned to the reference genome." Aligner: "STAR version 2.5.0a (default parameters)" onto the P. tetraurelia genome with "the annotation of the P51 version 2".
- C4 "A total of 31,830 genes with count per million >1 in ≥1 of the 27 samples were retained." (Results) ← fully specified filter; in scope.
- C5 "Of 32,252 genes with nonzero expression values across each of the experimental days, 1,037 and 1,145 have increasing and decreasing levels of expression, respectively, over the eight vegetative time points." (Results)
- C6 HSP70 enrichment (13 observed vs ~4 expected) — DAVID/PANTHER online.
The reported computational pipeline (Methods)
FastQC 0.11.57 (QC) → Trim Galore 0.4.5 (default) → STAR 2.5.0a (default, map to P. tetraurelia mac_51 + annotation P51 v2.0) → quantify (featureCounts) → CPM filter (C4) → DE trend over 8 days (C5; edgeR/DESeq2) → DAVID/PANTHER enrichment (C6).
In scope — what we attempt
| id | reported | pipeline step | plan |
|---|---|---|---|
| C1 | ">25M reads/sample (avg)" | sequencing/QC input | count read pairs from FASTQ; cross-check ENA read_count + dataset mean |
| C2 | Trim Galore 0.4.5 default | the linked tool | run trim_galore 0.4.5 (cutadapt 1.18, paper-era) on the data; capture report |
| C3 | "~90% aligned (avg)" | STAR | STAR index (mac_51 + annot v2.0) + align; total mapped % = uniq+multi |
| C4 | 31,830 genes CPM>1 in ≥1/27 | featureCounts + CPM filter | STRETCH: full 27-sample STAR→featureCounts→count genes CPM>1 in ≥1 sample |
C1–C3 run on a representative 3-sample subset (smallest/mid/largest by ENA read_count: ERR3627202 / ERR3627207 / ERR3627102) — enough for clear 1:1 data points. C4 requires the full 27-sample STAR+featureCounts run (well-specified filter, so worth the heavier compute — 80/20 is a floor, not a ceiling).
Out of scope (the harder ~20% / non-computational)
- C5 DE trend (1,037 up / 1,145 down over 8 vegetative time points): the trend-test method (tool, model, which day is excluded to get "8 vegetative time points" from 9 days) is under-specified. Attempt only as a far stretch; otherwise not claimed.
- C6 HSP70 enrichment: DAVID/PANTHER online tools on the full DE set — not a locally-reproducible pipeline step. Out of scope.
- All wet-lab phenotypes (autogamy %, heat-shock survival, fission rates, dietary restriction) — not pipeline-derived. Out of scope.
Version-drift note
Pinned: Trim Galore 0.4.5 (bioconda pkg trim-galore-0.4.5; its script
self-reports "0.4.4_dev" — a known bioconda quirk for the 0.4.5 build) +
cutadapt 1.18 (paper-era 1.x, on py
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.