Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

What's Genetic Variation Got to Do with It? Starvation-Induced Self-Fertilization Enhances Survival in Paramecium.

Genome Biol Evol · 2020
89/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
89/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 77% of all assessed papers rank 246 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Clean 1:1 reproduction of all four in-scope pipeline-derived claims, independently re-run end-to-end on «our HPC» compute nodes (the prior «infra» workdir had been reclaimed; this run rebuilt env+data+index+results from scratch). C1 read depth reproduced within-tol (counted read pairs == ENA read_count exactly on the 3-sample subset; mean 29.70M pairs/sample; 27/27 runs >25M reads). C2 trimming exact (trim-galore 0.4.5 + cutadapt 1.18 default, 99.96% pairs retained, Illumina adapter auto-detected). C3 alignment within-tol (STAR 2.7.11a uniquely-mapped mean 91.30% vs reported ~90%; minor STAR 2.7-vs-2.5 drift, noted). C4 within-tol: full 27-sample featureCounts->CPM>1 gave 31,848 genes unstranded vs reported 31,830 (+0.06%); 31,681 reverse-stranded (-0.47%). The paper describes its computational core well enough to reproduce, and the public data (PRJEB33070, 27/27 runs) delivers exactly what is promised. NOT attempted: C5 DE trend (under-specified trend test + ambiguous 8-of-9-day selection) and C6 HSP70 GO enrichment (online DAVID/PANTHER) - both honestly out of scope, not dropped. Robustness note: long single-node jobs repeatedly died at ~1-2h with clean logs (not a wall-time limit); splitting alignment into a 27-task SLURM array of short jobs resolved it. All grades provisional pending human audit.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 89
    assessed: 2026-06-20 ⛓ 6b810f84d3b2
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether self-fertilization (autogamy) in Paramecium tetraurelia can enhance survival (stress resistance) independent of new genetic variation, based on the hypothesis that the molecular machinery underlying sex overlaps with the cellular stress response.

Core claims
  • Self-fertilization enhances heat shock resistance in P. tetraurelia in the absence of new genetic variation finding
  • Cells that have recently achieved competence for self-fertilization also display increased heat shock resistance finding
  • Enhanced stress resistance is negatively coupled with cell proliferation, possibly shaped by withdrawal of growth factor(s) finding
  • Self-fertilization, onset of sexual maturation, stress response, and cell proliferation are interlinked in P. tetraurelia finding
  • Heat shock proteins with meiotic/fertilization functions (HSP_MF) may mechanistically link sexual reproduction to the stress response mechanism
  • A time-course transcriptomics study across the clonal life cycle post-autogamy provides insight into molecular mechanisms underlying the survival advantage method
  • Age at autogamy maturity relates to cell proliferation rate, assessed via meta-analysis across strains and diets finding
  • Dietary restriction (starvation) enhances survival under prolonged lethal heat stress compared with nutrient-rich fed cells finding
Experimental setups
Assay System Perturbation Readout Platform
Heat shock tolerance test Paramecium tetraurelia, strain d12 heat shock (43°C, 90s) binary cell survival (alive/dead) after 24h recovery Eppendorf Mastercycler pro S thermocycler
Autogamy reactivity test (DAPI staining) Paramecium tetraurelia, strain d12 starvation/nutritional deprivation fraction of autogamous cells (macronuclear fragmentation) ZEISS Axioskop 2 epifluorescence microscope
Bulk RNA-seq (time-course transcriptomics) Paramecium tetraurelia, strain d12 none (clonal aging post-autogamy, 9 time points) transcript abundance/differential gene expression Illumina HS4000
Cell proliferation and sexual maturation meta-analysis Paramecium tetraurelia, strains d12, d4-2, 51 diet (Alcaligenes faecalis vs Enterobacter aerogenes) cell proliferation rate and age at autogamy maturity
Dietary restriction and survival-to-stress assay Paramecium tetraurelia, postautogamous cells starvation vs nutrient-rich diet, then 40°C heat exposure cell survival/density until total cell death
Key results
  • Cell survival after heat shock increases in association with rising autogamy competence across the clonal life cycle
  • Rate of cell proliferation between days did not differ statistically significantly (Kruskal-Wallis test) P=0.10
  • Average cell proliferation rate across clonally aging mass cultures ~4 fissions/day
  • On average, RNA-seq reads aligned to the P. tetraurelia reference genome 90%
  • Genes retained for differential expression analysis (CPM>1 in ≥1 of 27 samples) 31,830 genes
  • Average RNA Integrity Number (RIN) of extracted total RNA 8.3
Key statistics
  • pvalue P=0.10 (Kruskal-Wallis test for differences in daily cell proliferation rate across the time course)
  • count 31,830 (genes retained for differential expression analysis (CPM>1 in ≥1 of 27 samples))
  • mean 8.3 (average RNA Integrity Number (RIN) of extracted total RNA)
  • other 90% (average percentage of reads aligned to the reference genome per sample)
  • count 27 (RNA samples sequenced (9 time points × 3 replicates))
  • pvalue <0.05 (adjusted P value threshold (Benjamini-Hochberg) for calling differentially expressed genes)
  • count ~10^5 (log-phase cells concentrated per RNA extraction sample)
  • other ~4 fissions/day (average cell proliferation rate across clonally aging mass cultures)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combined a nine-day time-course transcriptomics experiment in Paramecium tetraurelia (3 biological replicates per time point) with phenotypic assays (heat-shock tolerance, autogamy competence, dietary-restriction survival) to link self-fertilization to stress resistance. Differential gene expression was evaluated with two independent pipelines (DESeq2 and EdgeR), retaining only genes flagged as significant by both after Benjamini–Hochberg FDR correction. Phenotypic comparisons used the Kruskal–Wallis test and linear regression; dispersion was reported as standard error of the mean.

Replicationmixed Sample sizeSample sizes given per assay (90 cells/replicate for heat shock; ~56 cells/replicate for autogamy; 12 independent biological replicates per dietary group; 27 RNA-seq samples); no formal power analysis or a priori power calculation mentioned GroupsSelf-fertilizing / recently post-autogamous cells vs. clonally aging vegetative cells; nutritionally deprived vs. fed cells; multiple successive time points (days 0–8 post-autogamy) Pairingunclear Randomization/blindingstated DispersionSEM Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini–Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
Kruskal–Wallis test Comparison of daily cell-proliferation rates across days 1–8 post-autogamy (fig. 2A) 3 biological replicates per day, 9 days not stated
DESeq2 Wald test (negative binomial GLM) Pairwise differential gene expression analysis across time points 27 RNA-seq samples (9 time points × 3 biological replicates) not stated
EdgeR likelihood ratio / quasi-likelihood test (negative binomial GLM) Pairwise differential gene expression analysis across time points (parallel to DESeq2) 27 RNA-seq samples (9 time points × 3 biological replicates) not stated
Linear regression Autogamy competence proportion regressed on clonal age (DPA) to extrapolate age at 90% autogamy maturity 3 strains × ≥3 biological replicates per bacterium condition not stated
Principal component analysis (PCA) Visualization of between-sample transcriptomic distances (fig. 2B) 27 samples na
Approaches that could also have been used
  • Dispersion in phenotypic data (proliferation rate, survival) was summarized with standard error of the mean (SEM)
    Could also: Standard deviation (SD) or 95% confidence intervals could also describe the spread of the data — With small biological replicate numbers (n=3 per time point), SD directly conveys the variability among replicates, while 95% CIs make estimation uncertainty explicit and are increasingly preferred in biological reporting guidelines
  • Proliferation-rate differences across nine daily time points were evaluated with a Kruskal–Wallis test
    Could also: A linear mixed-effects model or repeated-measures ANOVA could also have been applied, treating the time-course structure explicitly — The three mass cultures were tracked longitudinally across nine days; a mixed-effects or repeated-measures approach would account for within-culture correlation across days and allow estimation of a time trend rather than only an omnibus difference test
  • The relationship between clonal age and autogamy competence (proportion of cells capable of autogamy) was modeled with linear regression
    Could also: Logistic regression or a beta regression could also model a bounded proportion outcome as a function of age — Autogamy competence is a proportion (bounded 0–1); logistic or beta regression naturally respects these bounds and can model the sigmoidal accumulation of competence with age without extrapolating beyond the feasible range
  • Differential gene expression was identified by taking the intersection of DESeq2 and EdgeR results (consensus filtering)
    Could also: A single pipeline with a more stringent FDR threshold, or a unified framework such as limma-voom, could also have been used — Consensus filtering increases specificity at the cost of sensitivity; limma-voom or a single well-validated pipeline with a tighter FDR threshold (e.g., 0.01) is a widely used alternative that provides a principled statistical justification rather than relying on the overlap of two correlated procedures
  • Survival of fed vs. starved cells under lethal heat was assessed by measuring cell density over time, with the statistical comparison not explicitly named in the available text
    Could also: A formal survival or time-to-event analysis (e.g., Kaplan–Meier curves with a log-rank test, or a parametric survival model) could also characterize the time-to-death profiles — When cell density declines over a heat-exposure period, survival analysis handles censoring and captures the full temporal shape of the mortality curve; it also provides a hazard ratio as a natural effect-size metric for the dietary-group comparison
  • Heat-shock survival (binary alive/dead outcome per well) was compared across time points from day 0 to day 8
    Could also: A generalized linear mixed model (binomial GLMM) could also have been applied to the binary per-cell outcome, with day as a fixed effect and biological replicate as a random effect — A GLMM on the raw binary data uses all individual-cell information rather than aggregating to proportions first, appropriately accounts for the nested structure (cells within replicates within days), and allows simultaneous estimation of the time-course effect with confidence intervals
Software: DESeq2 (R/Bioconductor) · EdgeR (R/Bioconductor) · STAR 2.5.oa · FastQC 0.11.57 · Trim Galore 0.4.5 · featureCounts (R/Bioconductor) · DAVID · PANTHER

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-32163147

Paper: Thind, Vitali, Guarracino, Catania (2020) Genome Biol Evol 12(4). "What's Genetic Variation Got to Do with It? Starvation-Induced Self-Fertilization Enhances Survival in Paramecium." DOI 10.1093/gbe/evaa052.

Linked code artifact (registry): https://github.com/FelixKrueger/TrimGalore — the third-party read-QC/adapter-trimming tool (Trim Galore, a wrapper around Cutadapt + FastQC) the paper used as its read-cleaning step. Per brief rule P16, applying this third-party tool to the paper's own data is a valid reproduction.

Data: ENA PRJEB33070 — 27 paired-end strand-specific RNA-seq runs of Paramecium tetraurelia strain d12 (sample design "9 days × 3 replicates"). All 27 runs public, FASTQ on ENA FTP. Confirmed via ENA portal: 27 runs, all RNA-Seq / PAIRED, total 801,990,715 read pairs, mean 29.70 M pairs/sample, range 23.36 M – 39.97 M, total ~127 GB compressed FASTQ.

Exact reported claims (verified against PMC7239694 full text)

  • C1 "On an average, more than 25 million reads per samples were produced from the sequencing." (Methods/Results)
  • C2 "Low-quality reads were removed using default options of trim galore version 0.4.5"; "Quality control … using the FastQC (version 0.11.57)".
  • C3 "On average, for each sample 90% of the reads aligned to the reference genome." Aligner: "STAR version 2.5.0a (default parameters)" onto the P. tetraurelia genome with "the annotation of the P51 version 2".
  • C4 "A total of 31,830 genes with count per million >1 in ≥1 of the 27 samples were retained." (Results) ← fully specified filter; in scope.
  • C5 "Of 32,252 genes with nonzero expression values across each of the experimental days, 1,037 and 1,145 have increasing and decreasing levels of expression, respectively, over the eight vegetative time points." (Results)
  • C6 HSP70 enrichment (13 observed vs ~4 expected) — DAVID/PANTHER online.

The reported computational pipeline (Methods)

FastQC 0.11.57 (QC) → Trim Galore 0.4.5 (default) → STAR 2.5.0a (default, map to P. tetraurelia mac_51 + annotation P51 v2.0) → quantify (featureCounts) → CPM filter (C4) → DE trend over 8 days (C5; edgeR/DESeq2) → DAVID/PANTHER enrichment (C6).

In scope — what we attempt

id reported pipeline step plan
C1 ">25M reads/sample (avg)" sequencing/QC input count read pairs from FASTQ; cross-check ENA read_count + dataset mean
C2 Trim Galore 0.4.5 default the linked tool run trim_galore 0.4.5 (cutadapt 1.18, paper-era) on the data; capture report
C3 "~90% aligned (avg)" STAR STAR index (mac_51 + annot v2.0) + align; total mapped % = uniq+multi
C4 31,830 genes CPM>1 in ≥1/27 featureCounts + CPM filter STRETCH: full 27-sample STAR→featureCounts→count genes CPM>1 in ≥1 sample

C1–C3 run on a representative 3-sample subset (smallest/mid/largest by ENA read_count: ERR3627202 / ERR3627207 / ERR3627102) — enough for clear 1:1 data points. C4 requires the full 27-sample STAR+featureCounts run (well-specified filter, so worth the heavier compute — 80/20 is a floor, not a ceiling).

Out of scope (the harder ~20% / non-computational)

  • C5 DE trend (1,037 up / 1,145 down over 8 vegetative time points): the trend-test method (tool, model, which day is excluded to get "8 vegetative time points" from 9 days) is under-specified. Attempt only as a far stretch; otherwise not claimed.
  • C6 HSP70 enrichment: DAVID/PANTHER online tools on the full DE set — not a locally-reproducible pipeline step. Out of scope.
  • All wet-lab phenotypes (autogamy %, heat-shock survival, fission rates, dietary restriction) — not pipeline-derived. Out of scope.

Version-drift note

Pinned: Trim Galore 0.4.5 (bioconda pkg trim-galore-0.4.5; its script self-reports "0.4.4_dev" — a known bioconda quirk for the 0.4.5 build) + cutadapt 1.18 (paper-era 1.x, on py

C1
Reported
on average more than 25 million reads per sample (PRJEB33070, 27 runs)
Reproduced
all 27 ENA runs: mean 29,703,360 read pairs/sample (59.4M reads), min 23.36M pairs, max 39.97M pairs; 27/27 runs >25M reads; subset ERR3627202/207/102 counted_read_pairs == ENA read_count EXACTLY (gzip read-count from FASTQ)
within tolerance
C2
Reported
Trim Galore 0.4.5, default options
Reproduced
trim-galore 0.4.5 + cutadapt 1.18 (py3.6) default --paired; all 27 samples trimmed; subset 3/3 retained 99.96-99.97% of pairs; Illumina adapter auto-detected (~35% reads), Phred-20 quality trim
exact
C3
Reported
on average 90% of reads aligned (STAR 2.5.0a, mac_51 + annotation v2.0)
Reproduced
STAR 2.7.11a, ptetraurelia_mac_51 (72.1 Mbp) + annotation v2.0 (41,533 genes): full 27-sample uniquely-mapped mean 91.30% (range 89.39-93.69%), total-mapped (uniq+multi) mean 97.94%
within tolerance
C4
Reported
31,830 genes with count-per-million >1 in >=1 of 27 samples
Reproduced
full 27-sample Trim Galore->STAR->featureCounts(gene,-p --countReadPairs)->CPM>1 filter: 31,848 genes unstranded (-s0, +0.06%) / 31,681 reverse-stranded (-s2, -0.47%); library is reverse-stranded (forward -s1 assigns 4.67M vs 726.9M read pairs)
within tolerance
C5
Reported
1,037 increasing / 1,145 decreasing over 8 vegetative time points (of 32,252 nonzero genes)
Reproduced
not attempted (out of scope: trend-test tool/model and the 8-of-9-day selection are under-specified in the Methods)
m.public.grade.out-of-scope
C6
Reported
HSP70 enrichment 13 observed vs ~4 expected
Reproduced
not attempted (out of scope: online DAVID/PANTHER GO enrichment, not a locally-reproducible pipeline step)
m.public.grade.out-of-scope

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

1.5 M
tokens (I/O) · 98.3 M incl. cache
752 min
runtime · 3.91 CPU-h
3.8 GB
peak RAM
4
HPC jobs
hummel
machine