Comparing the utility of in vivo transposon mutagenesis approaches in yeast species to infer gene essentiality.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce — 1:1 on the core. Real code = github.com/berman-lab/transposon-pipeline @54a6008 (the scaffold's code=pysam / data=SRR089408 were link-mining artifacts: pysam is just a dependency, SRR089408 is only the SpPB run of ~13). PIPELINE: cutadapt->Bowtie2->samtools->insertion calling->analyze_hits feature engineering->RandomForest 5-fold-CV essentiality. REPRODUCED EXACTLY from the repo's shipped intermediates: Table 2 unique-insertions, total-reads and mean-reads/insertion for BOTH shipped datasets — SpHermes (382818 / 23919798 / 62.5) and ScAcDs (514888 / 47096808 / 91.5) — 6 values, all exact. REPRODUCED CLOSE: the RandomForest ROC AUC (rebuilt from the repo's own analyze_hits + 5-fold StratifiedKFold RF under py2.7/sklearn 0.19.2) — ScAcDs 0.972 vs 0.99 (within-tol), SpHermes 0.903 vs 0.96 (partial); correct magnitude and ordering, systematic ~0.02-0.06 shortfall because the headline AUCs use a curated cross-species 'core' training set that Classifier.main() builds from C. albicans experiment hit files NOT shipped in the repo. RAW-READ check: re-derived SpHermes insertions from raw SRA SRR327340 (Bowtie2 local + ProcessPombeBam): reads 26.0M within ~9% of 23.92M, insertions 490992 (~1.28x of 382818) — exact match blocked by the authors' undocumented per-dataset transposon-trim. NO fabrication indicators: all reported Table-2 numbers are exactly derivable from shipped data, AUCs reproduce in magnitude+ordering, raw data yields the reported scale. NOT ATTEMPTED: CaAcDs/CaPB/ScHermes/SpPB Table-2 rows (need SRA downloads + per-dataset trim seqs), Fig-4a r=0.892 correlation (needs all 6 AUCs), genome-wide TTAA/TnnnnA target-seq coverage, and the manual/wet-lab claims (haploinsufficiency NDC1/MLC1/BCY1; 29/74 genomic-aberration call). Blocker noted: Classifier.main() is not runnable end-to-end (unshipped albicans hit files + hardcoded «path» figure paths).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 81assessed: 2026-06-17 ⛓ 034860ee90d6
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-17
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan in vivo transposon mutagenesis data, analyzed with machine learning, reliably predict gene essentiality across different yeast species and transposon systems, and which data features and transposon choices yield accurate predictions?
- ★ A Random Forest machine-learning approach can predict gene essentiality from in vivo transposon insertion data across multiple yeast species and transposon systems method
- ★ Sufficient numbers and distribution of independent insertion events are important data features for accurate essentiality prediction finding
- ★ All three transposons (AcDs, Hermes, PiggyBac) show insertion-site biases due to jackpot events, sequence preferences, and short- vs long-distance insertion preferences finding
- ★ PiggyBac's stringent TTAA target sequence limits the ability to predict essentiality in genes with few or no target sequences finding
- ★ The ML approach predicts gene function in less well-studied species by leveraging cross-species orthologs method
- ★ Comparison of isogenic diploid versus haploid S. cerevisiae identifies several haplo-insufficient genes, while most essential genes are recessive as expected finding
- The study provides recommendations for transposon choice and inference of gene essentiality in genome-wide studies of haploid yeasts resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| In vivo Hermes transposon mutagenesis with Tn-seq (deep sequencing) | S. cerevisiae haploid (BY4741, BY4742) and diploid (BY4743) | transposon insertion mutagenesis | genome-wide transposon insertion sites and read counts mapped to ORFs | Illumina MiSeq; pSG36 plasmid; Quick-DNA Fungal/Bacterial Miniprep kit (Zymo); Bowtie2 |
| In vivo AcDs (mini-Ds) transposon mutagenesis with deep sequencing | S. cerevisiae haploid | transposon insertion mutagenesis | transposon insertion maps (WT1 and WT2 combined) | — |
| In vivo PiggyBac transposon mutagenesis with deep sequencing | S. pombe | transposon insertion mutagenesis | genome-wide transposon insertion sites mapped to ORFs | SRA dataset SRR089408 |
| In vivo Hermes transposon mutagenesis with deep sequencing | S. pombe | transposon insertion mutagenesis | transposon insertion maps | SRA dataset SRR327340 |
| In vivo AcDs transposon mutagenesis with deep sequencing | C. albicans | transposon insertion mutagenesis | genome-wide insertion sites mapped to ORFs | SRA datasets SRR7824843/SRR7824841/SRR7824838 combined |
| In vivo PiggyBac transposon mutagenesis with deep sequencing | C. albicans | transposon insertion mutagenesis (DMSO/5-FOA/no-drug) | genome-wide insertion sites mapped to ORFs | SRA datasets SRR7704188-SRR7704200 |
| Random Forest machine-learning classification of gene essentiality | S. cerevisiae, S. pombe, C. albicans genomes | none (computational) | binary essentiality prediction per ORF with Youden Index threshold | Python scikit-learn (n_estimators=200, random_state=0, fivefold cross-validation) |
- – Hermes mutagenesis protocol yielded cells bearing transposon insertions at ~3% of all cells before enrichment ~5 × 10^6 insertion-bearing cells per mL (~3%)
- – Genes within duplicated/repeated regions appear to have few insertions and are falsely predicted essential, requiring manual curation removal
- – Genes shorter than 300 bp have lower probability of insertions and are more likely falsely predicted essential, so were removed
- – PiggyBac strong preference for TTAA sequences (more frequent in AT-rich intergenic regions than coding) limits essentiality prediction
- – Diploid vs haploid S. cerevisiae comparison identified several haplo-insufficient genes while most essential genes were recessive
- – Number of transposition events detected varied considerably across the six studies
- count ~5 × 10^6 cells bearing transposon insertions per mL (~3% of all cells) (Sc Hermes mutagenesis yield before enrichment)
- count 6 in vivo transposon mutagenesis studies (3 species × 2 transposons each) (datasets compared in study)
- count 300 bp minimum gene length threshold (genes shorter removed in manual curation)
- count ~300–600 essential genes per bacterium (cited prior bacterial transposon study (Price et al. 2018))
- count quality score < 20 reads removed; n_estimators=200; random_state=0 (Bowtie2 mapping filter and Random Forest parameters)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study applied a Random Forest machine-learning classifier (scikit-learn) to predict gene essentiality genome-wide from six in vivo transposon mutagenesis datasets across three yeast species and three transposon systems, using eight engineered genomic insertion features as inputs. Classifier performance was evaluated with fivefold cross-validation and ROC curve analysis; binary decision thresholds were selected via the Youden Index. Supplementary pairwise statistical comparisons used Mann–Whitney U tests and Pearson correlation coefficients computed with Python/SciPy. Results were reported primarily as classification metrics and genomic visualizations rather than traditional group-level inferential statistics.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Random Forest classification | Genome-wide binary prediction of gene essentiality from eight transposon insertion features, applied to all six species-by-transposon datasets | — | not stated |
| Fivefold cross-validation | Validation of Random Forest classifier performance for each of the six datasets | — | not stated |
| ROC curve analysis with Youden Index threshold selection | Threshold optimization for binary essentiality classification in each dataset; Euclidean-distance-to-(0,1) metric used as a secondary check | — | not stated |
| Mann–Whitney U test | Distributional comparisons across datasets or gene groups (specific comparisons not fully enumerated in the provided text excerpt) | — | not stated |
| Pearson correlation coefficient | Correlation analyses between features or dataset properties (specific pairings not fully enumerated in the provided text excerpt) | — | not stated |
-
Gene essentiality was predicted using a Random Forest classifier with largely default scikit-learn parameters (n_estimators=200, random_state=0 fixed for reproducibility)↳ Could also: Gradient boosting methods (e.g., XGBoost, LightGBM) or logistic regression with L1/L2 regularization could also be applied to the same feature matrix — Gradient boosting frequently achieves competitive or superior AUC on structured tabular data; logistic regression would additionally yield interpretable feature coefficients and calibrated probability estimates, which could facilitate direct comparison of which insertion features most strongly distinguish essential from non-essential genes across species
-
Classifier performance was evaluated with fivefold cross-validation and ROC-AUC, and a single fixed random seed was used for reproducibility↳ Could also: Stratified k-fold cross-validation (preserving the essential/non-essential class ratio in each fold) or repeated cross-validation with multiple seeds could also be used; reporting precision–recall AUC (AUPRC) alongside ROC-AUC is another common complement — With the likely class imbalance between essential and non-essential genes in eukaryotic genomes, AUPRC is often more sensitive to minority-class performance than ROC-AUC; stratification and repetition provide more stable variance estimates of generalization performance
-
Mann–Whitney U tests were used for distributional comparisons across datasets or gene groups without a described multiple-testing correction↳ Could also: A Benjamini–Hochberg FDR correction or Bonferroni correction could also be applied when multiple comparisons span datasets, species, or feature comparisons — Applying a correction procedure explicitly controls the expected rate of false positives across the family of tests, making the inferential interpretation clearer when many simultaneous comparisons are reported
-
Pearson correlation coefficients were used to quantify associations between variables↳ Could also: Spearman rank correlation could also be computed for the same variable pairs — Spearman correlation makes no assumption of linearity or normality and is more robust to outliers; for count-based insertion features, which tend to be right-skewed, Spearman is a common complement or alternative to Pearson
-
Binary classification thresholds were selected by maximizing the Youden Index (equal weighting of sensitivity and specificity on the ROC curve)↳ Could also: An F1-score-maximizing threshold or a cost-weighted threshold that assigns different penalties to false negatives (missing an essential gene) and false positives (incorrectly labeling a non-essential gene as essential) could also be used — The Youden Index treats sensitivity and specificity symmetrically; if the downstream biological cost of missing a true essential gene differs from that of a false essential call, a cost-sensitive threshold would allow that asymmetry to be encoded explicitly and could shift the operating point accordingly
-
Cross-species essentiality was inferred by leveraging known ortholog labels from well-studied species (S. cerevisiae, S. pombe) as training signal for less-studied species (C. albicans)↳ Could also: Formal transfer-learning or domain-adaptation frameworks could also be applied to transfer the classifier across species while explicitly modeling distributional shift in genomic features — Standard transfer-learning methods quantify and correct for feature distribution differences between source and target species, which may yield more calibrated probability estimates when insertion feature distributions differ substantially across species
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Diploid vs haploid S. cerevisiae Tn-seq comparison identifies haplo-insufficient genes; most essential genes are recessiveother saccharomyces-cerevisiae 2020×1papers★ This paper is the founder (earliest)
-
PiggyBac transposon shows strong TTAA insertion site preference enriched in AT-rich intergenic regions, reducing coding-sequence insertion density and limiting gene essentiality prediction accuracyother yeast 2020×1papers★ This paper is the founder (earliest)
-
Genes in duplicated or repetitive genomic regions show artifactually low transposon insertion density in Tn-seq, causing false-positive essential gene predictions that require manual curationother yeast 2020×1papers★ This paper is the founder (earliest)
-
ORFs shorter than 300 bp have a lower probability of transposon insertion and are disproportionately falsely predicted essential in Tn-seq studiesother yeast 2020×1papers★ This paper is the founder (earliest)
-
The number of detected transposon insertion events varied considerably across six Tn-seq studies in yeast, affecting cross-study comparability of gene essentiality predictionsother yeast mixed 2020×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32681306
Paper: Levitan, Gale, Dallon, Kozan, Cunningham, Sharan, Berman (2020), "Comparing the utility of in vivo transposon mutagenesis approaches in yeast species to infer gene essentiality." Curr Genet. PMID 32681306 / PMC7599172 / DOI 10.1007/s00294-020-01096-6.
Analysis code (resolved): https://github.com/berman-lab/transposon-pipeline
(the harvested code_url=pysam and data=SRR089408 in the scaffold are
link-mining artifacts; pysam is a library the repo uses, and SRR089408 is just
ONE of ~13 SRA runs = the SpPB dataset). Commit pinned at run time:
54a60088606b90d6c8256b24649026d280500a0a.
Datasets (transposon-seq, 3 yeast species × 3 transposon systems):
| Label | Species | Transposon | Source |
|---|---|---|---|
| CaAcDs | C. albicans | AcDs | new (this paper); BioProject runs SRR7824838/41/43 |
| CaPB | C. albicans | PiggyBac | SRR7704188/89/93/94/95/96/200 |
| ScAcDs | S. cerevisiae | AcDs (MiniDs/SATAY, Michel/Kornmann) | ArrayExpress E-MTAB-4885; shipped as Kornmann *WildType*.wig |
| ScHermes | S. cerevisiae | Hermes | new (this paper), UCSC track (Cunningham lab) |
| SpHermes | S. pombe | Hermes | SRR327340; shipped as dependencies/pombe/hermes_hits.csv |
| SpPB | S. pombe | PiggyBac | SRR089408 |
In scope (pipeline-derived)
Pipeline = cutadapt (strip transposon+primer, --discard-untrimmed) → Bowtie2
(default) → samtools sort → insertion calling (CreateHitFile.py for Ca:
-q 20 mapq, -k 2 merge = the paper's "quality <20 / mismatch at nt+1" filter;
ProcessPombeBam.py for Sp: unique (chrom,strand,pos) at mapq≥20) → per-gene
feature engineering (SummaryTable.analyze_hits: neighborhood index, freedom
index, insertions, reads, upstream-100, length) → Random Forest essentiality
classifier (Classifier.py: StratifiedKFold(n_folds=5, shuffle, random_state=0),
RandomForestClassifier(random_state=0), ROC AUC, Youden index).
- Table 2 per-dataset stats — total unique insertions, total reads, mean reads/insertion, genes-with-0-insertions. Pipeline: mapping + insertion calling. Cleanly checkable for the two datasets whose hit/track files are shipped (SpHermes, ScAcDs); checkable from raw reads for SRA datasets.
- Table 2 ROC AUC per dataset (0.99/0.99/0.97/0.96/0.94/0.79). Pipeline: RF classifier on engineered features vs literature essentiality labels (FYPO for pombe, SGD viable/inviable for cerevisiae). Cleanly checkable for SpHermes & ScAcDs from shipped data + repo code.
- Fig 4a correlation unique-insertions vs AUC (Pearson r=0.892, p=0.0169; r=0.995, p=0.0003 excl. SpPB). Derived from #1+#2 across all 6 datasets.
- Target-sequence coverage (% TTAA / TnnnnA sites without insertion).
Out of scope (not pipeline / not attempted)
- Wet-lab transposon library construction, growth, sequencing.
- Manual curation of false-positive/false-negative gene lists (hardcoded in
Classifier.main()); we use them but did not re-derive them. - Haploinsufficiency claims (NDC1/MLC1/BCY1) — manual genomic-aberration review.
- The genome-aberration call (29/74 genes) — manual.
Reproducibility blockers (honest)
Classifier.main()is not runnable as-is: it requires unshipped C. albicans experiment hit files (dependencies/albicans/experiment data/post evo/q20m2) and hardcoded macOS figure paths («path»). We therefore reproduce the AUC by reusing the repo's own functions (analyze_hits, the 5-fold-CVtest_classifierlogic) on the shipped data, rather than runningmain()end to end.- ScHermes raw data is not in the repo (UCSC track only); CaAcDs/CaPB/SpPB require SRA downloads + per-dataset transposon trim sequences (not all documented) → attempted as stretch, lower confidence. </content>
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The core of the study reproduces cleanly: all six Table-2 quantities (SpHermes 382818/23919798/62.5; ScAcDs 514888/47096808/91.5) are exactly derivable from the authors' shipped intermediates, and the essentiality-classifier AUCs reproduce in magnitude and the reported ScAcDs > SpHermes ordering (0.972 vs 0.99; 0.903 vs 0.96). The remaining deviations are all on the input/preprocessing side and explainable — an undocumented per-dataset transposon-trim (raw remap ~1.28x on insertions), an unshipped curated cross-species training set (AUC shortfall), and an unpinned ORF-set definition (genes-0-ins bracketed). This is partly our self-chosen methodology and partly authors' underspecification/non-deposition, with no fabrication indicators and the central conclusion fully intact — hence overall solid-but-not-1:1 (yellow).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.