Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comparing the utility of in vivo transposon mutagenesis approaches in yeast species to infer gene essentiality.

Curr Genet · 2020
L1 81/100 PQI 94
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
81/100
Reproducibility score
0.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 59% of all assessed papers rank 468 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce — 1:1 on the core. Real code = github.com/berman-lab/transposon-pipeline @54a6008 (the scaffold's code=pysam / data=SRR089408 were link-mining artifacts: pysam is just a dependency, SRR089408 is only the SpPB run of ~13). PIPELINE: cutadapt->Bowtie2->samtools->insertion calling->analyze_hits feature engineering->RandomForest 5-fold-CV essentiality. REPRODUCED EXACTLY from the repo's shipped intermediates: Table 2 unique-insertions, total-reads and mean-reads/insertion for BOTH shipped datasets — SpHermes (382818 / 23919798 / 62.5) and ScAcDs (514888 / 47096808 / 91.5) — 6 values, all exact. REPRODUCED CLOSE: the RandomForest ROC AUC (rebuilt from the repo's own analyze_hits + 5-fold StratifiedKFold RF under py2.7/sklearn 0.19.2) — ScAcDs 0.972 vs 0.99 (within-tol), SpHermes 0.903 vs 0.96 (partial); correct magnitude and ordering, systematic ~0.02-0.06 shortfall because the headline AUCs use a curated cross-species 'core' training set that Classifier.main() builds from C. albicans experiment hit files NOT shipped in the repo. RAW-READ check: re-derived SpHermes insertions from raw SRA SRR327340 (Bowtie2 local + ProcessPombeBam): reads 26.0M within ~9% of 23.92M, insertions 490992 (~1.28x of 382818) — exact match blocked by the authors' undocumented per-dataset transposon-trim. NO fabrication indicators: all reported Table-2 numbers are exactly derivable from shipped data, AUCs reproduce in magnitude+ordering, raw data yields the reported scale. NOT ATTEMPTED: CaAcDs/CaPB/ScHermes/SpPB Table-2 rows (need SRA downloads + per-dataset trim seqs), Fig-4a r=0.892 correlation (needs all 6 AUCs), genome-wide TTAA/TnnnnA target-seq coverage, and the manual/wet-lab claims (haploinsufficiency NDC1/MLC1/BCY1; 29/74 genomic-aberration call). Blocker noted: Classifier.main() is not runnable end-to-end (unshipped albicans hit files + hardcoded «path» figure paths).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 81
    assessed: 2026-06-17 ⛓ 034860ee90d6
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-17
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can in vivo transposon mutagenesis data, analyzed with machine learning, reliably predict gene essentiality across different yeast species and transposon systems, and which data features and transposon choices yield accurate predictions?

Core claims
  • A Random Forest machine-learning approach can predict gene essentiality from in vivo transposon insertion data across multiple yeast species and transposon systems method
  • Sufficient numbers and distribution of independent insertion events are important data features for accurate essentiality prediction finding
  • All three transposons (AcDs, Hermes, PiggyBac) show insertion-site biases due to jackpot events, sequence preferences, and short- vs long-distance insertion preferences finding
  • PiggyBac's stringent TTAA target sequence limits the ability to predict essentiality in genes with few or no target sequences finding
  • The ML approach predicts gene function in less well-studied species by leveraging cross-species orthologs method
  • Comparison of isogenic diploid versus haploid S. cerevisiae identifies several haplo-insufficient genes, while most essential genes are recessive as expected finding
  • The study provides recommendations for transposon choice and inference of gene essentiality in genome-wide studies of haploid yeasts resource
Experimental setups
Assay System Perturbation Readout Platform
In vivo Hermes transposon mutagenesis with Tn-seq (deep sequencing) S. cerevisiae haploid (BY4741, BY4742) and diploid (BY4743) transposon insertion mutagenesis genome-wide transposon insertion sites and read counts mapped to ORFs Illumina MiSeq; pSG36 plasmid; Quick-DNA Fungal/Bacterial Miniprep kit (Zymo); Bowtie2
In vivo AcDs (mini-Ds) transposon mutagenesis with deep sequencing S. cerevisiae haploid transposon insertion mutagenesis transposon insertion maps (WT1 and WT2 combined)
In vivo PiggyBac transposon mutagenesis with deep sequencing S. pombe transposon insertion mutagenesis genome-wide transposon insertion sites mapped to ORFs SRA dataset SRR089408
In vivo Hermes transposon mutagenesis with deep sequencing S. pombe transposon insertion mutagenesis transposon insertion maps SRA dataset SRR327340
In vivo AcDs transposon mutagenesis with deep sequencing C. albicans transposon insertion mutagenesis genome-wide insertion sites mapped to ORFs SRA datasets SRR7824843/SRR7824841/SRR7824838 combined
In vivo PiggyBac transposon mutagenesis with deep sequencing C. albicans transposon insertion mutagenesis (DMSO/5-FOA/no-drug) genome-wide insertion sites mapped to ORFs SRA datasets SRR7704188-SRR7704200
Random Forest machine-learning classification of gene essentiality S. cerevisiae, S. pombe, C. albicans genomes none (computational) binary essentiality prediction per ORF with Youden Index threshold Python scikit-learn (n_estimators=200, random_state=0, fivefold cross-validation)
Key results
  • Hermes mutagenesis protocol yielded cells bearing transposon insertions at ~3% of all cells before enrichment ~5 × 10^6 insertion-bearing cells per mL (~3%)
  • Genes within duplicated/repeated regions appear to have few insertions and are falsely predicted essential, requiring manual curation removal
  • Genes shorter than 300 bp have lower probability of insertions and are more likely falsely predicted essential, so were removed
  • PiggyBac strong preference for TTAA sequences (more frequent in AT-rich intergenic regions than coding) limits essentiality prediction
  • Diploid vs haploid S. cerevisiae comparison identified several haplo-insufficient genes while most essential genes were recessive
  • Number of transposition events detected varied considerably across the six studies
Key statistics
  • count ~5 × 10^6 cells bearing transposon insertions per mL (~3% of all cells) (Sc Hermes mutagenesis yield before enrichment)
  • count 6 in vivo transposon mutagenesis studies (3 species × 2 transposons each) (datasets compared in study)
  • count 300 bp minimum gene length threshold (genes shorter removed in manual curation)
  • count ~300–600 essential genes per bacterium (cited prior bacterial transposon study (Price et al. 2018))
  • count quality score < 20 reads removed; n_estimators=200; random_state=0 (Bowtie2 mapping filter and Random Forest parameters)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study applied a Random Forest machine-learning classifier (scikit-learn) to predict gene essentiality genome-wide from six in vivo transposon mutagenesis datasets across three yeast species and three transposon systems, using eight engineered genomic insertion features as inputs. Classifier performance was evaluated with fivefold cross-validation and ROC curve analysis; binary decision thresholds were selected via the Youden Index. Supplementary pairwise statistical comparisons used Mann–Whitney U tests and Pearson correlation coefficients computed with Python/SciPy. Results were reported primarily as classification metrics and genomic visualizations rather than traditional group-level inferential statistics.

Replicationmixed Sample sizeNumber of transposition events described as varying considerably across datasets; for the new Sc Hermes dataset, the diploid strain BY4743 was run three times and haploid strains BY4741 and BY4742 once and twice respectively, with runs combined prior to analysis. Specific per-dataset total insertion counts not fully enumerated in provided excerpt. GroupsEssential vs. non-essential genes; three yeast species (S. cerevisiae, S. pombe, C. albicans); three transposon types (AcDs, Hermes, PiggyBac); haploid vs. diploid S. cerevisiae isolates Pairingunpaired Randomization/blindingnot stated Dispersionnone Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Random Forest classification Genome-wide binary prediction of gene essentiality from eight transposon insertion features, applied to all six species-by-transposon datasets not stated
Fivefold cross-validation Validation of Random Forest classifier performance for each of the six datasets not stated
ROC curve analysis with Youden Index threshold selection Threshold optimization for binary essentiality classification in each dataset; Euclidean-distance-to-(0,1) metric used as a secondary check not stated
Mann–Whitney U test Distributional comparisons across datasets or gene groups (specific comparisons not fully enumerated in the provided text excerpt) not stated
Pearson correlation coefficient Correlation analyses between features or dataset properties (specific pairings not fully enumerated in the provided text excerpt) not stated
Approaches that could also have been used
  • Gene essentiality was predicted using a Random Forest classifier with largely default scikit-learn parameters (n_estimators=200, random_state=0 fixed for reproducibility)
    Could also: Gradient boosting methods (e.g., XGBoost, LightGBM) or logistic regression with L1/L2 regularization could also be applied to the same feature matrix — Gradient boosting frequently achieves competitive or superior AUC on structured tabular data; logistic regression would additionally yield interpretable feature coefficients and calibrated probability estimates, which could facilitate direct comparison of which insertion features most strongly distinguish essential from non-essential genes across species
  • Classifier performance was evaluated with fivefold cross-validation and ROC-AUC, and a single fixed random seed was used for reproducibility
    Could also: Stratified k-fold cross-validation (preserving the essential/non-essential class ratio in each fold) or repeated cross-validation with multiple seeds could also be used; reporting precision–recall AUC (AUPRC) alongside ROC-AUC is another common complement — With the likely class imbalance between essential and non-essential genes in eukaryotic genomes, AUPRC is often more sensitive to minority-class performance than ROC-AUC; stratification and repetition provide more stable variance estimates of generalization performance
  • Mann–Whitney U tests were used for distributional comparisons across datasets or gene groups without a described multiple-testing correction
    Could also: A Benjamini–Hochberg FDR correction or Bonferroni correction could also be applied when multiple comparisons span datasets, species, or feature comparisons — Applying a correction procedure explicitly controls the expected rate of false positives across the family of tests, making the inferential interpretation clearer when many simultaneous comparisons are reported
  • Pearson correlation coefficients were used to quantify associations between variables
    Could also: Spearman rank correlation could also be computed for the same variable pairs — Spearman correlation makes no assumption of linearity or normality and is more robust to outliers; for count-based insertion features, which tend to be right-skewed, Spearman is a common complement or alternative to Pearson
  • Binary classification thresholds were selected by maximizing the Youden Index (equal weighting of sensitivity and specificity on the ROC curve)
    Could also: An F1-score-maximizing threshold or a cost-weighted threshold that assigns different penalties to false negatives (missing an essential gene) and false positives (incorrectly labeling a non-essential gene as essential) could also be used — The Youden Index treats sensitivity and specificity symmetrically; if the downstream biological cost of missing a true essential gene differs from that of a false essential call, a cost-sensitive threshold would allow that asymmetry to be encoded explicitly and could shift the operating point accordingly
  • Cross-species essentiality was inferred by leveraging known ortholog labels from well-studied species (S. cerevisiae, S. pombe) as training signal for less-studied species (C. albicans)
    Could also: Formal transfer-learning or domain-adaptation frameworks could also be applied to transfer the classifier across species while explicitly modeling distributional shift in genomic features — Standard transfer-learning methods quantify and correct for feature distribution differences between source and target species, which may yield more calibrated probability estimates when insertion feature distributions differ substantially across species
Software: Python/scikit-learn (Random Forest) · Python/SciPy · Python/matplotlib · Python/seaborn · Bowtie2 · cutadapt · samtools · Biopython · pysam · gffutils

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
24
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

E-MTAB-4885 ArrayExpress in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-32681306

Paper: Levitan, Gale, Dallon, Kozan, Cunningham, Sharan, Berman (2020), "Comparing the utility of in vivo transposon mutagenesis approaches in yeast species to infer gene essentiality." Curr Genet. PMID 32681306 / PMC7599172 / DOI 10.1007/s00294-020-01096-6.

Analysis code (resolved): https://github.com/berman-lab/transposon-pipeline (the harvested code_url=pysam and data=SRR089408 in the scaffold are link-mining artifacts; pysam is a library the repo uses, and SRR089408 is just ONE of ~13 SRA runs = the SpPB dataset). Commit pinned at run time: 54a60088606b90d6c8256b24649026d280500a0a.

Datasets (transposon-seq, 3 yeast species × 3 transposon systems):

Label Species Transposon Source
CaAcDs C. albicans AcDs new (this paper); BioProject runs SRR7824838/41/43
CaPB C. albicans PiggyBac SRR7704188/89/93/94/95/96/200
ScAcDs S. cerevisiae AcDs (MiniDs/SATAY, Michel/Kornmann) ArrayExpress E-MTAB-4885; shipped as Kornmann *WildType*.wig
ScHermes S. cerevisiae Hermes new (this paper), UCSC track (Cunningham lab)
SpHermes S. pombe Hermes SRR327340; shipped as dependencies/pombe/hermes_hits.csv
SpPB S. pombe PiggyBac SRR089408

In scope (pipeline-derived)

Pipeline = cutadapt (strip transposon+primer, --discard-untrimmed) → Bowtie2 (default) → samtools sort → insertion calling (CreateHitFile.py for Ca: -q 20 mapq, -k 2 merge = the paper's "quality <20 / mismatch at nt+1" filter; ProcessPombeBam.py for Sp: unique (chrom,strand,pos) at mapq≥20) → per-gene feature engineering (SummaryTable.analyze_hits: neighborhood index, freedom index, insertions, reads, upstream-100, length) → Random Forest essentiality classifier (Classifier.py: StratifiedKFold(n_folds=5, shuffle, random_state=0), RandomForestClassifier(random_state=0), ROC AUC, Youden index).

  1. Table 2 per-dataset stats — total unique insertions, total reads, mean reads/insertion, genes-with-0-insertions. Pipeline: mapping + insertion calling. Cleanly checkable for the two datasets whose hit/track files are shipped (SpHermes, ScAcDs); checkable from raw reads for SRA datasets.
  2. Table 2 ROC AUC per dataset (0.99/0.99/0.97/0.96/0.94/0.79). Pipeline: RF classifier on engineered features vs literature essentiality labels (FYPO for pombe, SGD viable/inviable for cerevisiae). Cleanly checkable for SpHermes & ScAcDs from shipped data + repo code.
  3. Fig 4a correlation unique-insertions vs AUC (Pearson r=0.892, p=0.0169; r=0.995, p=0.0003 excl. SpPB). Derived from #1+#2 across all 6 datasets.
  4. Target-sequence coverage (% TTAA / TnnnnA sites without insertion).

Out of scope (not pipeline / not attempted)

  • Wet-lab transposon library construction, growth, sequencing.
  • Manual curation of false-positive/false-negative gene lists (hardcoded in Classifier.main()); we use them but did not re-derive them.
  • Haploinsufficiency claims (NDC1/MLC1/BCY1) — manual genomic-aberration review.
  • The genome-aberration call (29/74 genes) — manual.

Reproducibility blockers (honest)

  • Classifier.main() is not runnable as-is: it requires unshipped C. albicans experiment hit files (dependencies/albicans/experiment data/post evo/q20m2) and hardcoded macOS figure paths («path»). We therefore reproduce the AUC by reusing the repo's own functions (analyze_hits, the 5-fold-CV test_classifier logic) on the shipped data, rather than running main() end to end.
  • ScHermes raw data is not in the repo (UCSC track only); CaAcDs/CaPB/SpPB require SRA downloads + per-dataset transposon trim sequences (not all documented) → attempted as stretch, lower confidence. </content>
Figures / tables: Table
sphermes_insertions
Reported
382.82e3
Reproduced
382818
exact
sphermes_reads
Reported
23.92e6
Reproduced
23919798
exact
sphermes_avg_reads
Reported
62
Reproduced
62.48
exact
scacds_insertions
Reported
514.89e3
Reproduced
514888
exact
scacds_reads
Reported
47.10e6
Reproduced
47096808
exact
scacds_avg_reads
Reported
91
Reproduced
91.47
exact
scacds_auc
Reported
0.99
Reproduced
0.972
within tolerance
sphermes_auc
Reported
0.96
Reproduced
0.903
partial
scacds_genes0ins
Reported
261
Reproduced
138-336
partial
sphermes_genes0ins
Reported
99
Reproduced
41-185
partial
sphermes_remap_reads
Reported
23.92e6
Reproduced
26004183
within tolerance
sphermes_remap_insertions
Reported
382.82e3
Reproduced
490992
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 81/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

The core of the study reproduces cleanly: all six Table-2 quantities (SpHermes 382818/23919798/62.5; ScAcDs 514888/47096808/91.5) are exactly derivable from the authors' shipped intermediates, and the essentiality-classifier AUCs reproduce in magnitude and the reported ScAcDs > SpHermes ordering (0.972 vs 0.99; 0.903 vs 0.96). The remaining deviations are all on the input/preprocessing side and explainable — an undocumented per-dataset transposon-trim (raw remap ~1.28x on insertions), an unshipped curated cross-species training set (AUC shortfall), and an unpinned ORF-set definition (genes-0-ins bracketed). This is partly our self-chosen methodology and partly authors' underspecification/non-deposition, with no fabrication indicators and the central conclusion fully intact — hence overall solid-but-not-1:1 (yellow).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

371.1 k
tokens (I/O) · 39.1 M incl. cache
43 min
runtime · 0.69 CPU-h
7.2 GB
peak RAM
7 (2 failed)
HPC jobs
hummel
machine