Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

ConNIS and labeling instability: New statistical methods for improving the detection of essential genes in TraDIS libraries.

PLoS Comput Biol · 2026
L1 83/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
83/100
Reproducibility score
0.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 61% of all assessed papers rank 430 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1. ConNIS (bips-hb/ConNIS_results, R, commit 2fc3f0af) is reproducible directly from the repo: the three real-world TraDIS datasets are bundled in-repo (no Zenodo needed for them), and Table 1's MCC values come from a single (read_count_threshold, trimming_start) cell. Recomputing the authors' realworld_bw25113.R pipeline from the raw insertion counts on «our HPC» gave optimal MCC = ConNIS 0.641, Binomial 0.577, Exp.vs.Gamma 0.489, Geometric 0.519 -> all EXACT matches to Table 1's optimal row (0.64/0.58/0.49/0.52) for E. coli BW25113, including the paper's headline method ConNIS. Independently cross-checked: the recompute also equals the authors' shipped Performance RDS to 3 dp, so no fabrication indication. NOT attempted/finalized: MG1655 (target ConNIS 0.79) and 14028S (target 0.57) were the identical pipeline still computing at the operator's finalize-now cutoff (their RDS not yet written); the instability/selected MCC column (needs stabilities_*.R); the synthetic 160-combo and semi-synthetic studies (heavy, need Zenodo); and the two non-analytic methods InsDens (Bayesian MCMC) + Tn5Gaps (the deliberate 20%). Verdict: PARTIAL coverage, EXACT 1:1 where measured.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.16790977

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 83
    assessed: 2026-06-16 ⛓ c84e367099e7
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a statistically rigorous method that provides an exact probability distribution for intragenic insertion-free sequences, plus a data-driven criterion for setting parameters/thresholds, improve the detection of essential genes in Tn5-based TraDIS libraries compared to existing state-of-the-art methods?

Core claims
  • ConNIS provides an analytic solution for the probability of observing the longest insertion-free sequence within a gene given its length and number of insertion sites under non-essentiality. method
  • ConNIS outperforms five state-of-the-art Tn5 essentiality methods, particularly for low- and medium-density libraries. finding
  • Incorporating a weighting factor w for the genome-wide insertion density improves precision of existing methods by reducing false positives without losing many true positives. method
  • A subsample-based labeling instability criterion effectively selects well-suited parameter/threshold values across TIS methods. method
  • ConNIS reliably detects essential genes even among short genes that competing methods typically exclude. finding
  • The geometric distribution is the limiting distribution of ConNIS. mechanism
  • An R package and interactive web application are provided to facilitate application and reproducibility. resource
  • Gene essentiality is declared when the ConNIS probability is at or below significance level alpha, with Bonferroni-Holm (FWER) or Benjamini-Hochberg (FDR) multiple-testing correction. method
Experimental setups
Assay System Perturbation Readout Platform
Simulation study (synthetic TraDIS/Tn5 insertion-site data) in silico simulated bacterial genomes none (160 parameter combinations mimicking different data-generating processes) essential-gene classification performance
Semi-synthetic dataset analysis four semi-synthetic TraDIS datasets none essential-gene labeling performance
Real-world TraDIS (Tn5) data analysis three real bacterial datasets (e.g., Escherichia coli; Keio library reference) transposon mutagenesis (Tn5 insertion) identification/classification of essential genes
Key results
  • ConNIS was superior to five state-of-the-art Tn5 methods, especially in low- and medium-density libraries
  • Weighting factor w improved three competing methods by reducing false positives without losing too many true positives
  • Instability criterion successfully selected suitable parameter/threshold values across all methods in real and synthetic settings
  • ConNIS detected essential genes among short genes where competing methods could not distinguish signal from noise
Key statistics
  • count 160 parameter combinations (simulation study mimicking different data-generating processes)
  • count five state-of-the-art Tn5 analysis methods compared (Binomial, Exp. vs. Gamma, InsDens, Tn5Gaps, Geometric)
  • count four semi-synthetic datasets (additional evaluation datasets)
  • count three real datasets (real-world evaluation of methods)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This methods paper introduces ConNIS, a novel frequentist approach for identifying essential genes in Tn5-based TraDIS transposon insertion sequencing data by deriving an exact combinatorial probability distribution for the longest insertion-free run within a gene of given length and expected insertion count, yielding a per-gene p-value corrected for multiple testing via Bonferroni(-Holm) or Benjamini-Hochberg FDR. Performance was benchmarked against five state-of-the-art methods across an extensive simulation study of 160 parameter combinations, four semi-synthetic datasets, and three real-world datasets. A subsampling-based labeling instability criterion is additionally proposed as a data-driven approach to selecting threshold and tuning parameters for any TIS method. The paper text provided is truncated before the results section, so dispersion, exact p-value, and effect-size reporting practices cannot be confirmed from the available excerpt.

Replicationunclear Sample size160 simulation parameter combinations; 4 semi-synthetic datasets; 3 real-world datasets; m subsamples of size h_sub < h drawn without replacement for instability criterion (m not numerically fixed in available excerpt) GroupsEssential vs. non-essential gene classifications from ConNIS versus five competing Tn5 analysis methods, across a gradient of insertion densities and gene lengths Pairingna Randomization/blindingnot stated Dispersionunclear Multiplicity correctionBonferroni(-Holm) for FWER; Benjamini-Hochberg for FDR
Statistical tests used
Test Applied to n Assumptions
ConNIS: exact combinatorial survival probability P(L_j ≥ l_j) based on a novel discrete distribution for the longest insertion-free run, declared essential if ≤ significance level α Per-gene essentiality classification across all simulated and real datasets Genome-wide insertion site count h and gene length b_j (dataset-specific; not numerically stated in excerpt) stated
Binomial distribution test (TSAS 2.0) Competing method: per-gene essentiality classification using genome-wide insertion density as success probability not stated
Gumbel distribution approximation (Tn5Gaps, TRANSIT package) Competing method: essentiality declared by largest insertion-free gap within or partially overlapping a gene not stated
Bimodal Exponential vs. Gamma mixture distribution fit with log2 likelihood ratio threshold (Bio-TraDIS) Competing method: gene labeling as essential/non-essential/ambiguous via a priori log2 likelihood ratio threshold on gene-wise insertion density not stated
Bayesian posterior probability of essentiality with decision-theory threshold (InsDens) Competing method: gene labeling using posterior probability given prior hyperparameters not stated
Geometric distribution (limiting case of ConNIS; previously used in insertion-free region analysis) Competing method: probability of insertion-free genomic regions not stated
Approaches that could also have been used
  • ConNIS derives an exact parametric distribution for the longest insertion-free run, conditioning on the expected number of insertions under the genome-wide density θ weighted by scalar w
    Could also: A permutation-based null distribution could also be constructed by repeatedly randomising the observed genome-wide insertion positions and empirically recording the longest per-gene gap, yielding a non-parametric p-value — A permutation approach relaxes the parametric uniform-insertion assumption and automatically reflects the actual marginal insertion density; it would serve as an assumption-free reference point, especially in genomes with strong spatial clustering of insertion sites
  • The labeling instability criterion quantifies average variation in gene labels across m subsamples drawn without replacement from the observed insertion set
    Could also: Bootstrap resampling with replacement, or a leave-one-out jackknife over insertion sites, could also be used to quantify sensitivity of gene labels to the observed insertion set — Bootstrap methods have established asymptotic properties for variance estimation and may perform better in settings with very low insertion counts (h_sub close to h) where sampling without replacement approaches a fixed design; comparing both would test robustness of the instability criterion itself
  • The weight parameter w is selected by minimising average labeling instability over a grid of candidate values w_1, …, w_z
    Could also: Empirical Bayes shrinkage of the local insertion density (e.g., smoothing θ_j toward the genome-wide θ using a hierarchical model) or kernel-smoothed local density estimation could also be used to adjust for non-uniform insertion density without a scalar grid search — Local density estimation directly models spatial heterogeneity at base-pair resolution rather than applying a single global correction factor, and could reduce both false positives in coldspots and false negatives near hotspot boundaries without requiring a separate tuning step
  • Multiple testing correction for p-value-based methods is offered as a choice between Bonferroni(-Holm) (FWER) and Benjamini-Hochberg (FDR), with the selection left to the analyst
    Could also: Storey's q-value procedure, which estimates the proportion of true nulls π₀ from the p-value distribution, could also be applied; or the Benjamini-Yekutieli correction for positively dependent tests could be considered given that adjacent genomic loci share insertion context — Storey's q-value gains power over BH when π₀ is substantially below 1 (i.e., when many genes are truly non-essential), which is the typical case in bacterial genomes; BY provides formal FWER-like guarantees under dependency that BH does not
  • Simulation performance was evaluated across 160 fixed parameter combinations on a structured grid of insertion densities and genome configurations
    Could also: A Latin hypercube sampling (LHS) or quasi-random Sobol sequence design over the same parameter space could also be used to achieve broad, space-filling coverage with fewer combinations — Space-filling designs improve coverage of interaction regions between parameters (e.g., joint effects of insertion density and gene length on false-positive rates) and facilitate fitting response-surface meta-models that summarise each method's performance landscape across the full parameter space
  • Gene essentiality is determined by a hard binary decision (essential vs. non-essential) at a fixed significance threshold α after multiple testing correction
    Could also: A fully Bayesian formulation — placing a prior on the probability of essentiality and deriving a posterior for each gene using the exact ConNIS likelihood — could also be used to produce continuous posterior probabilities rather than a binary call — A Bayesian approach unifies threshold selection, uncertainty in θ, and prior knowledge about the proportion of essential genes in a single coherent model, naturally producing a ranked list of genes with calibrated uncertainty; this is the direction already explored by InsDens for the competing bimodal density approach
Software: R (ConNIS package, bips-hb/ConNIS) · insdens R package (Kevin-walters/insdens) commit 286f114 · Bio-TraDIS · TRANSIT (Tn5Gaps method) · TSAS 2.0 (Binomial method) · DBSCAN R package (mentioned as prior comparator)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41790830 (ConNIS)

Paper: Hanke M, Harten T, Foraita R. ConNIS and labeling instability: New statistical methods for improving the detection of essential genes in TraDIS libraries. PLoS Comput Biol 2026. PMID 41790830 · PMCID PMC12991369 · DOI 10.1371/journal.pcbi.1013428.

Code: https://github.com/bips-hb/ConNIS_results (R; commit pinned at run time). Data: Bundled inside the repo (real-world strains + reference essential-gene sets) + Zenodo 10.5281/zenodo.16790977 (synthetic/semi-synthetic). The three real-world datasets are self-contained in the repo — no external download is needed for the real-world reproduction.

What the paper reports (pipeline-derived)

The paper compares six essential-gene detection methods on TraDIS data: ConNIS (novel), Binomial (TSAS 2.0), Geometric, Exp. vs. Gamma (Bio-TraDIS), InsDens (Bayesian MCMC), Tn5Gaps (TRANSIT). Performance metric = MCC. Results: a synthetic simulation study (160 param combos), three real-world strains, a semi-synthetic study, and a "gene-labeling instability" tuning criterion. Headline numbers are the MCC values in Table 1 ("Tuning performance of the gene labeling instability criterion") and Fig 4 (real-world). Each Table 1 cell has a selected/instability value and a best/optimal value (max MCC over the tuning grid).

In scope (attempted) — real-world, fast analytic methods

For each real-world strain the repo ships the raw insertion data + reference essential genes, and realworld_<strain>.R recomputes all-method performance into performance/Performance_<strain>_realWorld.RDS. Table 1 reads from these with the filter read_count_threshold == {0|1} & trimming_start == 0.05.

We recompute, from the bundled raw data, the optimal (max) MCC of the four analytic, fast methods — ConNIS, Binomial, Geometric, Exp. vs. Gamma — for the exact filter the paper uses, and compare to Table 1's "optimal" column:

strain filter (rc, trim_start) reported optimal MCC: ConNIS / Binomial / ExpVsGamma / Geometric
BW25113 rc=0, 0.05 0.64 / 0.58 / 0.49 / 0.52
MG1655 rc=0, 0.05 0.79 / 0.79 / 0.79 / 0.74
14028S rc=1, 0.05 0.57 / 0.53 / 0.41 / 0.48

ConNIS is the paper's own method and headline; the other three analytic methods are reproduced as in-context competitors. Pipeline = the authors' R scripts (functions.R method implementations + realworld_<strain>.R driver), run unmodified except for (a) restricting the parameter loops to the single (read_count_threshold, trimming) cell that Table 1 actually uses, and (b) dropping the two non-analytic methods (see below).

Out of scope / deliberately skipped (the hard ~20%)

  • InsDens (Bayesian MCMC via the insdens CRAN package) and Tn5Gaps (TRANSIT port): slow, depend on an external MCMC package, and are not ConNIS. Dropped to keep compute bounded; documented, not reproduced.
  • Synthetic simulation study (160 param combos, 64-core workstation in the original) and semi-synthetic study: heavy compute, needs Zenodo data. Not attempted (80/20).
  • Instability-criterion (selected) MCC: requires the stabilities_*.R resampling pipeline. We target the optimal MCC column (max over tuning), which for ConNIS equals or closely brackets the reported instability value (BW25113 0.64=0.64; 14028S 0.57=0.57; MG1655 optimal 0.79). Not the selected column itself.
  • Wet-lab / manual content: none (paper is fully computational).

Reproduction host

All compute on «our HPC» («infra») via SLURM; R env built with conda on the compute node («infra» prefix). «host» holds results/pointers only.

Figures / tables: TableFig 4
rw_bw25113_connis_opt
Reported
0.64
Reproduced
0.641
exact
rw_bw25113_binom_opt
Reported
0.58
Reproduced
0.577
exact
rw_bw25113_expgam_opt
Reported
0.49
Reproduced
0.489
exact
rw_bw25113_geom_opt
Reported
0.52
Reproduced
0.519
exact
rw_mg1655_connis_opt
Reported
0.79
Reproduced
partial
rw_14028s_connis_opt
Reported
0.57
Reproduced
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 83/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

For the measured E. coli BW25113 real-world claims, this is a clean 1:1 reproduction: all four analytic methods (incl. the headline ConNIS 0.641 vs reported 0.64) match Table 1 to rounding, and an independent recompute from the raw bundled insertion counts also equals the authors' shipped RDS to 3 decimals — no fabrication concern, data fully derivable. The only limitation is partial coverage (MG1655, 14028S, and the synthetic/semi-synthetic studies were still computing at the finalize cutoff), which is a gap on our side, not a defect of the paper. Net: green where measured, with an honest coverage caveat.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

160.2 k
tokens (I/O) · 12.1 M incl. cache
32 min
runtime · 0.41 CPU-h
1.9 GB
peak RAM
1 (1 failed)
HPC jobs
hummel
machine