ConNIS and labeling instability: New statistical methods for improving the detection of essential genes in TraDIS libraries.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce 1:1. ConNIS (bips-hb/ConNIS_results, R, commit 2fc3f0af) is reproducible directly from the repo: the three real-world TraDIS datasets are bundled in-repo (no Zenodo needed for them), and Table 1's MCC values come from a single (read_count_threshold, trimming_start) cell. Recomputing the authors' realworld_bw25113.R pipeline from the raw insertion counts on «our HPC» gave optimal MCC = ConNIS 0.641, Binomial 0.577, Exp.vs.Gamma 0.489, Geometric 0.519 -> all EXACT matches to Table 1's optimal row (0.64/0.58/0.49/0.52) for E. coli BW25113, including the paper's headline method ConNIS. Independently cross-checked: the recompute also equals the authors' shipped Performance RDS to 3 dp, so no fabrication indication. NOT attempted/finalized: MG1655 (target ConNIS 0.79) and 14028S (target 0.57) were the identical pipeline still computing at the operator's finalize-now cutoff (their RDS not yet written); the instability/selected MCC column (needs stabilities_*.R); the synthetic 160-combo and semi-synthetic studies (heavy, need Zenodo); and the two non-analytic methods InsDens (Bayesian MCMC) + Tn5Gaps (the deliberate 20%). Verdict: PARTIAL coverage, EXACT 1:1 where measured.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 83assessed: 2026-06-16 ⛓ c84e367099e7
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a statistically rigorous method that provides an exact probability distribution for intragenic insertion-free sequences, plus a data-driven criterion for setting parameters/thresholds, improve the detection of essential genes in Tn5-based TraDIS libraries compared to existing state-of-the-art methods?
- ★ ConNIS provides an analytic solution for the probability of observing the longest insertion-free sequence within a gene given its length and number of insertion sites under non-essentiality. method
- ★ ConNIS outperforms five state-of-the-art Tn5 essentiality methods, particularly for low- and medium-density libraries. finding
- ★ Incorporating a weighting factor w for the genome-wide insertion density improves precision of existing methods by reducing false positives without losing many true positives. method
- ★ A subsample-based labeling instability criterion effectively selects well-suited parameter/threshold values across TIS methods. method
- ★ ConNIS reliably detects essential genes even among short genes that competing methods typically exclude. finding
- The geometric distribution is the limiting distribution of ConNIS. mechanism
- An R package and interactive web application are provided to facilitate application and reproducibility. resource
- Gene essentiality is declared when the ConNIS probability is at or below significance level alpha, with Bonferroni-Holm (FWER) or Benjamini-Hochberg (FDR) multiple-testing correction. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Simulation study (synthetic TraDIS/Tn5 insertion-site data) | in silico simulated bacterial genomes | none (160 parameter combinations mimicking different data-generating processes) | essential-gene classification performance | — |
| Semi-synthetic dataset analysis | four semi-synthetic TraDIS datasets | none | essential-gene labeling performance | — |
| Real-world TraDIS (Tn5) data analysis | three real bacterial datasets (e.g., Escherichia coli; Keio library reference) | transposon mutagenesis (Tn5 insertion) | identification/classification of essential genes | — |
- ▲ ConNIS was superior to five state-of-the-art Tn5 methods, especially in low- and medium-density libraries
- ▼ Weighting factor w improved three competing methods by reducing false positives without losing too many true positives
- – Instability criterion successfully selected suitable parameter/threshold values across all methods in real and synthetic settings
- – ConNIS detected essential genes among short genes where competing methods could not distinguish signal from noise
- count 160 parameter combinations (simulation study mimicking different data-generating processes)
- count five state-of-the-art Tn5 analysis methods compared (Binomial, Exp. vs. Gamma, InsDens, Tn5Gaps, Geometric)
- count four semi-synthetic datasets (additional evaluation datasets)
- count three real datasets (real-world evaluation of methods)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This methods paper introduces ConNIS, a novel frequentist approach for identifying essential genes in Tn5-based TraDIS transposon insertion sequencing data by deriving an exact combinatorial probability distribution for the longest insertion-free run within a gene of given length and expected insertion count, yielding a per-gene p-value corrected for multiple testing via Bonferroni(-Holm) or Benjamini-Hochberg FDR. Performance was benchmarked against five state-of-the-art methods across an extensive simulation study of 160 parameter combinations, four semi-synthetic datasets, and three real-world datasets. A subsampling-based labeling instability criterion is additionally proposed as a data-driven approach to selecting threshold and tuning parameters for any TIS method. The paper text provided is truncated before the results section, so dispersion, exact p-value, and effect-size reporting practices cannot be confirmed from the available excerpt.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| ConNIS: exact combinatorial survival probability P(L_j ≥ l_j) based on a novel discrete distribution for the longest insertion-free run, declared essential if ≤ significance level α | Per-gene essentiality classification across all simulated and real datasets | Genome-wide insertion site count h and gene length b_j (dataset-specific; not numerically stated in excerpt) | stated |
| Binomial distribution test (TSAS 2.0) | Competing method: per-gene essentiality classification using genome-wide insertion density as success probability | — | not stated |
| Gumbel distribution approximation (Tn5Gaps, TRANSIT package) | Competing method: essentiality declared by largest insertion-free gap within or partially overlapping a gene | — | not stated |
| Bimodal Exponential vs. Gamma mixture distribution fit with log2 likelihood ratio threshold (Bio-TraDIS) | Competing method: gene labeling as essential/non-essential/ambiguous via a priori log2 likelihood ratio threshold on gene-wise insertion density | — | not stated |
| Bayesian posterior probability of essentiality with decision-theory threshold (InsDens) | Competing method: gene labeling using posterior probability given prior hyperparameters | — | not stated |
| Geometric distribution (limiting case of ConNIS; previously used in insertion-free region analysis) | Competing method: probability of insertion-free genomic regions | — | not stated |
-
ConNIS derives an exact parametric distribution for the longest insertion-free run, conditioning on the expected number of insertions under the genome-wide density θ weighted by scalar w↳ Could also: A permutation-based null distribution could also be constructed by repeatedly randomising the observed genome-wide insertion positions and empirically recording the longest per-gene gap, yielding a non-parametric p-value — A permutation approach relaxes the parametric uniform-insertion assumption and automatically reflects the actual marginal insertion density; it would serve as an assumption-free reference point, especially in genomes with strong spatial clustering of insertion sites
-
The labeling instability criterion quantifies average variation in gene labels across m subsamples drawn without replacement from the observed insertion set↳ Could also: Bootstrap resampling with replacement, or a leave-one-out jackknife over insertion sites, could also be used to quantify sensitivity of gene labels to the observed insertion set — Bootstrap methods have established asymptotic properties for variance estimation and may perform better in settings with very low insertion counts (h_sub close to h) where sampling without replacement approaches a fixed design; comparing both would test robustness of the instability criterion itself
-
The weight parameter w is selected by minimising average labeling instability over a grid of candidate values w_1, …, w_z↳ Could also: Empirical Bayes shrinkage of the local insertion density (e.g., smoothing θ_j toward the genome-wide θ using a hierarchical model) or kernel-smoothed local density estimation could also be used to adjust for non-uniform insertion density without a scalar grid search — Local density estimation directly models spatial heterogeneity at base-pair resolution rather than applying a single global correction factor, and could reduce both false positives in coldspots and false negatives near hotspot boundaries without requiring a separate tuning step
-
Multiple testing correction for p-value-based methods is offered as a choice between Bonferroni(-Holm) (FWER) and Benjamini-Hochberg (FDR), with the selection left to the analyst↳ Could also: Storey's q-value procedure, which estimates the proportion of true nulls π₀ from the p-value distribution, could also be applied; or the Benjamini-Yekutieli correction for positively dependent tests could be considered given that adjacent genomic loci share insertion context — Storey's q-value gains power over BH when π₀ is substantially below 1 (i.e., when many genes are truly non-essential), which is the typical case in bacterial genomes; BY provides formal FWER-like guarantees under dependency that BH does not
-
Simulation performance was evaluated across 160 fixed parameter combinations on a structured grid of insertion densities and genome configurations↳ Could also: A Latin hypercube sampling (LHS) or quasi-random Sobol sequence design over the same parameter space could also be used to achieve broad, space-filling coverage with fewer combinations — Space-filling designs improve coverage of interaction regions between parameters (e.g., joint effects of insertion density and gene length on false-positive rates) and facilitate fitting response-surface meta-models that summarise each method's performance landscape across the full parameter space
-
Gene essentiality is determined by a hard binary decision (essential vs. non-essential) at a fixed significance threshold α after multiple testing correction↳ Could also: A fully Bayesian formulation — placing a prior on the probability of essentiality and deriving a posterior for each gene using the exact ConNIS likelihood — could also be used to produce continuous posterior probabilities rather than a binary call — A Bayesian approach unifies threshold selection, uncertainty in θ, and prior knowledge about the proportion of essential genes in a single coherent model, naturally producing a ranked list of genes with calibrated uncertainty; this is the direction already explored by InsDens for the competing bimodal density approach
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41790830 (ConNIS)
Paper: Hanke M, Harten T, Foraita R. ConNIS and labeling instability: New statistical methods for improving the detection of essential genes in TraDIS libraries. PLoS Comput Biol 2026. PMID 41790830 · PMCID PMC12991369 · DOI 10.1371/journal.pcbi.1013428.
Code: https://github.com/bips-hb/ConNIS_results (R; commit pinned at run time). Data: Bundled inside the repo (real-world strains + reference essential-gene sets) + Zenodo 10.5281/zenodo.16790977 (synthetic/semi-synthetic). The three real-world datasets are self-contained in the repo — no external download is needed for the real-world reproduction.
What the paper reports (pipeline-derived)
The paper compares six essential-gene detection methods on TraDIS data: ConNIS (novel), Binomial (TSAS 2.0), Geometric, Exp. vs. Gamma (Bio-TraDIS), InsDens (Bayesian MCMC), Tn5Gaps (TRANSIT). Performance metric = MCC. Results: a synthetic simulation study (160 param combos), three real-world strains, a semi-synthetic study, and a "gene-labeling instability" tuning criterion. Headline numbers are the MCC values in Table 1 ("Tuning performance of the gene labeling instability criterion") and Fig 4 (real-world). Each Table 1 cell has a selected/instability value and a best/optimal value (max MCC over the tuning grid).
In scope (attempted) — real-world, fast analytic methods
For each real-world strain the repo ships the raw insertion data + reference
essential genes, and realworld_<strain>.R recomputes all-method performance
into performance/Performance_<strain>_realWorld.RDS. Table 1 reads from these
with the filter read_count_threshold == {0|1} & trimming_start == 0.05.
We recompute, from the bundled raw data, the optimal (max) MCC of the four analytic, fast methods — ConNIS, Binomial, Geometric, Exp. vs. Gamma — for the exact filter the paper uses, and compare to Table 1's "optimal" column:
| strain | filter (rc, trim_start) | reported optimal MCC: ConNIS / Binomial / ExpVsGamma / Geometric |
|---|---|---|
| BW25113 | rc=0, 0.05 | 0.64 / 0.58 / 0.49 / 0.52 |
| MG1655 | rc=0, 0.05 | 0.79 / 0.79 / 0.79 / 0.74 |
| 14028S | rc=1, 0.05 | 0.57 / 0.53 / 0.41 / 0.48 |
ConNIS is the paper's own method and headline; the other three analytic methods
are reproduced as in-context competitors. Pipeline = the authors' R scripts
(functions.R method implementations + realworld_<strain>.R driver), run
unmodified except for (a) restricting the parameter loops to the single
(read_count_threshold, trimming) cell that Table 1 actually uses, and (b)
dropping the two non-analytic methods (see below).
Out of scope / deliberately skipped (the hard ~20%)
- InsDens (Bayesian MCMC via the
insdensCRAN package) and Tn5Gaps (TRANSIT port): slow, depend on an external MCMC package, and are not ConNIS. Dropped to keep compute bounded; documented, not reproduced. - Synthetic simulation study (160 param combos, 64-core workstation in the original) and semi-synthetic study: heavy compute, needs Zenodo data. Not attempted (80/20).
- Instability-criterion (selected) MCC: requires the
stabilities_*.Rresampling pipeline. We target the optimal MCC column (max over tuning), which for ConNIS equals or closely brackets the reported instability value (BW25113 0.64=0.64; 14028S 0.57=0.57; MG1655 optimal 0.79). Not the selected column itself. - Wet-lab / manual content: none (paper is fully computational).
Reproduction host
All compute on «our HPC» («infra») via SLURM; R env built with conda on the compute node («infra» prefix). «host» holds results/pointers only.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
For the measured E. coli BW25113 real-world claims, this is a clean 1:1 reproduction: all four analytic methods (incl. the headline ConNIS 0.641 vs reported 0.64) match Table 1 to rounding, and an independent recompute from the raw bundled insertion counts also equals the authors' shipped RDS to 3 decimals — no fabrication concern, data fully derivable. The only limitation is partial coverage (MG1655, 14028S, and the synthetic/semi-synthetic studies were still computing at the finalize cutoff), which is a gap on our side, not a defect of the paper. Net: green where measured, with an honest coverage caveat.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.