A four-gene signature from blood to exclude bacterial etiology of lower respiratory tract infection in adults.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough and reproduced ~1:1 for the paper's headline DISCOVERY result. The repo (ambaran3/LRTI_CV @ d9cba23) ships the discovery RNA-seq data as a SummarizedExperiment and the pipeline is deterministic (Leave-One-Batch-Out folds; no RNG), so a faithful re-run of the authors' relaxed-LASSO + hard-threshold logistic model exactly reproduced: the 4-gene signature {ITGB4, ITGA7, IFI27, FAM20A}; nested CV-AUC 0.90 (0.8999 stratified); cohort composition (224 any-bacterial vs 280 nonbacterial, N=504); and the -0.886 risk-score threshold giving 89.7% sensitivity, 71.4% specificity, 91.2% NPV@40% prevalence. All seven discovery claims grade 'exact'. NOT attempted: the 8 external validation cohorts (Table 2, AUC 0.73-0.98), including the brief-named GSE211567 (0.90/0.86), because they require external GEO data plus author-curated derivatives (a gene-name-annotated expression file and a curated outcome-metadata CSV) that are not shipped in the repo; phs001248 is dbGaP-restricted. Reconstructing those would be approximate and risk a misleading mismatch, so they were deferred rather than faked. No fabrication concerns: every reproduced value is directly derivable from shipped data+code and matches to rounding.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 94assessed: 2026-06-14 ⛓ ded761e15669
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a parsimonious blood-based host gene expression signature accurately discriminate bacterial (including mixed viral-bacterial) from nonbacterial etiology in adults hospitalized with acute respiratory infection, to help exclude bacterial cause and curb unnecessary antibiotic use?
- ★ A 4-gene whole-blood signature (ITGB4, ITGA7, IFI27, FAM20A) discriminates any bacterial from nonbacterial ARI with cross-validated AUC = 0.90 finding
- ★ The 4-gene signature validates across five independent adult RNAseq cohorts (AUC 0.89-0.98), two adult microarray cohorts (AUC 0.73-0.90), and one pediatric pneumonia RNAseq cohort (AUC 0.74) finding
- ★ A risk-score threshold tuned to 90% sensitivity yields 71% specificity and 91% negative predictive value for detecting bacterial infection finding
- ★ A hard-thresholded, mostly relaxed, LASSO-constrained logistic regression with nested leave-one-batch-out cross-validation selects the parsimonious gene set method
- Mixed viral-bacterial infection is categorized with bacterial infection because antibiotic therapy is warranted method
- 5401 genes were differentially expressed (FDR<0.05) between bacterial and nonbacterial groups, with 4584 (85%) consistent under leave-one-batch-out analysis finding
- ITGA7, ITGB4 and FAM20A are higher in bacterial infection while interferon-related IFI27 is higher in nonbacterial infection mechanism
- The signature represents a clinically usable diagnostic tool/resource to exclude bacterial etiology of LRTI resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk whole blood RNA sequencing | hospitalized adults with ARI (504 cases: 280 viral, 129 bacterial, 95 mixed viral-bacterial; plus 63 non-infected controls) | none (observational infection etiology groups) | gene expression / differential expression and diagnostic signature discrimination | — |
| RNA-seq validation | adult ARI cohort phs001248.v1.p1 (Rochester, 22 bacterial, 37 viral, 9 mixed) | none | 4-gene signature AUC for bacterial vs nonbacterial | RNA-Seq |
| RNA-seq validation | adult febrile illness cohort GSE211567 (101 bacterial, 123 viral, US/Sri Lanka) | none | signature AUC | RNA-Seq |
| RNA-seq validation | adult ARI cohort GSE161731 (24 bacterial, 78 viral) | none | signature AUC | RNA-Seq |
| RNA-seq validation | adult infection cohort GSE163151 (6 bacterial sepsis, 20 influenza/viral) | none | signature AUC | RNA-Seq |
| RNA-seq validation | pediatric pneumonia cohort GSE261482 (98 bacterial, 12 viral) and adult ARI cohort GSE282464 | none | signature AUC | RNA-Seq |
| microarray validation | adult ARI cohorts GSE60244 (154 adults) and GSE63990 (73 bacterial vs 117 viral, FAM20A imputed) | none | signature AUC | Microarray |
| clinical biomarker comparison | primary analysis adult ARI cohort | none | WBC count and serum procalcitonin (PCT) AUC vs 4-gene signature | — |
- – 4-gene signature discriminates any bacterial from nonbacterial ARI in primary cohort CV-AUC = 0.90
- – Signature performs equally when 63 non-infected controls added to nonbacterial group (B+VB vs V+NI) AUC = 0.90
- – Threshold of -0.886 (probability >29%) gives 90% sensitivity and 71% specificity in training data sensitivity 90%, specificity 71%
- – NPV across likely prevalences ranges from >91% at 40% prevalence to >95% at <25% prevalence >91% to >95%
- – Validation AUCs across independent RNAseq cohorts AUC 0.89-0.98
- – 5401 DEGs identified between bacterial and nonbacterial; 4584 (85%) consistent in leave-one-batch-out 85% (4584/5401)
- – Of 23 bacterial cases missed at the threshold, 74% were mixed viral-bacterial and 74% non-pneumonic ARI 74%
- – IFI27 higher in nonbacterial; ITGA7, ITGB4, FAM20A higher in bacterial infection
- other AUC = 0.90 (cross-validated AUC for any bacterial vs nonbacterial in primary cohort)
- count 5401 genes DEG at FDR<0.05 (differential expression bacterial vs nonbacterial)
- pvalue PCT 6.9 ± 18.6 vs 0.13 ± 0.15, P<0.0001 (serum procalcitonin bacterial vs nonbacterial)
- pvalue WBC 14.7 ± 7.7 vs 8.8 ± 3.8, P<0.0001 (white blood cell count bacterial vs nonbacterial)
- mean Age 63.1 ± 15.9 vs 59.8 ± 18.4, P=0.03 (mean age bacterial vs nonbacterial)
- other AUC = 0.98 (p<0.0001) (validation in GSE161731 adult ARI)
- other AUC = 0.74 (p=0.04) (validation in pediatric pneumonia GSE261482)
- count 40 ± 11 million reads, 84.2 ± 1.1% mapping rate (RNA-seq sequencing quality metrics)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This observational cohort study enrolled hospitalized adults with acute respiratory infection and assigned definitive microbiologic diagnoses (bacterial, viral, or mixed). Whole-blood RNA sequencing was performed, differential expression was tested with DESeq2 (BH FDR < 0.05), and a hard-thresholded relaxed LASSO-penalized logistic regression was used to select a 4-gene diagnostic signature. Model performance was estimated via nested leave-one-batch-out cross-validation (CV-AUC = 0.90) and externally validated in eight independent cohorts using ROC/AUC analysis, with sensitivity and specificity reported at a fixed risk-score threshold targeting ≥90% sensitivity.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Two-sided Student's t-test | Comparison of continuous clinical characteristics between bacterial and nonbacterial groups (Table 1) | bacterial n=224, nonbacterial n=280 | not stated |
| Chi-square or Fisher's exact test (two-sided) | Comparison of categorical clinical characteristics between bacterial and nonbacterial groups (Table 1) | bacterial n=224, nonbacterial n=280 | not stated |
| DESeq2 Wald test with Benjamini-Hochberg FDR correction | Differential gene expression: any bacterial (B+VB) vs nonbacterial (V) | n=224 bacterial, n=280 nonbacterial | not stated |
| Hard-thresholded relaxed LASSO-constrained logistic regression with nested leave-one-batch-out cross-validation | Diagnostic feature selection and model performance estimation (CV-AUC = 0.90) | n=504 primary analysis (224 bacterial, 280 nonbacterial) | not stated |
| One-sided Wilcoxon rank-sum test (for AUC significance) | AUC p-values in all external validation cohorts (Table 2) | varies by cohort (stated per cohort in Table 2) | not stated |
-
Leave-one-batch-out (LOBO) cross-validation was used to estimate model performance, exploiting the multi-batch structure of the sequencing data↳ Could also: Standard repeated stratified k-fold cross-validation (e.g., 10-fold, 5 repeats) could also have been used for performance estimation — Repeated k-fold CV produces lower-variance AUC estimates when batch membership is not the primary source of overfitting; LOBO is particularly well suited when batches represent distinct processing events that could inflate apparent performance if not held out, so the choice is directly informed by the study's multi-batch design
-
Table 1 clinical characteristic comparisons were explicitly stated as unadjusted for multiple comparisons across approximately 30 tests↳ Could also: A Bonferroni correction, Benjamini-Hochberg FDR, or a single omnibus test (e.g., MANOVA for continuous variables) could also have been applied to the Table 1 comparisons — Applying a multiplicity correction to a descriptive baseline table would reduce the chance that any single p-value below 0.05 is flagged as a signal purely by chance; however, many journals treat Table 1 as a descriptive summary rather than a confirmatory test, which is why unadjusted reporting is common
-
DESeq2 (negative binomial model) was used for differential gene expression analysis↳ Could also: limma-voom or edgeR could also have been used for RNA-seq differential expression analysis — Each of these three tools makes somewhat different distributional assumptions and uses different normalization strategies; comparing results across methods can increase confidence in the robustness of the identified gene list, and limma-voom in particular is sometimes preferred for large, well-powered studies because its linear modeling framework extends readily to complex designs
-
Elastic-net/LASSO penalization was used for feature selection within logistic regression to arrive at a 4-gene signature↳ Could also: Elastic net regularization (combining L1 and L2 penalties) or sparse group LASSO could also have been used for gene selection — The elastic net can handle correlated predictors (common among co-regulated genes) more stably than pure L1 LASSO, potentially yielding a more stable gene set across replications; the relaxed LASSO already partially addresses this, so the practical difference may be small here
-
AUC was used as the primary performance metric, and sensitivity/specificity were reported at a single pre-specified threshold targeting ≥90% sensitivity↳ Could also: Decision curve analysis (DCA) could also have been used alongside AUC to quantify the net clinical benefit of the signature across a range of decision thresholds — AUC summarizes discrimination across all thresholds but does not directly translate to clinical utility; DCA expresses benefit in terms of net true positives relative to treat-all or treat-none strategies, which may be more directly interpretable for antibiotic stewardship decision-making
-
Point estimates for AUC, sensitivity, and specificity were reported in validation cohorts without confidence intervals↳ Could also: Bootstrap or DeLong 95% confidence intervals for AUC, and exact (Clopper-Pearson) or Wilson CIs for sensitivity and specificity, could also have been reported — Several validation cohorts are small (e.g., n=26 in GSE163151, n=110 in the pediatric cohort), where sampling variability is substantial; CIs would convey the precision of the AUC estimate and help readers judge whether differences across cohorts reflect real heterogeneity or sampling uncertainty
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41285807
Paper: Falsey et al. (2025) A four-gene signature from blood to exclude bacterial
etiology of lower respiratory tract infection in adults. Nat Commun.
DOI 10.1038/s41467-025-65361-3 · PMCID PMC12644739
Code: https://github.com/ambaran3/LRTI_CV (commit d9cba23dde7ebed81172267877dbd14141b1ea58, main, pushed 2025-09-23; no LICENSE)
Brief-named data: GEO GSE211567 (one of the validation cohorts — NOT the discovery data)
Pipeline used
Leave-One-Batch-Out nested cross-validation tuning a hard-thresholded, mostly
relaxed, LASSO-constrained logistic regression (R glmnet::cv.glmnet,
relax=TRUE, gamma=seq(0,0.5,0.05), standardize=FALSE, type.measure="deviance",
foldid = LibPrepBatch), with a post-hoc hard threshold setting |β|<0.1 → 0.
Discovery input is whole-blood RNA-seq from 504 hospitalized adults
(B=129, VB=95, V=280), shipped in the repo as a SummarizedExperiment
(Definite ARI analysis data/DefiniteARI_CPMnorm_log2_SE.rda).
IN SCOPE (pipeline-derived, reproduced)
The discovery model is fully self-contained in the repo and deterministic
(folds fixed by sequencing batch — no RNG, no set.seed needed):
| id | result | paper location | script |
|---|---|---|---|
| C1 | 4-gene signature = {ITGB4, ITGA7, IFI27, FAM20A} | Abstract; Results "Fig 3 inset"; Discussion | prog_CV_relaxedLASSO_1SE_reduced_gamma_grid_HT.R |
| C2 | Nested CV-AUC = 0.90 (any-bacterial vs nonbacterial) | Abstract; Results; Fig 3A | prog_nestedCV_relaxedLASSO_1SE_reduced_gamma_grid_HT.R |
| C3 | Discovery cohort composition: V=280, B=129, VB=95 (any-bact 224 vs nonbact 280) | Abstract; Table 1 | shipped .rda colData |
| C4 | Risk-score threshold −0.886 (p>29%) → ~90% sensitivity, ~71% specificity | Results; Methods | prog_CV_...HT.R + prog_NPV_naivebased.R |
STRETCH (20%, attempted, uncertain — external data + author-curated derivatives)
| id | result | note |
|---|---|---|
| C5 | GSE211567 validation AUC = 0.90 (no-HC), 0.86 (+ill controls) — Table 2 | needs GEO download of GSE211567_normData_discovery_2021MAR24.txt.gz + a curated metadata_cc.csv and a gene-name-annotated _GeneName.txt derivative that are NOT shipped in the repo; reconstruction is approximate |
OUT OF SCOPE (not attempted)
- Other 7 validation cohorts (phs001248 dbGaP=restricted; GSE161731/163151/261482/282464/60244/63990) — each needs separate external data + its own curated derivatives. (Table 2 AUCs 0.73–0.98.)
- Wet-lab / clinical adjudication, microbiology, WBC/PCT comparisons (Supplementary).
- Figure cosmetics. Note: the discovery model script's plotting tail
(
prog_CV_...HT.R) contains a broken/interleaved ggplot block — computation (model fit, coef export, naive AUC) precedes it and is unaffected; our driver replicates that computation faithfully and skips the broken plotting code.
Reproduction approach
Faithful R driver repro_primary.R replicating the authors' exact cv.glmnet
calls + hard-threshold + AUC computation on the shipped .rda, run on «our HPC»
(«infra»), output compared to the printed values. Determinism expected → near-exact.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's discovery model — its headline result — reproduces essentially 1:1 from the repo's shipped SummarizedExperiment via a deterministic Leave-One-Batch-Out relaxed-LASSO pipeline (no RNG): the 4-gene signature, nested CV-AUC 0.8999→0.90, cohort composition (224 vs 280, N=504), the -0.886 threshold, and 89.7% sens / 71.4% spec / 91.2% NPV@40% all match to rounding. No fabrication concerns — every value is mutually consistent and derivable from shipped data+code. The only shortfall is the external Table-2 validation (GSE211567 0.90/0.86 and 7 other cohorts), which was not attempted because it requires external GEO data plus non-deposited author-curated derivatives (and dbGaP-restricted phs001248); this is a data-availability limit on our side, not an authors' defect, and does not undermine the reproduced core claim.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.