Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A four-gene signature from blood to exclude bacterial etiology of lower respiratory tract infection in adults.

Nat Commun · 2025
L1 94/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
94/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 87% of all assessed papers rank 133 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and reproduced ~1:1 for the paper's headline DISCOVERY result. The repo (ambaran3/LRTI_CV @ d9cba23) ships the discovery RNA-seq data as a SummarizedExperiment and the pipeline is deterministic (Leave-One-Batch-Out folds; no RNG), so a faithful re-run of the authors' relaxed-LASSO + hard-threshold logistic model exactly reproduced: the 4-gene signature {ITGB4, ITGA7, IFI27, FAM20A}; nested CV-AUC 0.90 (0.8999 stratified); cohort composition (224 any-bacterial vs 280 nonbacterial, N=504); and the -0.886 risk-score threshold giving 89.7% sensitivity, 71.4% specificity, 91.2% NPV@40% prevalence. All seven discovery claims grade 'exact'. NOT attempted: the 8 external validation cohorts (Table 2, AUC 0.73-0.98), including the brief-named GSE211567 (0.90/0.86), because they require external GEO data plus author-curated derivatives (a gene-name-annotated expression file and a curated outcome-metadata CSV) that are not shipped in the repo; phs001248 is dbGaP-restricted. Reconstructing those would be approximate and risk a misleading mismatch, so they were deferred rather than faked. No fabrication concerns: every reproduced value is directly derivable from shipped data+code and matches to rounding.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 94
    assessed: 2026-06-14 ⛓ ded761e15669
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a parsimonious blood-based host gene expression signature accurately discriminate bacterial (including mixed viral-bacterial) from nonbacterial etiology in adults hospitalized with acute respiratory infection, to help exclude bacterial cause and curb unnecessary antibiotic use?

Core claims
  • A 4-gene whole-blood signature (ITGB4, ITGA7, IFI27, FAM20A) discriminates any bacterial from nonbacterial ARI with cross-validated AUC = 0.90 finding
  • The 4-gene signature validates across five independent adult RNAseq cohorts (AUC 0.89-0.98), two adult microarray cohorts (AUC 0.73-0.90), and one pediatric pneumonia RNAseq cohort (AUC 0.74) finding
  • A risk-score threshold tuned to 90% sensitivity yields 71% specificity and 91% negative predictive value for detecting bacterial infection finding
  • A hard-thresholded, mostly relaxed, LASSO-constrained logistic regression with nested leave-one-batch-out cross-validation selects the parsimonious gene set method
  • Mixed viral-bacterial infection is categorized with bacterial infection because antibiotic therapy is warranted method
  • 5401 genes were differentially expressed (FDR<0.05) between bacterial and nonbacterial groups, with 4584 (85%) consistent under leave-one-batch-out analysis finding
  • ITGA7, ITGB4 and FAM20A are higher in bacterial infection while interferon-related IFI27 is higher in nonbacterial infection mechanism
  • The signature represents a clinically usable diagnostic tool/resource to exclude bacterial etiology of LRTI resource
Experimental setups
Assay System Perturbation Readout Platform
bulk whole blood RNA sequencing hospitalized adults with ARI (504 cases: 280 viral, 129 bacterial, 95 mixed viral-bacterial; plus 63 non-infected controls) none (observational infection etiology groups) gene expression / differential expression and diagnostic signature discrimination
RNA-seq validation adult ARI cohort phs001248.v1.p1 (Rochester, 22 bacterial, 37 viral, 9 mixed) none 4-gene signature AUC for bacterial vs nonbacterial RNA-Seq
RNA-seq validation adult febrile illness cohort GSE211567 (101 bacterial, 123 viral, US/Sri Lanka) none signature AUC RNA-Seq
RNA-seq validation adult ARI cohort GSE161731 (24 bacterial, 78 viral) none signature AUC RNA-Seq
RNA-seq validation adult infection cohort GSE163151 (6 bacterial sepsis, 20 influenza/viral) none signature AUC RNA-Seq
RNA-seq validation pediatric pneumonia cohort GSE261482 (98 bacterial, 12 viral) and adult ARI cohort GSE282464 none signature AUC RNA-Seq
microarray validation adult ARI cohorts GSE60244 (154 adults) and GSE63990 (73 bacterial vs 117 viral, FAM20A imputed) none signature AUC Microarray
clinical biomarker comparison primary analysis adult ARI cohort none WBC count and serum procalcitonin (PCT) AUC vs 4-gene signature
Key results
  • 4-gene signature discriminates any bacterial from nonbacterial ARI in primary cohort CV-AUC = 0.90
  • Signature performs equally when 63 non-infected controls added to nonbacterial group (B+VB vs V+NI) AUC = 0.90
  • Threshold of -0.886 (probability >29%) gives 90% sensitivity and 71% specificity in training data sensitivity 90%, specificity 71%
  • NPV across likely prevalences ranges from >91% at 40% prevalence to >95% at <25% prevalence >91% to >95%
  • Validation AUCs across independent RNAseq cohorts AUC 0.89-0.98
  • 5401 DEGs identified between bacterial and nonbacterial; 4584 (85%) consistent in leave-one-batch-out 85% (4584/5401)
  • Of 23 bacterial cases missed at the threshold, 74% were mixed viral-bacterial and 74% non-pneumonic ARI 74%
  • IFI27 higher in nonbacterial; ITGA7, ITGB4, FAM20A higher in bacterial infection
Key statistics
  • other AUC = 0.90 (cross-validated AUC for any bacterial vs nonbacterial in primary cohort)
  • count 5401 genes DEG at FDR<0.05 (differential expression bacterial vs nonbacterial)
  • pvalue PCT 6.9 ± 18.6 vs 0.13 ± 0.15, P<0.0001 (serum procalcitonin bacterial vs nonbacterial)
  • pvalue WBC 14.7 ± 7.7 vs 8.8 ± 3.8, P<0.0001 (white blood cell count bacterial vs nonbacterial)
  • mean Age 63.1 ± 15.9 vs 59.8 ± 18.4, P=0.03 (mean age bacterial vs nonbacterial)
  • other AUC = 0.98 (p<0.0001) (validation in GSE161731 adult ARI)
  • other AUC = 0.74 (p=0.04) (validation in pediatric pneumonia GSE261482)
  • count 40 ± 11 million reads, 84.2 ± 1.1% mapping rate (RNA-seq sequencing quality metrics)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This observational cohort study enrolled hospitalized adults with acute respiratory infection and assigned definitive microbiologic diagnoses (bacterial, viral, or mixed). Whole-blood RNA sequencing was performed, differential expression was tested with DESeq2 (BH FDR < 0.05), and a hard-thresholded relaxed LASSO-penalized logistic regression was used to select a 4-gene diagnostic signature. Model performance was estimated via nested leave-one-batch-out cross-validation (CV-AUC = 0.90) and externally validated in eight independent cohorts using ROC/AUC analysis, with sensitivity and specificity reported at a fixed risk-score threshold targeting ≥90% sensitivity.

Replicationbiological Sample sizeSample sizes described per microbiologic group (viral n=280, bacterial n=129, mixed n=95, non-infected n=63); no formal power calculation stated GroupsAny bacterial (B+VB, n=224) vs nonbacterial/viral-only (V, n=280); also extended to include non-infected controls (n=63) in secondary analyses Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR applied to DESeq2 differential expression results; no correction applied to Table 1 clinical comparisons (explicitly stated: 'unadjusted for multiple comparisons')
Statistical tests used
Test Applied to n Assumptions
Two-sided Student's t-test Comparison of continuous clinical characteristics between bacterial and nonbacterial groups (Table 1) bacterial n=224, nonbacterial n=280 not stated
Chi-square or Fisher's exact test (two-sided) Comparison of categorical clinical characteristics between bacterial and nonbacterial groups (Table 1) bacterial n=224, nonbacterial n=280 not stated
DESeq2 Wald test with Benjamini-Hochberg FDR correction Differential gene expression: any bacterial (B+VB) vs nonbacterial (V) n=224 bacterial, n=280 nonbacterial not stated
Hard-thresholded relaxed LASSO-constrained logistic regression with nested leave-one-batch-out cross-validation Diagnostic feature selection and model performance estimation (CV-AUC = 0.90) n=504 primary analysis (224 bacterial, 280 nonbacterial) not stated
One-sided Wilcoxon rank-sum test (for AUC significance) AUC p-values in all external validation cohorts (Table 2) varies by cohort (stated per cohort in Table 2) not stated
Approaches that could also have been used
  • Leave-one-batch-out (LOBO) cross-validation was used to estimate model performance, exploiting the multi-batch structure of the sequencing data
    Could also: Standard repeated stratified k-fold cross-validation (e.g., 10-fold, 5 repeats) could also have been used for performance estimation — Repeated k-fold CV produces lower-variance AUC estimates when batch membership is not the primary source of overfitting; LOBO is particularly well suited when batches represent distinct processing events that could inflate apparent performance if not held out, so the choice is directly informed by the study's multi-batch design
  • Table 1 clinical characteristic comparisons were explicitly stated as unadjusted for multiple comparisons across approximately 30 tests
    Could also: A Bonferroni correction, Benjamini-Hochberg FDR, or a single omnibus test (e.g., MANOVA for continuous variables) could also have been applied to the Table 1 comparisons — Applying a multiplicity correction to a descriptive baseline table would reduce the chance that any single p-value below 0.05 is flagged as a signal purely by chance; however, many journals treat Table 1 as a descriptive summary rather than a confirmatory test, which is why unadjusted reporting is common
  • DESeq2 (negative binomial model) was used for differential gene expression analysis
    Could also: limma-voom or edgeR could also have been used for RNA-seq differential expression analysis — Each of these three tools makes somewhat different distributional assumptions and uses different normalization strategies; comparing results across methods can increase confidence in the robustness of the identified gene list, and limma-voom in particular is sometimes preferred for large, well-powered studies because its linear modeling framework extends readily to complex designs
  • Elastic-net/LASSO penalization was used for feature selection within logistic regression to arrive at a 4-gene signature
    Could also: Elastic net regularization (combining L1 and L2 penalties) or sparse group LASSO could also have been used for gene selection — The elastic net can handle correlated predictors (common among co-regulated genes) more stably than pure L1 LASSO, potentially yielding a more stable gene set across replications; the relaxed LASSO already partially addresses this, so the practical difference may be small here
  • AUC was used as the primary performance metric, and sensitivity/specificity were reported at a single pre-specified threshold targeting ≥90% sensitivity
    Could also: Decision curve analysis (DCA) could also have been used alongside AUC to quantify the net clinical benefit of the signature across a range of decision thresholds — AUC summarizes discrimination across all thresholds but does not directly translate to clinical utility; DCA expresses benefit in terms of net true positives relative to treat-all or treat-none strategies, which may be more directly interpretable for antibiotic stewardship decision-making
  • Point estimates for AUC, sensitivity, and specificity were reported in validation cohorts without confidence intervals
    Could also: Bootstrap or DeLong 95% confidence intervals for AUC, and exact (Clopper-Pearson) or Wilson CIs for sensitivity and specificity, could also have been reported — Several validation cohorts are small (e.g., n=26 in GSE163151, n=110 in the pediatric cohort), where sampling variability is substantial; CIs would convey the precision of the AUC estimate and help readers judge whether differences across cohorts reflect real heterogeneity or sampling uncertainty
Software: DESeq2 · LASSO logistic regression (package not named; R/glmnet implied)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41285807

Paper: Falsey et al. (2025) A four-gene signature from blood to exclude bacterial etiology of lower respiratory tract infection in adults. Nat Commun. DOI 10.1038/s41467-025-65361-3 · PMCID PMC12644739 Code: https://github.com/ambaran3/LRTI_CV (commit d9cba23dde7ebed81172267877dbd14141b1ea58, main, pushed 2025-09-23; no LICENSE) Brief-named data: GEO GSE211567 (one of the validation cohorts — NOT the discovery data)

Pipeline used

Leave-One-Batch-Out nested cross-validation tuning a hard-thresholded, mostly relaxed, LASSO-constrained logistic regression (R glmnet::cv.glmnet, relax=TRUE, gamma=seq(0,0.5,0.05), standardize=FALSE, type.measure="deviance", foldid = LibPrepBatch), with a post-hoc hard threshold setting |β|<0.1 → 0. Discovery input is whole-blood RNA-seq from 504 hospitalized adults (B=129, VB=95, V=280), shipped in the repo as a SummarizedExperiment (Definite ARI analysis data/DefiniteARI_CPMnorm_log2_SE.rda).

IN SCOPE (pipeline-derived, reproduced)

The discovery model is fully self-contained in the repo and deterministic (folds fixed by sequencing batch — no RNG, no set.seed needed):

id result paper location script
C1 4-gene signature = {ITGB4, ITGA7, IFI27, FAM20A} Abstract; Results "Fig 3 inset"; Discussion prog_CV_relaxedLASSO_1SE_reduced_gamma_grid_HT.R
C2 Nested CV-AUC = 0.90 (any-bacterial vs nonbacterial) Abstract; Results; Fig 3A prog_nestedCV_relaxedLASSO_1SE_reduced_gamma_grid_HT.R
C3 Discovery cohort composition: V=280, B=129, VB=95 (any-bact 224 vs nonbact 280) Abstract; Table 1 shipped .rda colData
C4 Risk-score threshold −0.886 (p>29%) → ~90% sensitivity, ~71% specificity Results; Methods prog_CV_...HT.R + prog_NPV_naivebased.R

STRETCH (20%, attempted, uncertain — external data + author-curated derivatives)

id result note
C5 GSE211567 validation AUC = 0.90 (no-HC), 0.86 (+ill controls) — Table 2 needs GEO download of GSE211567_normData_discovery_2021MAR24.txt.gz + a curated metadata_cc.csv and a gene-name-annotated _GeneName.txt derivative that are NOT shipped in the repo; reconstruction is approximate

OUT OF SCOPE (not attempted)

  • Other 7 validation cohorts (phs001248 dbGaP=restricted; GSE161731/163151/261482/282464/60244/63990) — each needs separate external data + its own curated derivatives. (Table 2 AUCs 0.73–0.98.)
  • Wet-lab / clinical adjudication, microbiology, WBC/PCT comparisons (Supplementary).
  • Figure cosmetics. Note: the discovery model script's plotting tail (prog_CV_...HT.R) contains a broken/interleaved ggplot block — computation (model fit, coef export, naive AUC) precedes it and is unaffected; our driver replicates that computation faithfully and skips the broken plotting code.

Reproduction approach

Faithful R driver repro_primary.R replicating the authors' exact cv.glmnet calls + hard-threshold + AUC computation on the shipped .rda, run on «our HPC» («infra»), output compared to the printed values. Determinism expected → near-exact.

Figures / tables: Fig 3Fig 3ATable
C1
Reported
ITGB4, ITGA7, IFI27, FAM20A
Reproduced
ITGA7, IFI27, FAM20A, ITGB4 (same 4 genes)
exact
C2
Reported
CV-AUC = 0.90
Reproduced
0.8999 (stratified nested) / 0.8948 (pooled)
exact
C3
Reported
V=280, B=129, VB=95; any-bact 224 vs nonbact 280; N=504
Reproduced
V=280, B=129, VB=95; 224 vs 280; N=504
exact
C4a
Reported
threshold score > -0.886 (prob > 29%)
Reproduced
-0.8861609 (prob 0.292)
exact
C4b
Reported
~90% sensitivity
Reproduced
0.8973
exact
C4c
Reported
71% specificity
Reproduced
0.7143
exact
C4d
Reported
91% NPV @40% prev; >95% @<25%
Reproduced
0.9125 @40%; 0.9543 @25%; 0.9843 @10%
exact
C5
Reported
GSE211567 validation AUC 0.90 / 0.86
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 94/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

The paper's discovery model — its headline result — reproduces essentially 1:1 from the repo's shipped SummarizedExperiment via a deterministic Leave-One-Batch-Out relaxed-LASSO pipeline (no RNG): the 4-gene signature, nested CV-AUC 0.8999→0.90, cohort composition (224 vs 280, N=504), the -0.886 threshold, and 89.7% sens / 71.4% spec / 91.2% NPV@40% all match to rounding. No fabrication concerns — every value is mutually consistent and derivable from shipped data+code. The only shortfall is the external Table-2 validation (GSE211567 0.90/0.86 and 7 other cohorts), which was not attempted because it requires external GEO data plus non-deposited author-curated derivatives (and dbGaP-restricted phs001248); this is a data-availability limit on our side, not an authors' defect, and does not undermine the reproduced core claim.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

153.2 k
tokens (I/O) · 12.6 M incl. cache
16 min
runtime · 0.04 CPU-h
1 GB
peak RAM
1
HPC jobs
hummel
machine