Gene signature discovery and systematic validation across diverse clinical cohorts for TB prognosis and response to treatment.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> 1:1 reproduced (partial scope). Authors' own repo (wenhan-yu/tb-common-gene-signature @ 0b59e6c) deposits the per-sample model prediction scores for the pooled progression cohorts (6 signatures x 1183 samples) and diagnostic cohorts (5490 samples). I independently recomputed the deterministic scores->metric stage in clean-room pure-python (Mann-Whitney AUC, the authors' Hanley-McNeil CI, Youden top-left-corner cutoff) and compared to the shipped result CSVs and the paper. RESULT: every numeric claim in the abstract reproduced EXACTLY -- progression 2.5-year AUROC 0.85, ATB-vs-viral AUROC 0.93 (0.91-0.94), WHO sensitivity 74.2% / specificity 78.3% at the Youden cutoff -- and the entire combined-prognosis AUROC table (Table 2 full + reduced models plus RISK6/BATF2/Suliman4/Sweeney3 comparators) matched 72/72 cells including the 95% CIs. No fabrication detected at this stage: the reported numbers are exactly what the deposited scores yield. NOT attempted (documented 20%): regenerating the prediction scores from the deposited 34/84 MB RandomForest pickles, because the model input feature matrices (per-sample gene-pair expression ratios for the ~37 GEO/ArrayExpress datasets) are not deposited -- that needs the full GEO download + R microarray/RNA-seq QC+normalisation+feature-engineering pipeline; high effort/risk, deferred per 80/20. Also not attempted: the upstream 45-gene network-discovery step and treatment-monitoring AUCs (same un-deposited-input dependency). No SLURM job was used because the reproduced stage is metric recomputation over <=6x1183 rows (~1 s, stdlib only) -- «host» control-plane class, not heavy compute; the only «our HPC»-worthy step is the deferred model->scores inference. Caveat for the human reviewer: grades certify deposited-scores<->reported-metrics consistency, not the model inference that produced the scores.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-15 ⛓ ef20729a7dca
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a network-based meta-analysis combined with machine-learning modeling leverage heterogeneity across many clinical cohorts to identify a common blood gene signature specific to active tuberculosis and build a generalizable predictive model for TB risk estimation and treatment monitoring?
- ★ A network-based meta-analysis across studies identifies a common 45-gene signature specific to active TB disease that accounts for cohort/population heterogeneity finding
- ★ Optimized random forest regression models (full and reduced 45-gene sets) model the continuum from Mtb infection to disease and treatment response method
- ★ The model robustly predicts incipient-to-active TB risk over a 2.5-year period, approximating WHO target product profile criteria finding
- ★ The model strongly discriminates active TB from viral infection finding
- ★ Model-generated TB scores correlate with treatment response over time and are predictive of treatment outcomes even before treatment initiation finding
- ★ The network-based approach captures both differentially expressed genes and gene covariation reflecting functionally important biological processes (e.g., IFN-γ, IFN-α/β, IL-6 signaling, Toll-like receptor cascades) mechanism
- 71% (32/45) of signature genes are interconnected in a STRING protein-protein association network, with STAT1 forming associations with 17 proteins finding
- An end-to-end gene signature model development scheme and a probabilistic TB risk estimation tool are provided as a resource resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole blood transcriptome meta-analysis / differential gene expression | 27 published human TB cohorts (discovery dataset; ATB vs HC/LTBI/OLD/Tx) | none (observational disease-state comparison) | log fold-change of gene expression between disease conditions | qRT-PCR, microarray, or RNA-seq (mixed across studies) |
| Gene covariation network construction | N×M logFC matrix from M cohorts and N genes | none | node centrality (weighted degree) and edge weights (dot products) | — |
| Random forest regression modeling (ML) | Common 45-gene signature training data from TB cohorts | none | TB score modeling Mtb infection-to-disease continuum | — |
| Model validation (longitudinal TB progression prediction) | 10 independent longitudinal human cohorts | none | AUROC, sensitivity, specificity for incipient-to-active TB risk | — |
| Model specificity testing | 20 viral infection cohorts (influenza, RSV, rhinovirus, etc.) | none | AUROC discriminating ATB from viral infection | — |
| Treatment monitoring analysis | Longitudinal TB treatment cohorts (e.g., GSE89403, GSE67589, GSE157657) | standard TB drug treatment | TB score correlation with treatment response over time | — |
| Gene set enrichment analysis | 45-gene signature | none | enriched immune pathways | — |
| Protein-protein association network analysis | 45 candidate gene proteins | none | interconnection of genes/proteins | STRING database |
- ▲ Model predicts incipient-to-active TB risk over a 2.5-year period AUROC 0.85, 74.2% sensitivity, 78.3% specificity
- ▲ Model discriminates active TB from viral infection AUROC 0.93 (95% CI 0.91–0.94)
- – 45-gene signature identified as common to active TB across cohorts 45 genes
- ▲ Differential expression patterns correlate strongly across disease conditions r = 0.61–0.96, p ≤ 1e-05
- – Majority of signature genes interconnected in STRING PPI network 71% (32/45)
- – STAT1 forms associations with many proteins in the network n=17
- – TB scores correlate with and are predictive of treatment outcomes
- other AUROC 0.85, 74.2% sensitivity, 78.3% specificity (incipient-to-active TB risk prediction over 2.5 years)
- other AUROC 0.93 (95% CI 0.91–0.94) (discriminating active TB from viral infection)
- correlation r = 0.61–0.96, all p-values ≤ 1e-05 (correlation of averaged logFC expression patterns between disease comparisons)
- count 45 genes (common ATB-specific gene signature size)
- count 71% (32/45) (proportion of signature genes interconnected in STRING network)
- count n=17 (number of proteins STAT1 associates with in the network)
- count 57 studies (37 TB and 20 viral infections) (total transcriptome datasets used)
- other top 5% central genes, central in ≥2 of 4 networks, average logFC ≥ 0.5 or ≤ -0.5 (gene signature selection criteria)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper employs a two-stage computational design: first, a custom network-based meta-analysis across 27 discovery transcriptome cohorts builds gene covariation networks from per-cohort differential expression log fold-changes, using permutation-validated edge weights and node centrality to select a 45-gene TB-specific signature; second, two random forest regression models are trained on this signature and validated across 10 independent longitudinal TB cohorts and 20 viral infection cohorts. Discrimination performance is reported as AUROC (with 95% CI for at least one comparison), sensitivity, and specificity; cross-condition consistency of the gene signature is assessed via Pearson correlation of averaged logFC patterns; and pathway relevance is assessed via gene set enrichment analysis.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential gene expression analysis (log fold-change; specific package/method not named in excerpt) | Per-cohort comparisons of ATB vs. HC, LTBI, OLD, and Tx within each of the 27 discovery datasets | Varies by cohort (range approximately 27–537 per cohort; see Table 1) | not stated |
| Permutation test | Validation of the edge weight threshold (≥3) used to retain edges in the gene covariation network | null | not stated |
| Pearson correlation (r) | Comparison of average logFC patterns between pairs of disease-condition contrasts (ATB vs. HC, LTBI, OLD, Tx) for the 45-gene set; r reported as 0.61–0.96, all p ≤ 1e-05 | 45 genes | not stated |
| Gene set enrichment analysis | Pathway enrichment of the 45-gene signature (IFN-γ, IFN-α/β, IL-6, TLR cascades) | 45 genes | not stated |
| Random forest regression | ML model training (full and reduced 45-gene sets) for TB score generation | Pooled multi-cohort discovery dataset; exact training n not stated in excerpt | not stated |
| AUROC / ROC analysis with sensitivity and specificity at a fixed threshold | Validation: TB progression prediction (AUROC 0.85, sensitivity 74.2%, specificity 78.3%); ATB vs. viral infection (AUROC 0.93, 95% CI 0.91–0.94) | null | na |
-
Per-cohort logFC values were aggregated via a custom network covariation approach to identify cross-cohort signal↳ Could also: A formal random-effects meta-analysis (e.g., using R packages metafor or limma with dream/voom) could also synthesize per-cohort effect sizes weighted by their standard errors — Random-effects meta-analysis explicitly estimates and reports between-study heterogeneity (τ²) and yields confidence intervals for pooled effect sizes, giving a standardized quantification of how consistently each gene responds across cohorts and allowing direct comparison with other meta-analyses
-
The top-5% node centrality cutoff (combined with logFC ≥ 0.5) was used to select the 45-gene signature from the network↳ Could also: A permutation-based FDR threshold on centrality scores (e.g., Benjamini-Hochberg q < 0.05 applied to a null distribution of weighted degrees from randomized networks) could also define the selection boundary — A data-driven FDR cutoff provides a principled, reproducible selection criterion whose false-discovery rate is explicitly controlled, which aids replication and helps reviewers interpret how conservative or liberal the threshold is relative to chance
-
Pearson correlation was used to compare average logFC patterns across the four disease-condition contrasts for the 45-gene set↳ Could also: Spearman rank correlation could also be used, as it does not assume linearity or normality of the logFC distribution — With a modest gene set (n = 45) and a few flagged outlier genes (HP, CEACAM1 noted as inconsistent), a rank-based measure is more robust to influential observations without requiring distributional assumptions
-
Random forest regression was the sole ML framework evaluated for TB score generation↳ Could also: Regularized linear regression (LASSO or elastic net) or gradient boosting could also be applied, with cross-validated hyperparameter tuning and formal comparison between methods — LASSO/elastic net yield explicit feature coefficients that simplify clinical interpretation and implementation (e.g., as a weighted sum score); evaluating multiple model classes with a held-out comparison allows the best-generalizing approach to be selected on principled grounds rather than assumed
-
Model discrimination at validation was reported as AUROC and a single fixed sensitivity/specificity operating point benchmarked against WHO targets↳ Could also: Calibration metrics (Brier score, calibration curves, or reliability diagrams) could also be reported alongside AUROC — For a risk-estimation tool intended for clinical screening, calibration—how well predicted probabilities match observed event rates—complements discrimination and is needed to evaluate whether the score functions as an absolute risk estimate rather than a rank-ordering device
-
Longitudinal validation cohort performance was summarized with a single binary AUROC (progressor vs. non-progressor at a fixed endpoint)↳ Could also: Time-dependent ROC analysis (e.g., R package timeROC) or the concordance statistic from a Cox model could also evaluate performance across the full 2.5-year prediction horizon — Standard AUROC treats the outcome as binary at one time point, which can obscure how predictive accuracy changes as the prediction horizon varies; time-dependent metrics explicitly account for the censoring structure inherent in longitudinal TB progression data
-
No multiplicity correction is mentioned for the multiple per-cohort DEG analyses or the six pairwise cross-condition correlations↳ Could also: A Benjamini-Hochberg FDR correction applied across all pairwise correlation tests, or a Bonferroni correction across the DEG analyses, could also be reported — With six correlation tests and many per-cohort DEG comparisons run in parallel, reporting a multiplicity-corrected threshold alongside nominal p-values makes explicit how many findings would survive a family-wise error control and contextualizes the strength of the cross-condition concordance result
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
TB-score predicts incipient-to-active TB progression over 2.5 years with AUROC 0.85 (74.2% sensitivity, 78.3% specificity) across independent longitudinal cohorts.other human blood up 2023×1papers★ This paper is the founder (earliest)
-
TB-score declines over the course of standard TB drug treatment and correlates with treatment response in longitudinal cohorts.other human blood down 2023×1papers★ This paper is the founder (earliest)
-
TB-score discriminates active TB from viral infections (influenza, RSV, rhinovirus) with AUROC 0.93 across 20 independent viral cohorts.other human blood up 2023×1papers★ This paper is the founder (earliest)
-
71% (32/45) of the TB transcriptome signature genes form an interconnected cluster in the STRING protein-protein interaction network, indicating functional coherence.other human blood 2023×1papers★ This paper is the founder (earliest)
-
A 45-gene whole-blood transcriptome signature is identified as commonly dysregulated in active TB across 27 diverse published human cohorts by meta-analysis.other human blood 2023×1papers★ This paper is the founder (earliest)
-
STAT1 is the highest-connectivity hub in the 45-gene TB signature protein-protein interaction network, associating with 17 other network proteins.other human blood 2023×1papers★ This paper is the founder (earliest)
-
TB differential expression patterns are highly correlated across disease conditions and cohorts (r=0.61–0.96, p≤1e-05), supporting a conserved host transcriptomic response to Mtb.other human blood 2023×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37471455
Paper: Vargas R, Abbott L, Bower D, Frahm N, Shaffer M, Yu WH. Gene signature
discovery and systematic validation across diverse clinical cohorts for TB
prognosis and response to treatment. PLoS Comput Biol 2023. PMCID PMC10393163.
Repo (authors' own, P16 N/A): https://github.com/wenhan-yu/tb-common-gene-signature
@ commit 0b59e6c810ac8d0fff640f73cd91d638978ae720 (master, pushed 2022-12-07).
Pipeline (from Methods + repo)
- DE analysis across 27 discovery datasets (R:
fun/R/*,r-df-process.ipynb). - Gene covariation networks → 45-gene common signature (
py-network-building,py-gene-signature). - Feature engineering: pairwise expression ratios (990 features from 45 genes).
- Two-stage feature selection (mutual-info + LASSO) → Full=41 pairs, Reduced=12 pairs.
- ML model selection via 5×5 nested CV (7 algos) → RandomForestRegressor chosen (
py-predictive-model). - Validation on longitudinal progression cohorts + diagnostic cohorts + viral-infection cohorts
(
fun/validation.py,py-model-validation-*).
In scope (attempted) — scores → reported-metric stage
The repo deposits the per-sample model prediction scores (Y_pred) for the
pooled progression cohorts (final-ML-model/opt_score_rocauc/Com_progress_pred_*.csv,
6 signatures × 1183 samples) and the diagnostic cohorts
(data/TB_disease_prediction_*.csv, 5490 samples). The reported AUROCs,
sensitivity, specificity and their CIs are the deterministic scores→metric stage.
We independently recompute that stage (clean-room pure-python: Mann-Whitney
AUC, the authors' Hanley-McNeil CI roc_auc_ci, and the Youden/top-left-corner
cutoff utilities.rocauc) and compare to the shipped result CSVs and to the paper.
Targets:
- Table 2 / combined-prognosis figure: full AUROC table, 6 signatures (Full fs2, Reduced fs3, RISK6, BATF2, Suliman4, Sweeney3) × 12 time-interval columns (exclusive + cumulative) — 72 cells incl. CIs.
- Abstract claim 1: incipient→active-TB 2.5-year AUROC 0.85.
- Abstract claim 2: ATB vs viral infection AUROC 0.93 (0.91–0.94).
- Abstract claim 3: WHO sens/spec at Youden cutoff over 30m: 74.2% / 78.3%.
Out of scope / NOT attempted (the documented 20%)
- Regenerating
Y_predfrom the deposited RF model pickles (34 MB fs2 / 84 MB fs3). The model input feature matrices (per-sample gene-pair expression ratios for the ~37 GEO/ArrayExpress datasets) are not deposited; only model outputs (scores, figures) and the trained pickles are. Regenerating scores needs the full GEO download + R microarray/RNA-seq QC/normalisation + probe→gene mapping + pairwise-ratio feature build (r-df-process,r-validate-data-process). High effort, high risk of subtle normalisation mismatch → deferred per the brief's 80/20 rule. - The 45-gene network-discovery step itself (network construction over 27 datasets) — upstream of the deposited signature; not attempted.
- Treatment-response monitoring AUCs (Catalysis/SA cohorts) — same dependency on un-deposited feature matrices; not attempted.
- Wet-lab / manual content: none (paper is fully computational).
Compute note
The reproduced stage is metric recomputation over ≤6×1183 rows — milliseconds,
pure stdlib. No heavy compute and therefore no SLURM job was required («host»
control-plane class, same as screening). The only «our HPC»-worthy step (model→scores
on raw GEO data) is the deferred 20% above. «infra» work dir
«path»
was therefore not populated (no heavy data downloaded).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
What deviates: nothing within the reproduced scope — all five abstract metrics (prog AUROC 0.85, ATB-vs-viral 0.93 [0.91–0.94], sens 0.742, spec 0.783) and the full 72-cell Table 2 (point AUCs + 95% CIs) reproduce bit-for-bit at 2 d.p. from the authors' deposited per-sample scores. Whose side / severity: no authors' defect and no fabrication signal; the only limitation is data availability — the per-sample input feature matrices for the ~37 GEO datasets are not deposited, so the model→scores inference, the 45-gene discovery step, and treatment-monitoring AUCs could not be independently re-run. Net: a genuinely exact but partial reproduction that certifies deposited-scores↔reported-metrics internal consistency, not the predictive model that produced the scores — hence q1 yellow (inputs unavailable) and q8 yellow (solid, explainable scope gap) while q5/q7 stay green.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.