Gene Set Enrichment Analysis Reveals Individual Variability in Host Responses in Tuberculosis Patients.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction, described-well-enough = mostly yes (method clearly a standard per-patient GSA / CERNO on Blood Transcriptional Modules). The authors' repo (github.com/terkaterka/immune-response-to-TB) is DELETED (HTTP 404), so under brief rule P16 we re-implemented the described method (tmod::tmodCERNOtest on built-in LI+DC BTMs) and ran it on the paper's own GEO data on «our HPC» (SLURM 2218103, exit 0). The paper's CENTRAL claims reproduce: individual per-patient IFN-response variability incl. IFN-negative patients (C3, both cohorts), dominant IFN/inflammation/immune-activation module classes (C4), and significantly higher BATF2 in IFN+ vs IFN- (C5, p=0.001). The headline 70%/30% IFN+/- split (C1) is a 7-cohort pooled figure; reproduced per-cohort it is 71.7% in the independent Gambia cohort (within-tol of 70%) and 92.6% in the canonical London cohort (more IFN-skewed) -> graded partial. C2 (267/319 IFN-I vs IFN-II split) NOT attempted: it requires the custom Interferome-derived module definition that lived only in the now-deleted repo. NOTE: a prior room had already filled identical claim values; this room independently re-ran the whole pipeline from scratch (prior «infra» workdir + conda env had been reclaimed) and regenerated the same numbers to 3 decimals, which is strong evidence against fabrication. GEOquery 2.70.0 cannot parse these large series matrices (parseGSEMatrix bug); bypassed with a deterministic manual parser (prep_eset.R). All grades provisional; a human reviewer signs off (AUDIT.md).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-16 ⛓ 53d16b8d1b1a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe authors hypothesized that published whole-blood transcriptomic TB signatures are averaged representations masking multiple distinct individual host-response patterns, and set out to test this by analyzing individual (rather than group-aggregated) transcriptional profiles across independent TB cohorts.
- ★ TB patients show substantial individual variability in the intensity of hallmark IFN responses, as well as in complement system, metabolic, and other pathway responses. finding
- ★ This variability cannot be sufficiently explained by covariates such as gender or age, and defines molecular endotypes reproducible across studies and populations. finding
- ★ In a Cynomolgus macaque TB model, transcriptional signatures of the different molecular endotypes did not depend on time since infection (TB progression). finding
- ★ Patients with IFN-rich molecular endotypes suffered from more severe lung pathology than those with IFN-low endotypes. finding
- ★ Machine learning (Random Forest)-derived gene signatures can classify IFN-rich and IFN-low TB endotypes, and the IFN-low signature classified TB vs. non-TB patients slightly more reliably than the IFN-rich signature. finding
- ★ A novel individual-level gene set analysis (GSA) approach combined with clustering of transcriptomic profiles across seven integrated datasets was developed to identify molecular TB endotypes. method
- ★ Complement system response is strongly correlated with the IFN response and contributes to defining the IFN-rich/IFN-low endotype distinction. mechanism
- Additional endotype-relevant variability was identified in D-arginine/D-ornithine metabolic pathways and in insulin- and calcium-metabolism-related modules, independent of IFN response. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Gene Set Analysis (CERNO test via tmod, using blood transcriptional modules) | human whole blood, meta-dataset (MDS) from 7 integrated TB cohorts | none (untreated TB patients vs. healthy/LTBI/OD controls) | per-individual gene set enrichment p-values on z-score-ranked gene lists | mRNA microarray (various platforms), R package tmod |
| Differential expression analysis | human whole blood, individual study datasets | none | differentially expressed genes between TB and healthy/control groups | R package limma |
| Principal Component Analysis (PCA) and eigenvector analysis | human whole blood MDS (all samples and TB-only subset) | none | fraction of variance explained by TB status, IFN status, study, ethnicity, residence, HIV, OD, array technology | R packages stats, pca3d, tmod |
| Random Forest machine learning classification with 10-fold cross-validation | human whole blood MDS | none | classification of TB IFN+/IFN- vs. non-TB (uninfected, LTBI, OD); AUC via ROC curves | R packages randomForest, caret, pROC |
| Longitudinal whole-blood transcriptomics (GSA of IFN status over time) | Cynomolgus macaque Mtb infection model (GSE84152, 38 macaques) | Mtb infection, multiple time points (pre-infection and days 3-180 p.i.) | IFN I+/IFN I- status per timepoint vs. clinical TB/LTBI diagnosis and lung inflammation severity | mRNA-array |
| GSA-based IFN status correlated with radiographic disease severity | human whole blood, GSE19491 dataset (TB patients and healthy individuals) | none | IFN status (transcriptomic) vs. lung X-ray severity grading (healthy/minimal/moderate/advanced) | — |
| Validation of TB IFN+/IFN- transcriptional signatures | external human whole blood datasets (Cai et al.; Blankley et al.) and sepsis patient datasets | none | classification performance (ROC/AUC) of TB IFN+ and IFN- signatures for TB vs. non-TB and vs. sepsis | — |
| Pathway-level GSA (KEGG and Hallmark/MSigDB gene sets) with PCA and Spearman correlation | human whole blood, active TB patients from MDS | none | identification of pathways/endotypes uncorrelated with IFN modules | CERNO enrichment method |
- – Individual TB patients showed markedly variable enrichment in IFN-response and other modules both within and between cohorts, despite overall trends toward T-cell, IFN, and inflammation enrichment.
- – IFN status variability was not sufficiently explained by sex, diabetes, HIV, smoking status, or age (tested via chi-square/Cramér's V, odds ratio, Mann-Whitney/rank biserial correlation).
- ▲ IFN-rich TB patients had more severe lung pathology on X-ray than IFN-low patients in the GSE19491 cohort.
- ▲ The IFN-low signature (50 top-ranked transcripts) achieved slightly more reliable overall classification of TB vs. non-TB patients than the IFN-rich signature (20 top-ranked transcripts).
- – In Cynomolgus macaques, IFN response peaked between days 20-42 post-infection, but onset and duration of IFN response varied substantially between individual animals.
- – Complement system response, D-arginine/D-ornithine metabolism, and insulin/calcium metabolism modules showed variable enrichment across TB patients, some correlating with IFN response and others representing independent endotype-defining pathways.
- count 61 TB patients, 105 healthy individuals (69 LTBI + 36 non-LTBI), 274 OD patients (GSE19491 cohort composition used for IFN status vs. disease severity analysis)
- count 72 individuals with X-ray diagnosis: 34 healthy, 14 minimal disease, 13 moderate disease, 11 advanced disease (subgroup of GSE19491 with lung X-ray severity grading)
- count 20 top-ranked transcripts (size of the defined TB IFN+ transcriptional signature (Signature Model 1))
- count 50 top-ranked transcripts (size of the defined TB IFN- transcriptional signature (Signature Model 2))
- count 38 macaques (Cynomolgus macaque longitudinal Mtb infection dataset (GSE84152))
- other ~15% false negative rate (IGRA test failure rate among Mtb-infected individuals, cited as background)
- count 10 million new TB cases and 1.4 million deaths in 2019; 1.7 billion estimated infected individuals (global TB burden background statistics)
- count seven datasets used to build MDS; two independent validation datasets; three sepsis datasets (study design/data acquisition criteria)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper performs individual-level gene set analysis (GSA) using the CERNO test on z-score-transformed whole-blood transcriptomic data from seven integrated TB datasets, with IQR-based nonparametric standardization across studies to minimize batch effects. IFN-rich and IFN-low TB endotypes were identified by per-patient enrichment profiles against Blood Transcriptional Modules (BTMs) and KEGG/Hallmark collections; Spearman rank correlation was used to identify pathways varying independently of IFN status. Clinical covariate associations were tested with chi-square (discrete variables) and Mann-Whitney U (continuous), and Random Forest classifiers with 10-fold cross-validation were trained to derive and validate TB endotype gene signatures, with performance evaluated by ROC/AUC.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| CERNO test (R/tmod, tmodCERNOtest) | Per-individual gene set enrichment against BTMs and KEGG/Hallmark collections across all patients in the MDS and validation sets | Not stated for total MDS; each constituent dataset required ≥8 TB and ≥8 healthy samples | not stated |
| Chi-square test with Cramér's V effect size | Association of discrete clinical covariates (sex, diabetes, HIV, smoking status) with IFN status | not stated | not stated |
| Mann-Whitney U test with rank biserial correlation | Association of age (continuous variable) with IFN status | not stated | not stated |
| Odds ratio with 95% CI (test H0: OR=1) | Association of discrete clinical covariates with IFN status | not stated | not stated |
| Spearman rank correlation | Correlation between per-patient pathway AUC values (KEGG/Hallmark) and IFN module AUC values, to identify pathways varying independently of IFN status | not stated | not stated |
| Chi-square test for independence; Random Forest with 10-fold cross-validation (R/randomForest, R/caret); ROC/AUC (R/pROC) | Chi-square: co-enrichment of gene sets with IFN gene sets; RF: classification of IFN-rich vs non-TB and IFN-low vs non-TB; ROC: signature performance on training MDS, test MDS, and two external validation datasets | 80% of MDS for training, 20% for test; macaque dataset: 38 animals; GSE19491 X-ray subset: 72 individuals (n=34 healthy, n=14 minimal, n=13 moderate, n=11 advanced disease) | not stated |
-
Individual-level gene set scoring used the CERNO test on z-score–ranked gene lists, with z-scores calculated relative to healthy individuals within each cohort↳ Could also: Single-sample GSEA methods such as ssGSEA or GSVA could also produce per-sample pathway enrichment scores without requiring a reference healthy population — ssGSEA and GSVA generate continuous enrichment scores directly from each sample's expression profile; they are widely benchmarked for individual-level pathway quantification and would allow downstream clustering or regression without the assumption that a matched healthy reference is available in every dataset
-
Batch effects across seven datasets were harmonized using IQR-based nonparametric standardization per gene↳ Could also: Empirical Bayes batch correction (ComBat) or surrogate variable analysis (SVA) could also be applied to multi-cohort microarray data — ComBat and SVA explicitly model study-level variance components and are frequently used as benchmarks in multi-cohort transcriptomic integration, providing a basis for comparing how much residual batch structure remains after correction
-
Multiple separate chi-square tests were used to assess independence between co-enrichment of individual pathways and the IFN gene set↳ Could also: A single logistic regression or generalized linear model with pathway co-enrichment as predictors, followed by one multiplicity correction step, could also address these associations jointly — A joint model accounts for correlations among pathway enrichment scores and provides a unified multiplicity framework; separate tests each require their own correction, and the correction method and family scope applied to these chi-square tests are not stated in the excerpt
-
Random Forest model performance was summarized as a point-estimate AUC from 10-fold cross-validation↳ Could also: Reporting 95% confidence intervals for AUC (e.g., via bootstrap or the pROC CI function) would also be standard when comparing two signatures — A confidence interval around AUC communicates estimation uncertainty and supports formal comparison of whether the IFN-rich and IFN-low signatures differ significantly in discriminative performance
-
GSA p-values for individual patients were depicted by color intensity in heatmap figures rather than reported as exact numerical values↳ Could also: Reporting exact enrichment p-values or the underlying CERNO AUC enrichment statistics in supplementary tables would also be standard — Exact values allow readers and future meta-analysts to apply their own thresholds, assess effect magnitudes quantitatively, and reproduce or compare findings across studies
-
The X-ray disease-severity analysis compared IFN status groups descriptively across four ordinal severity categories (healthy, minimal, moderate, advanced)↳ Could also: An ordinal logistic regression or Jonckheere-Terpstra trend test could also formally test for a monotone association between IFN status and ordered lung pathology severity — Ordinal methods use the natural ordering of severity categories and provide a single, interpretable test statistic for trend, whereas cell-by-cell comparisons across four groups require additional multiplicity correction
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.