Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Gene signature discovery and systematic validation across diverse clinical cohorts for TB prognosis and response to treatment.

PLoS Comput Biol · 2023
L1 100/100 PQI 99
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score 0
✓ What held up
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> 1:1 reproduced (partial scope). Authors' own repo (wenhan-yu/tb-common-gene-signature @ 0b59e6c) deposits the per-sample model prediction scores for the pooled progression cohorts (6 signatures x 1183 samples) and diagnostic cohorts (5490 samples). I independently recomputed the deterministic scores->metric stage in clean-room pure-python (Mann-Whitney AUC, the authors' Hanley-McNeil CI, Youden top-left-corner cutoff) and compared to the shipped result CSVs and the paper. RESULT: every numeric claim in the abstract reproduced EXACTLY -- progression 2.5-year AUROC 0.85, ATB-vs-viral AUROC 0.93 (0.91-0.94), WHO sensitivity 74.2% / specificity 78.3% at the Youden cutoff -- and the entire combined-prognosis AUROC table (Table 2 full + reduced models plus RISK6/BATF2/Suliman4/Sweeney3 comparators) matched 72/72 cells including the 95% CIs. No fabrication detected at this stage: the reported numbers are exactly what the deposited scores yield. NOT attempted (documented 20%): regenerating the prediction scores from the deposited 34/84 MB RandomForest pickles, because the model input feature matrices (per-sample gene-pair expression ratios for the ~37 GEO/ArrayExpress datasets) are not deposited -- that needs the full GEO download + R microarray/RNA-seq QC+normalisation+feature-engineering pipeline; high effort/risk, deferred per 80/20. Also not attempted: the upstream 45-gene network-discovery step and treatment-monitoring AUCs (same un-deposited-input dependency). No SLURM job was used because the reproduced stage is metric recomputation over <=6x1183 rows (~1 s, stdlib only) -- «host» control-plane class, not heavy compute; the only «our HPC»-worthy step is the deferred model->scores inference. Caveat for the human reviewer: grades certify deposited-scores<->reported-metrics consistency, not the model inference that produced the scores.

💻 Code ↗ 🗄 Data: GSE19439

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-15 ⛓ ef20729a7dca
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a network-based meta-analysis combined with machine-learning modeling leverage heterogeneity across many clinical cohorts to identify a common blood gene signature specific to active tuberculosis and build a generalizable predictive model for TB risk estimation and treatment monitoring?

Core claims
  • A network-based meta-analysis across studies identifies a common 45-gene signature specific to active TB disease that accounts for cohort/population heterogeneity finding
  • Optimized random forest regression models (full and reduced 45-gene sets) model the continuum from Mtb infection to disease and treatment response method
  • The model robustly predicts incipient-to-active TB risk over a 2.5-year period, approximating WHO target product profile criteria finding
  • The model strongly discriminates active TB from viral infection finding
  • Model-generated TB scores correlate with treatment response over time and are predictive of treatment outcomes even before treatment initiation finding
  • The network-based approach captures both differentially expressed genes and gene covariation reflecting functionally important biological processes (e.g., IFN-γ, IFN-α/β, IL-6 signaling, Toll-like receptor cascades) mechanism
  • 71% (32/45) of signature genes are interconnected in a STRING protein-protein association network, with STAT1 forming associations with 17 proteins finding
  • An end-to-end gene signature model development scheme and a probabilistic TB risk estimation tool are provided as a resource resource
Experimental setups
Assay System Perturbation Readout Platform
Whole blood transcriptome meta-analysis / differential gene expression 27 published human TB cohorts (discovery dataset; ATB vs HC/LTBI/OLD/Tx) none (observational disease-state comparison) log fold-change of gene expression between disease conditions qRT-PCR, microarray, or RNA-seq (mixed across studies)
Gene covariation network construction N×M logFC matrix from M cohorts and N genes none node centrality (weighted degree) and edge weights (dot products)
Random forest regression modeling (ML) Common 45-gene signature training data from TB cohorts none TB score modeling Mtb infection-to-disease continuum
Model validation (longitudinal TB progression prediction) 10 independent longitudinal human cohorts none AUROC, sensitivity, specificity for incipient-to-active TB risk
Model specificity testing 20 viral infection cohorts (influenza, RSV, rhinovirus, etc.) none AUROC discriminating ATB from viral infection
Treatment monitoring analysis Longitudinal TB treatment cohorts (e.g., GSE89403, GSE67589, GSE157657) standard TB drug treatment TB score correlation with treatment response over time
Gene set enrichment analysis 45-gene signature none enriched immune pathways
Protein-protein association network analysis 45 candidate gene proteins none interconnection of genes/proteins STRING database
Key results
  • Model predicts incipient-to-active TB risk over a 2.5-year period AUROC 0.85, 74.2% sensitivity, 78.3% specificity
  • Model discriminates active TB from viral infection AUROC 0.93 (95% CI 0.91–0.94)
  • 45-gene signature identified as common to active TB across cohorts 45 genes
  • Differential expression patterns correlate strongly across disease conditions r = 0.61–0.96, p ≤ 1e-05
  • Majority of signature genes interconnected in STRING PPI network 71% (32/45)
  • STAT1 forms associations with many proteins in the network n=17
  • TB scores correlate with and are predictive of treatment outcomes
Key statistics
  • other AUROC 0.85, 74.2% sensitivity, 78.3% specificity (incipient-to-active TB risk prediction over 2.5 years)
  • other AUROC 0.93 (95% CI 0.91–0.94) (discriminating active TB from viral infection)
  • correlation r = 0.61–0.96, all p-values ≤ 1e-05 (correlation of averaged logFC expression patterns between disease comparisons)
  • count 45 genes (common ATB-specific gene signature size)
  • count 71% (32/45) (proportion of signature genes interconnected in STRING network)
  • count n=17 (number of proteins STAT1 associates with in the network)
  • count 57 studies (37 TB and 20 viral infections) (total transcriptome datasets used)
  • other top 5% central genes, central in ≥2 of 4 networks, average logFC ≥ 0.5 or ≤ -0.5 (gene signature selection criteria)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper employs a two-stage computational design: first, a custom network-based meta-analysis across 27 discovery transcriptome cohorts builds gene covariation networks from per-cohort differential expression log fold-changes, using permutation-validated edge weights and node centrality to select a 45-gene TB-specific signature; second, two random forest regression models are trained on this signature and validated across 10 independent longitudinal TB cohorts and 20 viral infection cohorts. Discrimination performance is reported as AUROC (with 95% CI for at least one comparison), sensitivity, and specificity; cross-condition consistency of the gene signature is assessed via Pearson correlation of averaged logFC patterns; and pathway relevance is assessed via gene set enrichment analysis.

Replicationbiological Sample sizeSample sizes described per cohort in Table 1; no formal a priori power calculation is mentioned in the excerpt GroupsATB vs. HC, LTBI, OLD, Tx (discovery); LTBI progressors vs. non-progressors over 2.5 years; ATB vs. viral infection (validation) Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Differential gene expression analysis (log fold-change; specific package/method not named in excerpt) Per-cohort comparisons of ATB vs. HC, LTBI, OLD, and Tx within each of the 27 discovery datasets Varies by cohort (range approximately 27–537 per cohort; see Table 1) not stated
Permutation test Validation of the edge weight threshold (≥3) used to retain edges in the gene covariation network null not stated
Pearson correlation (r) Comparison of average logFC patterns between pairs of disease-condition contrasts (ATB vs. HC, LTBI, OLD, Tx) for the 45-gene set; r reported as 0.61–0.96, all p ≤ 1e-05 45 genes not stated
Gene set enrichment analysis Pathway enrichment of the 45-gene signature (IFN-γ, IFN-α/β, IL-6, TLR cascades) 45 genes not stated
Random forest regression ML model training (full and reduced 45-gene sets) for TB score generation Pooled multi-cohort discovery dataset; exact training n not stated in excerpt not stated
AUROC / ROC analysis with sensitivity and specificity at a fixed threshold Validation: TB progression prediction (AUROC 0.85, sensitivity 74.2%, specificity 78.3%); ATB vs. viral infection (AUROC 0.93, 95% CI 0.91–0.94) null na
Approaches that could also have been used
  • Per-cohort logFC values were aggregated via a custom network covariation approach to identify cross-cohort signal
    Could also: A formal random-effects meta-analysis (e.g., using R packages metafor or limma with dream/voom) could also synthesize per-cohort effect sizes weighted by their standard errors — Random-effects meta-analysis explicitly estimates and reports between-study heterogeneity (τ²) and yields confidence intervals for pooled effect sizes, giving a standardized quantification of how consistently each gene responds across cohorts and allowing direct comparison with other meta-analyses
  • The top-5% node centrality cutoff (combined with logFC ≥ 0.5) was used to select the 45-gene signature from the network
    Could also: A permutation-based FDR threshold on centrality scores (e.g., Benjamini-Hochberg q < 0.05 applied to a null distribution of weighted degrees from randomized networks) could also define the selection boundary — A data-driven FDR cutoff provides a principled, reproducible selection criterion whose false-discovery rate is explicitly controlled, which aids replication and helps reviewers interpret how conservative or liberal the threshold is relative to chance
  • Pearson correlation was used to compare average logFC patterns across the four disease-condition contrasts for the 45-gene set
    Could also: Spearman rank correlation could also be used, as it does not assume linearity or normality of the logFC distribution — With a modest gene set (n = 45) and a few flagged outlier genes (HP, CEACAM1 noted as inconsistent), a rank-based measure is more robust to influential observations without requiring distributional assumptions
  • Random forest regression was the sole ML framework evaluated for TB score generation
    Could also: Regularized linear regression (LASSO or elastic net) or gradient boosting could also be applied, with cross-validated hyperparameter tuning and formal comparison between methods — LASSO/elastic net yield explicit feature coefficients that simplify clinical interpretation and implementation (e.g., as a weighted sum score); evaluating multiple model classes with a held-out comparison allows the best-generalizing approach to be selected on principled grounds rather than assumed
  • Model discrimination at validation was reported as AUROC and a single fixed sensitivity/specificity operating point benchmarked against WHO targets
    Could also: Calibration metrics (Brier score, calibration curves, or reliability diagrams) could also be reported alongside AUROC — For a risk-estimation tool intended for clinical screening, calibration—how well predicted probabilities match observed event rates—complements discrimination and is needed to evaluate whether the score functions as an absolute risk estimate rather than a rank-ordering device
  • Longitudinal validation cohort performance was summarized with a single binary AUROC (progressor vs. non-progressor at a fixed endpoint)
    Could also: Time-dependent ROC analysis (e.g., R package timeROC) or the concordance statistic from a Cox model could also evaluate performance across the full 2.5-year prediction horizon — Standard AUROC treats the outcome as binary at one time point, which can obscure how predictive accuracy changes as the prediction horizon varies; time-dependent metrics explicitly account for the censoring structure inherent in longitudinal TB progression data
  • No multiplicity correction is mentioned for the multiple per-cohort DEG analyses or the six pairwise cross-condition correlations
    Could also: A Benjamini-Hochberg FDR correction applied across all pairwise correlation tests, or a Bonferroni correction across the DEG analyses, could also be reported — With six correlation tests and many per-cohort DEG comparisons run in parallel, reporting a multiplicity-corrected threshold alongside nominal p-values makes explicit how many findings would survive a family-wise error control and contextualizes the strength of the cross-condition concordance result
Software: Custom code (GitHub: wenhan-yu/tb-common-gene-signature) · STRING database (protein-protein interaction network)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
20
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE28623 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE34608 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE54992 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE73408 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE83456 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE101702 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE101705 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE103842 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE107993 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE107994 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE107995 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE111368 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE112104 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE116014 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE117827 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE157657 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE17156 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE19439 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE19442 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE19444 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE20346 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE21802 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE25504 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE29536 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE31348 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE36238 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE37250 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE38900 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE39939 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE39940 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE40012 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE40553 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE41055 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE42825 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE42826 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE42830 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE4607 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE50834 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE56153 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE61754 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE61821 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE62147 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE62525 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE6269 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE66099 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE67059 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE67589 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE68004 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE68310 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE69581 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE73072 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE77087 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE79362 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE84076 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE89403 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE94438 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37471455

Paper: Vargas R, Abbott L, Bower D, Frahm N, Shaffer M, Yu WH. Gene signature discovery and systematic validation across diverse clinical cohorts for TB prognosis and response to treatment. PLoS Comput Biol 2023. PMCID PMC10393163. Repo (authors' own, P16 N/A): https://github.com/wenhan-yu/tb-common-gene-signature @ commit 0b59e6c810ac8d0fff640f73cd91d638978ae720 (master, pushed 2022-12-07).

Pipeline (from Methods + repo)

  1. DE analysis across 27 discovery datasets (R: fun/R/*, r-df-process.ipynb).
  2. Gene covariation networks → 45-gene common signature (py-network-building, py-gene-signature).
  3. Feature engineering: pairwise expression ratios (990 features from 45 genes).
  4. Two-stage feature selection (mutual-info + LASSO) → Full=41 pairs, Reduced=12 pairs.
  5. ML model selection via 5×5 nested CV (7 algos) → RandomForestRegressor chosen (py-predictive-model).
  6. Validation on longitudinal progression cohorts + diagnostic cohorts + viral-infection cohorts (fun/validation.py, py-model-validation-*).

In scope (attempted) — scores → reported-metric stage

The repo deposits the per-sample model prediction scores (Y_pred) for the pooled progression cohorts (final-ML-model/opt_score_rocauc/Com_progress_pred_*.csv, 6 signatures × 1183 samples) and the diagnostic cohorts (data/TB_disease_prediction_*.csv, 5490 samples). The reported AUROCs, sensitivity, specificity and their CIs are the deterministic scores→metric stage. We independently recompute that stage (clean-room pure-python: Mann-Whitney AUC, the authors' Hanley-McNeil CI roc_auc_ci, and the Youden/top-left-corner cutoff utilities.rocauc) and compare to the shipped result CSVs and to the paper.

Targets:

  • Table 2 / combined-prognosis figure: full AUROC table, 6 signatures (Full fs2, Reduced fs3, RISK6, BATF2, Suliman4, Sweeney3) × 12 time-interval columns (exclusive + cumulative) — 72 cells incl. CIs.
  • Abstract claim 1: incipient→active-TB 2.5-year AUROC 0.85.
  • Abstract claim 2: ATB vs viral infection AUROC 0.93 (0.91–0.94).
  • Abstract claim 3: WHO sens/spec at Youden cutoff over 30m: 74.2% / 78.3%.

Out of scope / NOT attempted (the documented 20%)

  • Regenerating Y_pred from the deposited RF model pickles (34 MB fs2 / 84 MB fs3). The model input feature matrices (per-sample gene-pair expression ratios for the ~37 GEO/ArrayExpress datasets) are not deposited; only model outputs (scores, figures) and the trained pickles are. Regenerating scores needs the full GEO download + R microarray/RNA-seq QC/normalisation + probe→gene mapping + pairwise-ratio feature build (r-df-process, r-validate-data-process). High effort, high risk of subtle normalisation mismatch → deferred per the brief's 80/20 rule.
  • The 45-gene network-discovery step itself (network construction over 27 datasets) — upstream of the deposited signature; not attempted.
  • Treatment-response monitoring AUCs (Catalysis/SA cohorts) — same dependency on un-deposited feature matrices; not attempted.
  • Wet-lab / manual content: none (paper is fully computational).

Compute note

The reproduced stage is metric recomputation over ≤6×1183 rows — milliseconds, pure stdlib. No heavy compute and therefore no SLURM job was required («host» control-plane class, same as screening). The only «our HPC»-worthy step (model→scores on raw GEO data) is the deferred 20% above. «infra» work dir «path» was therefore not populated (no heavy data downloaded).

Figures / tables: Table
prog_auroc_30m
Reported
0.85
Reproduced
0.85 (0.82 to 0.88)
exact
diag_auroc_viral
Reported
0.93 (0.91 to 0.94)
Reproduced
0.93 (0.91 to 0.94)
exact
sens_30m
Reported
0.742
Reproduced
0.742 (0.693 - 0.792)
exact
spec_30m
Reported
0.783
Reproduced
0.783 (0.756 - 0.810)
exact
prog_table_full_table
Reported
Table 2 + 4 comparator signatures (72 cells, AUROC + 95% CI)
Reproduced
72/72 cells match at 2 d.p.
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score 0

What deviates: nothing within the reproduced scope — all five abstract metrics (prog AUROC 0.85, ATB-vs-viral 0.93 [0.91–0.94], sens 0.742, spec 0.783) and the full 72-cell Table 2 (point AUCs + 95% CIs) reproduce bit-for-bit at 2 d.p. from the authors' deposited per-sample scores. Whose side / severity: no authors' defect and no fabrication signal; the only limitation is data availability — the per-sample input feature matrices for the ~37 GEO datasets are not deposited, so the model→scores inference, the 45-gene discovery step, and treatment-monitoring AUCs could not be independently re-run. Net: a genuinely exact but partial reproduction that certifies deposited-scores↔reported-metrics internal consistency, not the predictive model that produced the scores — hence q1 yellow (inputs unavailable) and q8 yellow (solid, explainable scope gap) while q5/q7 stay green.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

113.3 k
tokens (I/O) · 7 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.