Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Blood RNA signature RISK4LEP predicts leprosy years before clinical onset.

EBioMedicine · 2021
L1 53/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
53/100
Reproducibility score
1.2 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 14% of all assessed papers rank 1005 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the pipeline-derived DE results from the deposited GEO GSE163498 count matrix; PARTIAL overall. EXACT: the longitudinal null (progressor First vs Second = 0 DEGs) reproduces exactly under a paired edgeR design; sample N (120 = 40/40/40) exact. PARTIAL: the headline First-vs-HHC DEG count reproduces at the same order of magnitude (1390 vs reported 1613, ~86%, robust across two filters) and the RF predictive signature reproduces qualitatively (compact 13-19 gene RF reaches AUC ~0.89 vs reported 0.95-0.97). MISMATCH: the up/down DEG split differs (547/843 vs reported 836/777). Differences are explained by edgeR version (4.0.16 vs authors' 2021), unspecified QC (authors used 39/group vs our 40), unspecified exact filter/test params, and a positional column->group mapping (internal s-IDs absent from public metadata) validated to 87.5% by sex markers. No fabrication signal: all in-scope numbers are reconstructible from shipped data in the right regime; top DEGs (MT-CO1, MT-ND5, G0S2) match the paper's biology. NOT ATTEMPTED: upstream HISAT2->HTSeq alignment from FASTQ (deposited matrix is that pipeline's output; also gated by storage), and the RT-qPCR-based final 4-gene RISK4LEP model + validation AUC (C9/C10, wet-lab, out of scope). INFRA NOTE: «infra» + NFS-home quotas were exhausted; all compute ran on the separate writable «our HPC» filesystem /usw (never «host»).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 53
    assessed: 2026-06-18 ⛓ 4c6e7cbdc075
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a prospective whole-blood transcriptomic signature in household contacts of leprosy patients predict who will develop leprosy before clinical symptoms appear?

Core claims
  • A 4-gene blood RNA signature (RISK4LEP: MT-ND2, REX1BD, TPGS1, UBC) predicts leprosy development 4–61 months before clinical diagnosis finding
  • 1,613 genes are differentially expressed in progressors before diagnosis compared to household contacts who remain disease-free finding
  • A 13-gene prospective RNA-Seq risk signature discriminates progressors from contacts with AUC 95.2% finding
  • No significant intra-individual longitudinal variation in gene expression occurs within leprosy progressors between pre-diagnosis (t=1) and diagnosis (t=2) finding
  • A Random Forest machine learning approach can select gene features and build a model to predict leprosy progression from blood expression data method
  • RISK4LEP is based on unstimulated whole blood and only a few genes, giving it characteristics suitable for rapid point-of-care diagnostic tests resource
  • RT-qPCR validation of the RNA-Seq signature in an independent set reduced it to a final 4-gene signature with AUC 86.4% finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-Seq (whole transcriptome, poly(A) enrichment, globin reduction) whole blood (PAXgene) from leprosy progressors (n=40) and household contacts (n=40), Bangladesh none (natural M. leprae exposure/progression) differential gene expression (TMM-normalized CPM) Illumina NovaSeq6000, 2x150bp paired-end; NEBNext Ultra II Directional RNA Library Prep
RT-qPCR (validation) whole blood from progressors (n=43) and household contacts (n=43), Bangladesh none gene expression as ΔCt relative to GAPDH reference Biomark HD system, 48.48 Dynamic Array IFC, TaqMan Assays (Fluidigm)
Random Forest machine learning classification RNA-Seq discovery set (n=78 after QC) and RT-qPCR validation set none prediction of leprosy progression (AUC, sensitivity, specificity, accuracy) mlr R package v2.17.1
RNA isolation and quality control PAXgene whole blood none RNA concentration and integrity (RIN) QIAcube/PAXgene Blood RNA kit; Qubit RNA BR; Agilent Fragment Analyzer
Functional/pathway analysis RNA-Seq differentially expressed genes none Gene Ontology terms and canonical pathways ClueGO plugin; Ingenuity Pathway Analysis (IPA)
Key results
  • 13-gene prospective RNA-Seq risk signature discriminates progressors from HC AUC=95.2%
  • Final RT-qPCR-validated 4-gene signature RISK4LEP predicts leprosy AUC=86.4%
  • Differentially expressed genes in progressors before diagnosis vs HC 1,613 genes (FDR<0.05)
  • No significant longitudinal DGE within progressors between t=1 and t=2
  • Validation by RT-qPCR identified 4 significantly differentially expressed genes (Mann-Whitney U test) n=4 genes
Key statistics
  • other AUC=95.2% (13-gene RNA-Seq prospective risk signature)
  • other AUC=86.4% (RISK4LEP 4-gene RT-qPCR signature)
  • count 1,613 genes differentially expressed (progressors t=1 vs HC, FDR<0.05)
  • count household contacts recruited (whole blood cohort, Bangladesh)
  • count progressors identified (contacts diagnosed with leprosy 4–61 months after recruitment)
  • pvalue adjusted p-value (FDR) < 0.05 (threshold for differential expression in edgeR)
  • count estimated 5% of M. leprae-exposed people develop disease (leprosy incidence after exposure)
  • other incubation period 2 to >10 years (leprosy disease latency)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper used RNA-Seq differential gene expression (DGE) analysis with edgeR and TMM normalization to compare leprosy progressors before symptom onset (t=1, n=40) with household contacts (HC, n=40) in an unpaired design, and longitudinally within progressors (t=1 vs t=2, paired design, n=40), applying FDR<0.05 as the significance threshold. A Random Forest classifier with leave-one-out cross-validation was then used for feature selection and prediction of leprosy progression, yielding a 13-gene RNA-Seq signature (AUC=95.2%). RT-qPCR validation in an independent set (43 progressors vs 43 HC) used Mann-Whitney U tests at p<0.05 and a second Random Forest model, producing the final 4-gene RISK4LEP signature (AUC=86.4%).

Replicationbiological Sample sizeDiscovery set: 40 progressors and 40 matched HC; validation set: 43 progressors and 43 matched HC; controls were optimally matched on age, sex, date of recruitment, follow-up time, and BCG vaccination; no formal a priori power calculation mentioned GroupsLeprosy progressors before clinical diagnosis (t=1) vs disease-free household contacts; also progressors at t=1 vs at diagnosis (t=2) Pairingmixed Randomization/blindingnot stated Dispersionnone Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR (via edgeR) for RNA-Seq DGE; no multiple-testing correction stated for RT-qPCR Mann-Whitney U tests
Statistical tests used
Test Applied to n Assumptions
edgeR negative binomial GLM with TMM normalization, unpaired design DGE between leprosy progressors at t=1 vs HC (discovery set, RNA-Seq) 40 progressors vs 40 HC not stated
edgeR negative binomial GLM with TMM normalization, paired design Longitudinal DGE within progressors at t=1 vs t=2 (discovery set, RNA-Seq) 40 paired samples not stated
Linear model within edgeR DGE modelled as a linear function of months elapsed between t=1 and t=2 within progressors 40 paired samples not stated
Mann-Whitney U test DGE of RT-qPCR delta-Ct values between progressors and HC (validation set) 43 progressors vs 43 HC not stated
Random Forest classifier with chi-squared feature selection and LOOCV; accuracy, sensitivity, specificity, and AUC reported Prediction of leprosy progression from RNA-Seq expression (discovery) and RT-qPCR delta-Ct (validation) RNA-Seq: training n=62, test n=16; RT-qPCR: training n=73 (including 8 discovery samples), test n=19 na
Approaches that could also have been used
  • RT-qPCR DGE was evaluated with Mann-Whitney U tests at nominal p<0.05 across multiple genes without a stated multiple-testing correction
    Could also: Apply Benjamini-Hochberg FDR correction across the set of RT-qPCR genes tested, consistent with the approach already used in the RNA-Seq stage — When several genes are tested simultaneously, FDR control bounds the expected proportion of false discoveries; applying it to both platforms would make the multiplicity framework consistent and may reduce the chance of spurious inclusions entering the RT-qPCR Random Forest
  • Classifier performance was reported as a single point-estimate AUC (95.2% for RNA-Seq; 86.4% for RT-qPCR)
    Could also: Report 95% confidence intervals around AUC using the DeLong method or nonparametric bootstrap, and complement with calibration metrics such as Brier score or a calibration plot — CIs convey estimation uncertainty for the sample sizes used; calibration metrics indicate whether predicted risk scores track observed event rates, which is especially informative for a clinical early-warning tool
  • Random Forest was the sole machine-learning method used for feature selection and classification
    Could also: LASSO-penalized or elastic-net logistic regression as a complementary approach — Regularized regression yields direct, interpretable coefficient estimates and a built-in feature-selection path; cross-validating both methods and examining the gene-set overlap can strengthen confidence in the selected signature
  • DESeq2 was not used; edgeR was the sole tool for RNA-Seq DGE
    Could also: DESeq2 with median-of-ratios normalization and shrinkage-based fold-change estimation, run in parallel with edgeR — DESeq2 is an equally standard and widely cited RNA-Seq DGE framework; examining the intersection of genes significant in both tools is a common strategy for identifying high-confidence candidates before downstream feature selection
  • A single 80/20 training-test split was used before LOOCV within the training set for both Random Forest models
    Could also: Repeated k-fold (e.g., 10-fold repeated 10 times) or nested cross-validation with an outer test loop and inner tuning loop — A single split can yield variable performance estimates depending on which samples are held out; repeated or nested CV produces a more stable, less optimistic generalization estimate, particularly relevant when total n is under 100
  • Covariates (age, sex, BCG status) were addressed through matching of controls rather than inclusion in the DGE model
    Could also: Include age, sex, and BCG vaccination as covariates in the edgeR design matrix alongside the group factor — Covariate adjustment within the regression model retains all available subjects (no discarding of unmatched individuals), can account for residual confounding after matching, and is a common approach in blood transcriptomic studies where host characteristics affect gene expression
Software: R/edgeR · R/stats (Mann-Whitney U) 3.6.3 · R/mlr (Random Forest) 2.17.1 · RStudio 1.2.5033 · BIOWDL RNA-Seq pipeline 2.0 · HISAT2 2.1.0 · HTSeq-count 0.9.1 · Cutadapt 2.4 · Fluidigm Real-Time PCR Analysis 4.5.2 · ClueGO / Ingenuity Pathway Analysis (IPA)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34090257 (RISK4LEP)

Paper: Tió-Coma et al. 2021, EBioMedicine 68:103379. "Blood RNA signature RISK4LEP predicts leprosy years before clinical onset." PMID 34090257 · PMCID PMC8182229 · DOI 10.1016/j.ebiom.2021.103379.

Data: GEO GSE163498 (BioProject PRJNA686397, SRA SRP298484). Bulk whole-blood RNA-seq, Illumina NovaSeq 6000, 2×150 PE, poly(A)+GLOBINclear, Homo sapiens. 120 samples = 40 leprosy-progressor First (baseline, pre-symptom) + 40 progressor Second (at diagnosis) + 40 household-contact controls (HHC). Code: biowdl/RNA-seq pipeline (paper states v2.0) — HISAT2 v2.1.0 → GRCh38, HTSeq-count v0.9.1 (--stranded reverse), Ensembl 94.

Reported pipeline-derived results (candidate targets)

# Result Reported value Paper loc Pipeline In scope?
R1 DEGs progressor First vs HHC 1613 (836 up / 777 down), edgeR, TMM, BH FDR<0.05 Results / Fig 2 edgeR on count matrix YES (primary)
R2 DEGs progressor First vs Second 0 (none significant) Results edgeR YES
R3 19-gene RNA-seq RF signature; AUC 96.7%, acc 87.5% 19 genes Results / Table random forest on CPM YES (stochastic)
R4 13-gene coding RF signature; AUC 95.2%, acc 87.5% 13 genes (SNHG32, MT-ND4/5/2, MT-CO1, TAOK3, REPS1, MT-CYB, TPGS1, MMRN1, UBC, REX1BD, CCDC85B) Results random forest YES (stochastic)
R5 4-gene RISK4LEP (MT-ND2, REX1BD, TPGS1, UBC); RT-qPCR AUC 86.4% 4 genes Results derived from RF + RT-qPCR OUT (validation is wet-lab RT-qPCR)
R6 RT-qPCR validation cohort metrics (acc 79.0%, sens 87.5%, spec 72.7%) Results RT-qPCR OUT (wet-lab)
R7 TB-signature cross-application AUCs (Sweeney3 51.6%, Suliman2 58.7%, RISK6 78.3%) Results scoring on RNA-seq maybe (secondary)

In scope (computational, attempted here)

  • R1, R2 — differential expression with edgeR (TMM, BH FDR<0.05) on the deposited HTSeq fragments-per-gene matrix. This is the most directly reproducible result.
  • R3, R4 — random-forest signature derivation (stochastic; expect approximate reproduction of gene-count tiers and AUC magnitude, not exact gene lists).

Out of scope (not attempted)

  • Upstream alignment/quantification from FASTQ (HISAT2→HTSeq) is reproducible in principle but the deposited matrix is the pipeline's published output; re-running alignment on 120 NovaSeq 2×150 libraries (~hundreds of GB FASTQ) is gated by the storage blocker below and adds little over validating the deposited matrix. Recorded as reproducible-in-principle, not executed.
  • R5/R6: RT-qPCR validation and the final 4-gene model selection (wet-lab assay).
  • Sample collection, RNA extraction, clinical follow-up (wet-lab / observational).

Key reproduction caveat (auditable)

The deposited count-matrix columns are internal library IDs (s103830-…) that do not appear in any public GEO/SRA/ENA metadata, so columns cannot be mapped to group labels by accession. Mapping is therefore positional (assume count-matrix column order == GSM/series-matrix order: 40 First, 40 Second, 40 HHC). This assumption is independently validated by sex-marker (XIST vs RPS4Y1/DDX3Y) concordance of each subject's paired First/Second samples (see reproduction/AUDIT.md).

Infrastructure deviation (documented)

«infra» and NFS-home user quotas for the compute account were exhausted at run time (writes failed everywhere on «infra»/home; /tmp is a full 50 MB ramdisk). Work was therefore run on the separate writable «our HPC» filesystem /usw (still on «our HPC», never on «host»). Work dir: /usw/u/«user»/humangenetik/ag/schlein/schlein_christian/reproductions_usw/pmid-34090257.

Figures / tables: Fig 2Fig 4
C1
Reported
1613 DEGs progressor-First vs HHC (edgeR, TMM, BH FDR<0.05)
Reproduced
1390 (exactTest, filterByExpr); 1367 (CPM>1 in >=39)
partial
C2
Reported
836 upregulated
Reproduced
547 up in First
did not match
C3
Reported
777 downregulated
Reproduced
843 down in First
did not match
C4
Reported
0 DEGs progressor First vs Second timepoint
Reproduced
0 (paired design, glmQLF and glmLRT)
exact
C7
Reported
13-gene signature AUC 0.952
Reproduced
0.894 (RF LOOCV, top-13 DE genes)
partial
C8
Reported
19-gene signature AUC 0.967
Reproduced
0.892 (RF LOOCV, top-19 DE genes)
partial
C11
Reported
120 samples (40 First + 40 Second + 40 HHC)
Reproduced
120 (40/40/40) in deposited matrix + series matrix
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 53/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is a solid, partial reproduction from the public GSE163498 count matrix. The headline DE signal reproduces at the same order of magnitude (1390 vs 1613 DEGs, ~86%, robust to filter), the longitudinal null is exact (0 DEGs First vs Second, paired design), N is exact (40/40/40), and top DEGs match the paper's MT-heavy biology — no fabrication signal. Deviations sit on our/availability side: edgeR version drift, unspecified QC/filter/test params, and a positional sample mapping (87.5% validated) that plausibly drives the flipped up/down split (547/843 vs 836/777). The predictive core holds qualitatively (AUC ~0.89 vs 0.95–0.97) but not 1:1, and the final RT-qPCR 4-gene RISK4LEP model was out of scope — hence yellow overall.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

266.1 k
tokens (I/O) · 21.9 M incl. cache
43 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.