Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Integration of multi-omics and machine learning strategies identifies immune related candidate biomarkers in inflammation-associated hypertrophic cardiomyopathy

Front Immunol · 2025
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH for the core biomarker claims, and they reproduce 1:1 on the paper's own named dataset (GEO GSE141910) using the named third-party tools limma + pROC (the registry 'code' link MRCIEU/TwoSampleMR is a generic MR package, not author code; paper ships no own repo). EXACT: GSE141910 sample split 166 NF / 28 HCM. EXACT: 6 of the 7 named candidate biomarkers (RNF165, SNCA up; SRGN, MARCO, STEAP4, TKT down) are differentially expressed in the reported direction at FDR<0.05; the 7th (SIGLEC9, ENSG00000129385) is absent from GSE141910's processed 20,781-gene matrix so could not be tested. Single-gene ROC-AUCs (0.729-0.890) overlap the reported 0.728-0.778 at the lower bound (TKT 0.729 vs 0.728) but run higher for several genes -> GSE141910 alone is larger/cleaner than the paper's Fig-4 cohort. DEG count differs (3908 vs 472) ONLY because the paper merged 3 datasets (GSE141910+GSE160997+GSE36961) with batch correction whereas we used the single named dataset -> a scope difference, not a discrepancy in method. No fabrication indicators: the named genes are genuinely DE in the expected directions in the real public data. NOT ATTEMPTED (hard 20% / under-specified): Mendelian randomization via TwoSampleMR (exposure eQTL source has no specific accession), the Random-Forest AUC 0.939 + 10-algorithm ML panel + SHAP (needs merged matrix + unspecified hyperparameters), CIBERSORT immune infiltration, qRT-PCR (wet-lab), ceRNA network and drug targeting (downstream DB lookups).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 85
    assessed: 2026-06-14 ⛓ f0c56ff3a5a9
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether immune dysregulation contributes to hypertrophic cardiomyopathy (HCM) progression and aims to identify immune-related candidate biomarkers and therapeutic targets by integrating Mendelian randomization, multi-dataset transcriptomics, machine learning, and experimental validation.

Core claims
  • Seven key immune-related genes (RNF165, SNCA, SRGN, MARCO, STEAP4, SIGLEC9, TKT) were identified for HCM by intersecting DEGs with MR-derived HCM-associated eQTLs. finding
  • Key genes are enriched in immune-related pathways including cytokine activity, leukocyte migration, and JAK-STAT signaling. finding
  • A Random Forest machine learning model achieved the highest diagnostic performance (AUC 0.939), with SHAP analysis identifying MARCO as the top contributor. method
  • HCM samples show altered immune cell infiltration: increased CD4+ T cells and M0 macrophages, decreased M2 macrophages and neutrophils. finding
  • Immune dysregulation contributes to HCM pathogenesis beyond the traditional sarcomere mutation model. mechanism
  • A ceRNA network of 5 mRNAs, 40 miRNAs, and 152 lncRNAs and drug-target predictions (SRGN-heparin; SNCA-33 drugs) were constructed as resources. resource
  • Integration of MR, transcriptomics, SHAP-interpretable ML, and ceRNA/drug-target networks provides a novel framework for immunogenomic analysis of HCM. method
  • qRT-PCR validation supported the transcriptomic expression trends of the identified key genes. finding
Experimental setups
Assay System Perturbation Readout Platform
Bulk transcriptomic differential expression analysis (limma) Human left ventricular myocardium (210 healthy controls, 152 HCM) from GSE141910, GSE160997, GSE36961 none (disease vs control) differentially expressed genes (|log2FC|>0.5, FDR<0.05) GEO datasets; RNA-seq (TPM) and microarray (RMA); sva/ComBat batch correction
Two-sample Mendelian randomization / eQTL-GWAS analysis European cohort: 19,942 eQTLs (exposure); HCM GWAS 507 cases / 489,220 controls (ebi-a-GCST90018861) genetic instrumental variables causal association (IVW p<0.05) for HCM-associated eQTLs TwoSampleMR R package
Machine learning diagnostic modeling (10 algorithms) with SHAP Merged transcriptomic cohort (70% train / 30% test) none classification AUC and feature contribution (SHAP values) caret R package; permshap/shapviz
GO/KEGG functional enrichment analysis DEGs from merged HCM vs control cohort none enriched pathways (qvalue<0.05, top 10) clusterProfiler, org.Hs.eg.db
Immune cell infiltration deconvolution (CIBERSORT) HCM transcriptomic samples none relative proportions of 22 immune cell types; Spearman correlation with key genes CIBERSORT
ceRNA network construction (mRNA-miRNA-lncRNA) Key gene targets none predicted miRNA and lncRNA interactions miRanda, miRDB, miRWalk, TargetScan, spongeScan; Cytoscape 3.10.1
Drug-target prediction HCM key genes none predicted gene-drug interactions DGIdb; Cytoscape
qRT-PCR validation PBMC samples from HCM patients and healthy controls none relative key gene mRNA expression (2^-ΔΔCt, GAPDH reference) SYBR Premix Ex Taq II, CFX96 Real-Time PCR Detection System (Bio-Rad)
Key results
  • 472 DEGs identified between HCM and control samples 472 DEGs
  • 205 HCM-associated eQTLs selected from MR analysis (from 5,430 eQTLs / 25,472 SNPs) 205 eQTLs
  • Seven key genes obtained by intersection of DEGs and MR eQTLs 7 genes
  • Random Forest model had highest diagnostic performance AUC 0.939
  • SHAP analysis identified MARCO as top feature contributor
  • HCM samples showed increased CD4+ T cells and M0 macrophages, decreased M2 macrophages and neutrophils
  • ceRNA network comprised 5 mRNAs, 40 miRNAs, and 152 lncRNAs 5/40/152
  • SRGN identified as target for heparin; SNCA for 33 other drugs 33 drugs
Key statistics
  • other AUC: 0.939 (Random Forest diagnostic model performance)
  • count 472 DEGs (DEGs HCM vs control)
  • count 205 eQTLs (HCM-associated loci after MR/IVW screening)
  • count 19,942 eQTLs (exposure eQTLs analyzed in MR)
  • count 507 HCM cases / 489,220 controls; 24,199,797 SNPs (HCM GWAS outcome dataset ebi-a-GCST90018861)
  • count 210 healthy controls, 152 HCM (transcriptomic samples across three GEO datasets)
  • other |log2FC|>0.5, FDR<0.05 (DEG screening thresholds)
  • pvalue p<5e-08, clump_r2=0.001, clump_Kb=10000 (SNP instrument selection criteria)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

Three GEO transcriptomic datasets (210 healthy controls, 152 HCM patients) were merged with ComBat batch correction and analyzed using limma for differential expression (|logFC|>0.5, FDR<0.05). Two-sample Mendelian randomization with IVW as the primary estimator and four sensitivity methods was applied to 19,942 eQTL exposures against a GWAS HCM outcome; intersection of significant causal eQTLs with DEGs yielded seven key genes. Ten machine learning algorithms were evaluated on a 70/30 random split, the best-performing model (Random Forest, AUC=0.939) was interpreted via SHAP values, and immune cell proportions estimated by CIBERSORT were correlated with key genes using Spearman correlation. qRT-PCR on PBMC samples provided experimental expression validation.

Replicationbiological Sample sizeThree GEO cohorts totaling 210 healthy controls and 152 HCM patients; qRT-PCR validation sample size described only as 'a small number'; MR outcome GWAS: 507 HCM cases and 489,220 European controls GroupsHCM patients vs. healthy controls Pairingunpaired Randomization/blindingnot stated Dispersionnone Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR for DEG discovery; adjusted p-value (qvalueFilter) for GO/KEGG enrichment; no multiplicity correction stated for the seven downstream Wilcoxon tests on key genes
Statistical tests used
Test Applied to n Assumptions
limma moderated empirical-Bayes t-statistic Differential expression analysis between healthy controls and HCM patients across merged GEO datasets 362 total (210 controls, 152 HCM) not stated
Inverse variance weighting (IVW) two-sample Mendelian randomization Primary causal inference linking 19,942 eQTL exposures to HCM GWAS outcome (ebi-a-GCST90018861); significance threshold p<0.05 with heterogeneity p>0.05 507 HCM cases and 489,220 controls for outcome GWAS stated
MR-Egger, weighted median, weighted modal, simple modal (sensitivity MR analyses) Complementary two-sample MR estimators applied alongside IVW to assess pleiotropy 507 HCM cases and 489,220 controls for outcome GWAS stated
Wilcoxon rank-sum test Comparison of seven key gene expression levels between HCM samples and healthy controls 362 total merged cohort (exact per-gene n not restated in methods) not stated
Spearman rank correlation Associations between seven key gene expression levels and CIBERSORT-estimated proportions of 22 immune cell types null not stated
ROC curve / AUC (pROC package) Diagnostic performance of individual key genes and of ten machine learning model outputs in distinguishing HCM from controls 70% training / 30% test random partition of merged cohort na
Approaches that could also have been used
  • Machine learning model performance was estimated on a single random 70/30 train/test split of ~362 samples
    Could also: Repeated k-fold cross-validation (e.g., 10-fold repeated 10 times) or nested cross-validation could also estimate generalization performance — A single split yields a point estimate with high variance at this sample size; cross-validation averages over multiple partitions, providing a more stable AUC estimate with an associated uncertainty range
  • The seven pre-selected key genes were each compared between HCM and controls with individual Wilcoxon tests; no multiplicity adjustment was stated for this family of seven simultaneous tests
    Could also: Bonferroni or Benjamini-Hochberg correction applied across the seven simultaneous Wilcoxon comparisons could also control the family-wise or false-discovery error rate — When multiple tests share a common scientific question, adjustment is a standard option; reporting both raw and corrected thresholds allows readers to assess findings at both levels
  • DEG analysis was performed using limma on a merged cross-platform matrix (RNA-seq converted to TPM plus RMA-normalized microarray) after ComBat batch correction
    Could also: Performing separate differential expression analyses per dataset and combining results via meta-analysis (e.g., metaMA, RankProd, or Fisher's combined p-value) could also integrate evidence across cohorts — Meta-analytic combination explicitly models cross-study heterogeneity and does not require a shared expression scale, complementing the integrated-matrix approach as a sensitivity check
  • Immune cell infiltration was estimated exclusively with CIBERSORT (LM22 signature matrix)
    Could also: Additional deconvolution algorithms such as xCell, MCP-counter, or TIMER could also estimate immune cell proportions from the same expression data — Different algorithms use different reference signatures and statistical assumptions; concordance of immune infiltration findings across methods conveys robustness of the reported associations
  • qRT-PCR validation results were described in terms of directional expression trends without reported dispersion measures or sample size
    Could also: Reporting the validation n, mean fold change, SD (or 95% CI), and a formal test statistic alongside directional conclusions would also characterize the experimental validation quantitatively — These elements are recommended by MIQE guidelines for qPCR reporting; providing them allows readers to assess the precision and replicability of the experimental findings
  • The IVW method served as the primary MR estimator, with four sensitivity methods applied but without explicit reporting of a weighted or consensus result across estimators
    Could also: A structured comparison table of all five MR estimates with confidence intervals, or use of MR-PRESSO for outlier-robust estimation, could also summarize the causal evidence across methods — Displaying point estimates and CIs from all five estimators together facilitates assessment of directional consistency and the influence of potential pleiotropy on the primary IVW result
Software: R/limma 4.3.1 · R/sva (ComBat) · R/TwoSampleMR · R/caret · R/pROC · R/clusterProfiler · R/ggplot2 · R/shapviz + permshap · CIBERSORT · Cytoscape 3.10.1

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Authors · 5
1Qingzhu Liang 2Jinfeng Wang 3Qingxiao Nong 4Shouwen Tao 5Dalang Fang
Citations
4
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41080564

Paper: Liang Q, et al. "Integration of multi-omics and machine learning strategies identifies immune related candidate biomarkers in inflammation-associated hypertrophic cardiomyopathy." Front Immunol 2025. PMID 41080564 · PMCID PMC12510942 · DOI 10.3389/fimmu.2025.1645382

Registry "Code" link: https://github.com/MRCIEU/TwoSampleMR — this is the generic third-party Mendelian-randomization R package, not the authors' own analysis code. The paper ships no own repository. Per BRIEF rule 2 (P16), we reproduce by applying the named third-party tools (limma, pROC) to the paper's own GEO data.

Data: GEO GSE141910 (MAGNet consortium, human left-ventricular RNA-seq). GEO ships per-sample CSVs (one per GSM) of log2-normalized expression keyed by Ensembl gene IDs, plus a series matrix with etiology per sample (Non-Failing Donor / Hypertrophic cardiomyopathy / Dilated CM / Peripartum CM). The paper's GSE141910 subset = 166 Non-Failing controls + 28 HCM.

In scope (pipeline-derived, clearly specified → attempted on GSE141910)

id reported result (paper) pipeline note
C1 472 DEGs, ` logFC >0.5 & FDR<0.05` (Fig 1C,D)
C2 7 candidate biomarkers + direction: up RNF165, SNCA; down SRGN, MARCO, STEAP4, SIGLEC9, TKT (Fig 1/2, Results) limma logFC sign clean 1:1 check — does each named gene show the reported up/down direction in GSE141910?
C3 single-gene ROC-AUC of the 7 genes = 0.728–0.778 (Fig 4A–G) pROC per-gene AUC (HCM vs control) on GSE141910; compare to reported range.

Out of scope / not attempted (and why)

  • Mendelian randomization (TwoSampleMR). Outcome GWAS is named (ebi-a-GCST90018861, HCM) but the exposure is "19,942 eQTLs from the GWAS database" with no specific eQTL accession — the exposure dataset is not identifiable from the text, so the 205-significant-eQTL result is not reproducible as specified. (Hard 20% / under-specified.)
  • Random-Forest AUC 0.939 + 10-algorithm ML panel + SHAP (Fig 5). Requires the merged 3-dataset matrix and many unspecified hyper-parameters/splits. (Hard 20%.)
  • CIBERSORT immune infiltration (Fig 6). Doable but more involved; deferred to keep to a few clear data points.
  • qRT-PCR (Fig 9, Table 2) — wet-lab, not a pipeline output.
  • ceRNA network (Fig 7), drug targeting (Fig 8) — downstream database lookups.

Reproduction target

Run limma + pROC on GSE141910 (the room's named dataset) for C1–C3 above. All compute on «our HPC» SLURM; data stays on «infra».

Figures / tables: Fig 1CFig 4A
C1-sample-split
Reported
166 Non-Failing + 28 HCM (GSE141910)
Reproduced
NF=166, HCM=28
exact
C2-DEG-count
Reported
472 DEGs (|logFC|>0.5 & FDR<0.05, 3 merged datasets)
Reproduced
3908 DEGs (2484 up / 1424 down) on GSE141910 alone
did not match
C3a-RNF165-dir
Reported
up
Reproduced
up (logFC +1.10, FDR 2e-11)
exact
C3b-SNCA-dir
Reported
up
Reproduced
up (logFC +0.79, FDR 6e-06)
exact
C3c-SRGN-dir
Reported
down
Reproduced
down (logFC -0.71, FDR 1e-08)
exact
C3d-MARCO-dir
Reported
down
Reproduced
down (logFC -2.08, FDR 1e-06)
exact
C3e-STEAP4-dir
Reported
down
Reproduced
down (logFC -1.08, FDR 4e-11)
exact
C3f-SIGLEC9-dir
Reported
down
Reproduced
untestable (gene absent from GSE141910 matrix)
partial
C3g-TKT-dir
Reported
down
Reproduced
down (logFC -0.50, FDR 2e-03)
exact
C4-singlegene-AUC
Reported
ROC-AUC 0.728-0.778 (Fig 4A-G)
Reproduced
AUC 0.729-0.890 (6 genes; TKT 0.729 ... RNF165 0.890)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

The core biomarker claims reproduce cleanly on the paper's own named dataset GSE141910: the 166 NF/28 HCM split is exact, 6/6 testable candidate genes (RNF165, SNCA up; SRGN, MARCO, STEAP4, TKT down) are DE in the reported direction at FDR<0.05, and single-gene AUCs (0.729–0.890) overlap the reported 0.728–0.778 at the lower bound — no fabrication indicators. The notable deviations (472 vs 3908 DEGs; AUC ceiling 0.778 vs up to 0.890) sit on the input/cohort side and on our own method choice: the paper merged 3 batch-corrected datasets while we used the single named cohort, which is larger/cleaner. One gene (SIGLEC9) was absent from the matrix and the merged matrix plus MR/ML/SHAP/CIBERSORT steps were under-specified, so those remain unverified rather than refuted. Overall a solid partial reproduction with scope-explained deviations, not an authors' defect.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

123.4 k
tokens (I/O) · 8.7 M incl. cache
16 min
runtime · 0.01 CPU-h
2.5 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine