Integrated multi-omics analysis combined with clinical validation reveals that HLA-DRB5 and ODAPH are causal risk genes for keratoconus.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the upstream transcriptomic pipeline; effectively 1:1 there. The paper's headline DEG number (2884 = union of upregulated DEGs across GSE151631+GSE77938) reproduces BYTE-EXACT from public NCBI GEO raw counts via GEO2R-equivalent DESeq2 (padj<0.05 & |log2FC|>1.5): all four sub-counts (1353,608,1793,186) and both per-dataset totals (1961,1979) match to the gene. GO/KEGG enrichment qualitatively confirmed (TNF, IL-17, cytokine-cytokine receptor, cell adhesion, immune response). NOT attempted (hard-20%): SMR/TSMR/colocalization causal-gene claims (HLA-DRB5, ODAPH) - require KC GWAS GCST90435979 + GTEx v8/eQTLGen besd + SMR binary + coloc with under-specified tissue/LD/liftover glue. ODAPH OR=202.851 is biologically implausible and flagged for the human auditor as a possible instability/fabrication. Out of scope: IHC/qPCR wet-lab validation. NOTES: (1) the BRIEF's code link github.com/xinqi0702/mstate is a text-mining FALSE POSITIVE - it is code for an unrelated UK-Biobank CVD/depression paper; this paper ships no analysis code, so reproduced via described standard tools on the paper's own data (P16). (2) Paper internal inconsistency: Methods say padj<0.01 but Results say padj<0.05; 2884 matches only at 0.05.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 64assessed: 2026-06-14 ⛓ 686863fbcd51
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan an integrated multi-omics approach (transcriptomic DEG analysis plus Mendelian randomization and colocalization) combined with clinical validation identify key causal risk genes for keratoconus (KC), with the hypothesis that HLA-DRB5 and ODAPH are causal risk genes for KC?
- ★ HLA-DRB5 and ODAPH are causal risk genes for keratoconus, supported by SMR and Bayesian colocalization (HLA-DRB5 PP4=0.844, SMR p=0.001, OR=1.768; ODAPH PP4=1.0, SMR p=0.013, OR=202.851). finding
- ★ 2,884 differentially expressed (upregulated) genes were identified in KC, enriched in cell adhesion, immune response, and TNF, IL-17, and MAPK signaling pathways. finding
- ★ Twenty-four genes met the strong causal colocalization criterion (PP4 > 0.8) with KC. finding
- ★ Clinical validation confirmed significantly elevated expression of HLA-DRB5, ODAPH, and MMP-9 in KC cornea and whole blood. finding
- ★ Integration of transcriptome DEG analysis with SMR, TSMR, and Bayesian colocalization is an effective method to identify causal genes for KC. method
- Meplazumab, an HLA-DRB5 inhibitor, is a candidate etiology-targeted therapy for KC identified via drug-target screening (STRING/DrugBank). resource
- Abnormal ODAPH expression may disrupt stable cross-linking of collagen fibers and compromise corneal structural stability; HLA-DRB5 dysfunction may trigger aberrant inflammatory/immune responses. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (transcriptome, reanalyzed via GEO2R/DESeq2) | human corneal tissue (epithelium and stroma), GSE151631: 19 KC patients and 7 controls | none (disease vs control) | differentially expressed gene expression (padj<0.05, |log2FC|>1.5) | Illumina HiSeq 2500, TruSeq Stranded RNA Library Prep Kit |
| bulk RNA-seq (transcriptome, reanalyzed via GEO2R/DESeq2) | human corneal tissue, GSE77938: 25 KC patients and 25 controls (European origin) | none (disease vs control) | differentially expressed gene expression (padj<0.05, |log2FC|>1.5) | Illumina HiSeq 1500, TruSeq Stranded Total RNA LT with Ribo-Zero Human/Mouse/Rat Kit |
| Summary-data Mendelian randomization (SMR) with HEIDI test | KC GWAS (GCST90435979) as outcome; eQTL data (GTEx v8, eQTLGen blood/multi-tissue) as exposure | none (genetic instrumental variables) | causal association statistics (SMR p-value, OR, log2OR) | — |
| Two-sample Mendelian randomization (TSMR / reverse MR) | KC GWAS SNPs as IVs; risk DEGs from GSE151631 and GSE77938 as outcome | none (genetic instrumental variables) | causal direction consistency (MR-PRESSO, MR-Egger intercept) | — |
| Bayesian colocalization | KC GWAS data and DEG eQTL data (±1 Mb window) | none | posterior probability PP4 (colocalization) | — |
| RT-qPCR | corneal tissue and whole blood from 5 KC patients (aged 20-30) vs 5 age-matched donor controls | none (disease vs control) | RNA expression of HLA-DRB5, ODAPH, MMP-9 | — |
| Immunohistochemical staining | KC corneal tissue vs control | none (disease vs control) | protein expression of key genes | — |
| Drug target / protein interaction screening | STRING and DrugBank databases | none | drug-target interactions for core risk genes | — |
- ▲ ODAPH showed near-complete causal confidence for KC by colocalization and SMR PP4=1.0, OR=202.851 (95% CI 3.130–13,146.857)
- ▲ HLA-DRB5 showed strong causal association with KC PP4=0.844, OR=1.768 (95% CI 1.245–2.509)
- ▲ Union of upregulated DEGs across both datasets identified as candidate genes 2,884 genes
- – SMR identified risk and protective genes for KC using 3,882 SNP instrumental variables 35 risk genes and 34 protective genes
- – Genes meeting strong causal colocalization criterion PP4 > 0.8 24 genes
- ▲ HLA-DRB5, ODAPH, and MMP-9 significantly elevated in KC cornea and whole blood by clinical validation
- – GSE151631 DEGs: downregulated and upregulated in KC 608 downregulated, 1,353 upregulated
- – GSE77938 DEGs: downregulated and upregulated in KC 186 downregulated, 1,793 upregulated
- other PP4 = 1.0 (ODAPH Bayesian colocalization posterior probability (H4) with KC)
- other PP4 = 0.844 (HLA-DRB5 Bayesian colocalization posterior probability (H4) with KC)
- pvalue SMR p = 0.013 (ODAPH SMR causal association with KC; OR 202.851 (95% CI 3.130–13,146.857), log2OR 7.664)
- pvalue SMR p = 0.001 (HLA-DRB5 SMR causal association with KC; OR 1.768 (95% CI 1.245–2.509), log2OR 0.822)
- count 2,884 (total upregulated DEGs (union of two datasets) included in study)
- count 3,882 (SNPs meeting eQTL p < 5e-8 threshold used as instrumental variables in SMR)
- count 35 risk genes, 34 protective genes (SMR analysis results with GWAS eQTL as exposure)
- count 24 (genes meeting strong causal colocalization criterion PP4 > 0.8)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper uses an integrated multi-omics design combining DESeq2-based differential expression analysis of two public RNA-seq datasets (GEO), followed by summary-data Mendelian randomization (SMR) and two-sample Mendelian randomization (TSMR) with eQTL and GWAS summary statistics to infer causal gene–disease relationships, and Bayesian colocalization to confirm shared genetic signals. Multiple-testing in SMR was addressed with FDR correction, and causal candidates were further validated in a small clinical cohort (n=5 per group) via RT-qPCR and immunohistochemical staining. Results were reported as ORs with 95% CIs, exact SMR p-values, and Bayesian posterior probabilities (PP4).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 Wald test (negative binomial model, via GEO2R) | Differential expression between KC and control in GSE151631 and GSE77938 independently | GSE151631: 19 KC + 7 controls = 26; GSE77938: 25 KC + 25 controls = 50 | not stated |
| Summary-data Mendelian Randomization (SMR) | Causal association between each DEG (eQTL exposure) and KC GWAS outcome; reported as OR with 95% CI | 3,882 SNPs identified as IVs at eQTL p < 5×10⁻⁸ | stated |
| HEIDI test (heterogeneity in dependent instruments) | Pleiotropy/heterogeneity check for each SMR result (threshold: HEIDI p > 0.05) | null | not stated |
| Bayesian colocalization (coloc, PP4 posterior probability) | Shared causal variant assessment between KC GWAS and DEG eQTL signals; ±1 Mb window | null | not stated |
| Two-sample Mendelian randomization (TSMR, reverse direction) | Reverse causal direction test: KC GWAS SNPs as IVs, DEG expression as outcome; p ≥ 0.05 supports original direction | null | stated |
| MR-Egger regression intercept test | Detection of directional horizontal pleiotropy in TSMR (intercept ≠ 0 at p < 0.05 flags bias) | null | not stated |
| MR-PRESSO outlier test | Detection and removal of outlier SNPs exhibiting horizontal pleiotropy across MR analyses | null | not stated |
| Hypergeometric enrichment test (GO and KEGG; specific test not named) | Functional enrichment of 2,884 unioned DEGs across BP, CC, MF, and KEGG pathways | 2,884 DEGs (union of upregulated genes from both datasets) | not stated |
| RT-qPCR quantification (statistical comparison test not named) | Clinical validation of HLA-DRB5, ODAPH, and MMP-9 expression in KC vs control corneal tissue and blood | 5 KC patients vs 5 age-matched controls | not stated |
-
DEGs from the two datasets were combined by taking the union of upregulated genes, yielding 2,884 DEGs for downstream analyses↳ Could also: The intersection of DEGs replicated across both datasets could also have been used, or a fixed-effects meta-analysis approach (e.g., via the metaMA or RankProd package) pooling effect estimates across datasets — The intersection or meta-analysis approach would prioritize genes consistently dysregulated across both datasets and ancestry backgrounds, potentially increasing specificity and reducing the multiple-testing burden in subsequent MR analyses; the union maximizes sensitivity but includes dataset-specific signals
-
SMR served as the primary causal inference method, with TSMR providing supporting evidence; a single primary IV estimator was not named for TSMR↳ Could also: Standard TSMR estimators such as inverse-variance weighted (IVW), weighted median, and weighted mode could also be reported alongside SMR as a triangulation strategy — Reporting multiple MR estimators with different assumptions about pleiotropy (IVW assumes no pleiotropy; weighted median tolerates up to 50% invalid IVs; MR-Egger allows directional pleiotropy) provides a richer sensitivity framework and is a common practice in two-sample MR reporting guidelines
-
FDR correction was applied to SMR p-values across all tested genes, with the specific FDR algorithm not named↳ Could also: Bonferroni correction or a pre-specified family-wise error rate (FWER) approach could also have been applied, and the specific algorithm (e.g., Benjamini-Hochberg) could be named explicitly — Naming the specific FDR procedure aids reproducibility; Bonferroni would be more conservative and appropriate if independence between tests cannot be assumed given LD structure across tested gene regions
-
Bayesian colocalization was performed using a ±1 Mb genomic window centered on each DEG, with sensitivity analyses at ±500 kb↳ Could also: Alternative colocalization tools such as eCAVIAR (which models multiple causal variants) or SuSiE-coloc (which handles fine-mapped credible sets) could also have been applied — The standard coloc tool (PP4) assumes a single causal variant per region; eCAVIAR and SuSiE-based colocalization relax this assumption and can be more robust in regions with complex LD structure, which is particularly relevant for the HLA region on chromosome 6
-
Clinical validation of gene expression differences was conducted in n=5 KC patients vs n=5 age-matched controls, with the comparison described as 'significantly elevated' without a named statistical test↳ Could also: For two independent groups of n=5, a two-tailed Mann-Whitney U test (non-parametric) or two-tailed Student's t-test with the specific test named and the resulting test statistic, exact p-value, and a dispersion measure (e.g., median [IQR] or mean ± SD) could also be reported explicitly — Naming the test, providing the test statistic, and reporting a dispersion measure alongside the p-value allows readers to assess the magnitude and variability of the observed differences; with n=5 per group, the Mann-Whitney U is often preferred as normality assumptions cannot be well-assessed
-
The two GEO transcriptome datasets differed in ancestry composition (GSE151631: multi-ethnic; GSE77938: European) and were analyzed separately before pooling DEGs↳ Could also: A formal cross-dataset heterogeneity assessment (e.g., Cochran's Q or I² applied to log fold-change estimates) could also have been performed before pooling, or ancestry-stratified analyses could be reported — Assessing heterogeneity between datasets before pooling their DEGs provides information on whether the transcriptomic signal is consistent across ancestries and sequencing platforms, which is relevant to the generalizability of the identified candidate genes
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
21 downstream papers · 4 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Collagen synthesis disruption and downregulation of... 2017 · 88 cites
- Further evaluation of differential expression of ker... 2020 · 21 cites
- Identification of the immune-associated characterist... 2023 · 14 cites
- Co-Expression of Mitochondrial Genes and ACE2 in Cor... 2020 · 13 cites
- RNA-sequencing in ophthalmology research: considerat... 2019 · 7 cites
- Deciphering mitochondrial dysfunction in keratoconus... 2025 · 6 cites
- RNA sequencing of corneas from two keratoconus patie... 2020 · 60 cites
- Identification of the immune-associated characterist... 2023 · 14 cites
- Gene expression profile analyses to identify potenti... 2023 · 10 cites
- Identification of potential biomarkers of myopia bas... 2023 · 6 cites
- Deciphering mitochondrial dysfunction in keratoconus... 2025 · 6 cites
- Comprehensive Bioinformatics Analysis to Reveal Key... 2022 · 5 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 41803193 (keratoconus multi-omics; HLA-DRB5 & ODAPH)
Title: Integrated multi-omics analysis combined with clinical validation reveals that HLA-DRB5 and ODAPH are causal risk genes for keratoconus. Sci Rep 2026. DOI 10.1038/s41598-026-41037-w · PMCID PMC13179358.
Code link in BRIEF is a FALSE POSITIVE
github.com/xinqi0702/mstate (the only "Code" link) is the analysis code for an
unrelated paper — "The Risk of social isolation and loneliness on progression
from incident CVD to subsequent depression" (UK Biobank multistate Cox model,
data265794.xlsx). It has nothing to do with keratoconus, GSE151631, SMR, or
colocalization. So this paper effectively ships no authors' analysis code.
Per BRIEF rule P16 we may still reproduce by applying the described standard
tools to the paper's own public data — which is exactly what we do.
Pipeline-derived results (what the paper actually computes)
| # | Result | Pipeline / tool | Data | In scope? |
|---|---|---|---|---|
| C1 | 2,884 DEGs (union of two datasets) | GEO2R → DESeq2, padj<0.01 & |log2FC|>1.5 | GSE151631 (19 KC/7 ctrl) + GSE77938 (25 KC/25 ctrl), public GEO | YES — primary anchor |
| C2 | DEGs enriched in cell adhesion, immune response, TNF & IL-17 signaling | GO/KEGG (clusterProfiler-equivalent) | DEG list from C1 | YES — qualitative |
| C3 | 24 genes PP4>0.8 causal; HLA-DRB5 PP4=0.844, SMR p=0.001, OR=1.768; ODAPH PP4=1.0, SMR p=0.013, OR=202.851 | SMR + TwoSampleMR + coloc (Bayesian) | KC GWAS GCST90435979 + eQTL (GTEx v8, eQTLGen) | PARTIAL / hard-20% (see below) |
| — | IHC, qPCR protein/RNA of HLA-DRB5/ODAPH/MMP-9 in patient corneas/blood | wet-lab | 5 patients, hospital | OUT (manual/wet-lab) |
80/20 decision
- Primary (do now): C1 DEG union count = 2,884. Cleanly specified (exact tool, exact thresholds, exact public datasets, exact group labels). This is the load-bearing upstream number the entire paper depends on.
- Secondary: C2 enrichment pathway sanity check (cheap, same env).
- Hard 20% (attempt only if cheap, else documented skip): C3 SMR/coloc. Requires GWAS .ma + eQTL besd (GTEx v8 + eQTLGen, multi-GB), the SMR binary, TwoSampleMR + coloc, MR-PRESSO/HEIDI. Under-specified glue (which tissue, exact liftover, clumping ref panel). Red flag: ODAPH OR = 202.851 is a biologically implausible point estimate, the classic signature of a single weak rare instrument — recorded as a possible-instability / possible-fabrication note for the human auditor, not "reproduced".
Data resolves (all checked, control-plane)
- GSE151631 NCBI raw counts: HTTP 200, 1.22 MB. groups:
disease: healthy control/disease: Keratoconus. - GSE77938 NCBI raw counts: HTTP 200, 2.68 MB. groups:
disease state: KTCN/non-KTCN. - KC GWAS GCST90435979 (EBI GWAS Catalog dir): HTTP 200.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The reproducible upstream is excellent: the headline DEG union 2884 (1961 + 1979, sub-counts 1353/608 and 1793/186) reproduces byte-exact from public GEO raw counts, and all five enrichment themes (TNF, IL-17, cytokine-cytokine receptor, cell adhesion, immune response) qualitatively confirm. However, the paper's title claim — that HLA-DRB5 and ODAPH are causal risk genes — rests on SMR/coloc that could not be attempted (no code shipped, the cited code repo is an unrelated UK-Biobank false positive, and tissue/LD/liftover steps are unspecified). Crucially, ODAPH OR=202.851 with PP4=1.0 is biologically implausible and too-perfect — a possible-fabrication/weak-instrument-instability signature on the central gene — so the headline causal values are neither derivable nor confirmed here. This places the defect on the authors'/data-availability side, makes the central conclusion only weakly supported (upstream dysregulation, not causality), and warrants a critical, fabrication-suspect overall grade despite the flawless DEG reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.