Genetics of circulating inflammatory proteins identifies drivers of immune-mediated disease risk and therapeutic targets.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡Could not use the authors’ exact input data
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH -> 1:1 REPRODUCED. The brief's enriched code/data were wrong (kauralasoo/eQTL-Catalogue-resources + GSE16879 belong to neither this paper); the real artifacts are github.com/jinghuazhao/INF (SCALLOP-INF pipeline) and GWAS Catalog full summary statistics GCST90274758-GCST90274848 (public, GRCh37, N~14736). Two reproductions: (A) the shipped sentinel table INF1.merge.cis.vs.trans reproduces the abstract numbers EXACTLY -- 180 pQTLs, 59 cis, 121 trans, 70 proteins, 59 with a cis-pQTL; (B) for 6 showcase cis-pQTLs (OPG, CXCL11, TNFB/LTA, CD40, CXCL5, TRAIL) we independently re-derived -log10 p from the deposited beta/SE at each reported sentinel -- all 6 exist at the exact GRCh37 position, all are genome-wide significant (p<5e-10), and the recomputed strengths match the reported values to <0.5% on the log scale; 5/6 sentinels are the marginal-min lead in the +/-1Mb cis window (CXCL5's lead rs425535 is ~6kb from the reported rs450373, same locus, essentially tied). No fabrication concern: the reported pQTL signals are fully derivable from the public deposited data. NOT ATTEMPTED (hard last ~20% / restricted): full pQTL discovery meta-analysis from raw cohorts (individual-level biobank data restricted), conditional/COJO 227-signal count (needs LD ref + GCTA), and colocalization/Mendelian-randomization (Figs 3-6, multi-stage).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 93assessed: 2026-06-15 ⛓ bf757fa4650e
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusIdentifying the genetic determinants (pQTLs) of inflammation-related circulating plasma proteins will yield insights into physiology and the etiology of immune-mediated diseases and help prioritize therapeutic drug targets.
- ★ Genome-wide pQTL mapping of 91 inflammation-related plasma proteins in 14,824 participants identified 180 pQTLs (59 cis, 121 trans). finding
- ★ This pQTL resource enables drug target prioritization by linking proteins between genotype and disease phenotype. resource
- ★ Mendelian randomization and colocalization identify proteins causally involved in immune-mediated disease etiology, revealing shared and distinct effects across diseases. method
- ★ CD40 has directionally discordant effects on risk of rheumatoid arthritis versus multiple sclerosis and inflammatory bowel disease. finding
- ★ Lymphotoxin-α (LTA) is implicated in multiple sclerosis pathogenesis via integration of pQTL with disease GWAS. mechanism
- ★ CXCL5 is implicated in the etiology of ulcerative colitis, supported by elevated gut CXCL5 transcript expression in UC patients. finding
- Genetic variation in NLRC4 alters its expression and inflammasome activity, affecting circulating IL-18 levels (trans-pQTL via caspase-1 cleavage of pro-IL-18). mechanism
- At least 50% of cis-pQTLs may be driven by underlying cognate cis-eQTLs, identified via colocalization across tissues. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Genome-wide pQTL mapping (GWAS of plasma protein levels) | Plasma from 14,824 European-ancestry participants across 11 cohorts | none | Plasma abundance of 91 inflammation-related proteins; genetic associations | Olink Target Inflammation panel |
| pQTL replication | Plasma from 1,585 ARISTOTLE participants | none | Replication/concordance of pQTL effect estimates | Olink |
| pQTL replication | Plasma from 35,556 Icelanders (deCODE) | none | Replication of locus–protein associations | SomaScan (aptamer-based) |
| cis-eQTL colocalization analysis | Whole blood (eQTLGen, n=31,684) | none | Colocalization of eQTL with pQTL; gene expression vs protein | — |
| Multi-tissue cis-eQTL colocalization analysis | Multiple tissues/cell types (GTEx v.8, eQTL Catalogue) | none | Colocalization of cis-eQTLs with cis-pQTLs across tissues | — |
| trans-pQTL mediator identification (bioinformatic) | pQTL meta-analysis data plus annotation databases | none | Prioritized candidate mediating genes | ProGeM; STRINGdb |
| Mendelian randomization and colocalization | pQTL data integrated with disease GWAS | none (genetic instruments) | Causal effects of proteins on immune-mediated disease risk | — |
| Gut transcript expression analysis | Gut tissue from ulcerative colitis patients | disease (UC vs control) | CXCL5 transcript expression level | — |
- – 180 significant pQTLs identified between 108 genomic regions and 70 proteins (59 cis, 121 trans) 180 pQTLs
- – Conditional analyses revealed 47 additional independent signals, raising total to 227 (99 cis, 128 trans) 227 signals
- ▲ Strong correlation between pQTL effect estimates in ARISTOTLE and discovery meta-analysis r=0.97 (cis r=0.99, trans r=0.94)
- – 126 (71%) of 178 testable pQTLs replicated in ARISTOTLE or deCODE at P≤2.8×10^-4 71%
- – trans-pQTL hotspot at SH2B3 (rs3184504) associated with six proteins (CXCL9, CXCL10, CXCL11, CD5, CD244, IL-12B) 6 proteins
- – For 70 of 91 proteins ≥1 pQTL detected; 59 had a cis-pQTL 77% / 65%
- – Only 6 of 40 blood cis-eQTLs colocalized with cognate cis-pQTLs; IL18 eQTL-pQTL directionally discordant 6 of 40
- ▲ Elevated gut CXCL5 transcript expression in patients with ulcerative colitis
- correlation r = 0.97 (Pearson correlation of pQTL effect estimates ARISTOTLE vs discovery)
- correlation cis r=0.99, trans r=0.94 (Replication effect-size correlation by pQTL type)
- pvalue P ≤ 5 × 10^-10 (Genome-wide significance threshold for discovery pQTLs)
- count 180 pQTLs (59 cis, 121 trans) (Total significant locus–protein associations)
- count 227 signals (99 cis, 128 trans) (Total after conditional analysis (+47 independent))
- other 0.003 (NTF3) to 0.285 (CCL8) (Proportion of variance explained by sentinel variants)
- count 32 / 75 significant at P≤5×10^-10 (Replicated pQTLs in ARISTOTLE (32 of 174) and deCODE (75 of 158))
- count 40 of 59 cis-pQTLs had a significant blood cis-eQTL; 32 colocalized in ≥1 tissue (cis-eQTL support for cis-pQTLs)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a genome-wide protein quantitative trait locus (pQTL) study of 91 inflammation-related plasma proteins (Olink Target platform) in 14,824 European-ancestry participants across 11 cohorts. Genetic associations were estimated with linear regression per cohort and combined by fixed-effect meta-analysis, with significance set at a stringent genome-wide threshold (P ≤ 5 × 10^-10); downstream integration used statistical colocalization, eQTL comparison and Mendelian randomization to infer causality. Results were reported as effect estimates with two-sided P values, and findings were assessed for replication in independent cohorts (ARISTOTLE and deCODE).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Linear regression (per-cohort genetic association), combined by fixed-effect meta-analysis | Genome-wide pQTL discovery across 91 proteins; significance P ≤ 5 × 10^-10 | 14,824 participants | not stated |
| Pearson correlation | Correlation of pQTL effect estimates between discovery meta-analysis and ARISTOTLE replication (r = 0.97 overall; cis r = 0.99, trans r = 0.94) | 174 testable pQTL signals (ARISTOTLE n = 1,585) | not stated |
| Linear regression (replication association testing) | Replication in ARISTOTLE (n = 1,585) and deCODE (n = 35,556) at P ≤ 5 × 10^-10 and Bonferroni P ≤ 2.8 × 10^-4 | ARISTOTLE 1,585; deCODE 35,556 | not stated |
| Statistical colocalization (posterior probability ≥ 0.8) | eQTL–pQTL pairs (eQTLGen, GTEx v8, eQTL Catalogue) for the 59 cis-pQTLs | — | not stated |
| Mendelian randomization | Assessing causal effects of proteins on immune-mediated disease risk (e.g., CD40, LTA, CXCL5) | — | not stated |
| Conditional analysis | Identifying additional independent pQTL signals (180 raised to 227) | — | not stated |
-
Per-cohort estimates were combined using fixed-effect meta-analysis.↳ Could also: A random-effects meta-analysis (or reporting between-cohort heterogeneity statistics such as I^2/Cochran's Q alongside the fixed-effect result) could also have been used. — Random-effects models additionally accommodate between-cohort heterogeneity arising from differing populations or assay conditions, and reporting heterogeneity helps readers gauge consistency of effects across the 11 cohorts.
-
Discovery significance used a fixed P-value threshold of 5 × 10^-10.↳ Could also: A false discovery rate approach (e.g., Benjamini–Hochberg) could also have been applied across the protein–variant tests. — An FDR framework expresses the expected proportion of false discoveries among reported signals and can offer a complementary view of the discovery set's reliability alongside a fixed family-wise threshold.
-
Replication concordance was summarized with Pearson's correlation of effect estimates.↳ Could also: Spearman's rank correlation or a Deming/weighted regression that accounts for measurement error in both axes could also have been used. — Rank-based correlation is robust to outliers and non-linearity, while errors-in-variables regression accounts for sampling uncertainty in both the discovery and replication estimates, which can refine the slope of agreement.
-
Causal inference for disease etiology relied on Mendelian randomization with colocalization support.↳ Could also: Sensitivity analyses such as MR-Egger, weighted median, or formal pleiotropy/heterogeneity diagnostics could also accompany the primary MR estimates. — These complementary estimators relax different assumptions about horizontal pleiotropy and help characterize the robustness of a causal estimate to violations of the standard MR assumptions.
-
Colocalization used a posterior probability threshold of ≥ 0.8.↳ Could also: Reporting the full posterior probability distribution or using methods that allow multiple causal variants (e.g., conditional/SuSiE-based colocalization) could also have been presented. — Showing the continuous posterior and allowing multiple causal signals can convey the strength of shared-variant evidence more granularly, especially in regions with allelic heterogeneity.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
A trans-pQTL hotspot at SH2B3 (rs3184504) regulates plasma levels of six inflammation proteins (CXCL9, CXCL10, CXCL11, CD5, CD244, IL12B).other human plasma 2023×1papers★ This paper is the founder (earliest)
-
CXCL5 transcript expression is elevated in gut tissue of ulcerative colitis patients vs controls.RNA-seq human gut up 2023×1papers★ This paper is the founder (earliest)
-
IL18 cis-eQTL and cis-pQTL effects are directionally discordant, indicating limited eQTL-pQTL colocalization.RNA-seq human whole blood mixed 2023×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37563310
Paper: Zhao JH et al. (2023) Genetics of circulating inflammatory proteins identifies drivers of immune-mediated disease risk and therapeutic targets. Nat Immunol 24:1540–1551. PMID 37563310 · PMCID PMC10457199 · DOI 10.1038/s41590-023-01588-w. SCALLOP-INF consortium.
Metadata correction (brief enrichment was WRONG)
The scaffolded manifest.json lists code_url = kauralasoo/eQTL-Catalogue-resources
and data_accession = GSE16879. Neither belongs to this paper (GSE16879 is the
Arijs IBD/infliximab microarray set; the eQTL-Catalogue repo is Alasoo's resource).
The auto-enrichment mismatched. Corrected pointers:
- Code: https://github.com/jinghuazhao/INF (SCALLOP-INF analysis, branch
master, active; the paper's own pipeline). cis/trans logic:cardio/cis.vs.trans.classification.R; sentinel detection:cardio/sentinels.R; shipped final pQTL table:rsid/10-4-2024/INF1.merge.cis.vs.trans. - Data: GWAS Catalog full summary statistics, public, one study per protein:
GCST90274758–GCST90274848 (91 Olink Target-96 Inflammation proteins).
GRCh37, GWAS-SSF v1.0, European, N≈14,736. FTP under
.../summary_statistics/GCST90274001-GCST90275000/<GCST>/.
Study design (from abstract/results)
- 91 plasma inflammatory proteins (Olink Target platform), 14,824 participants, meta-analysis across 11 cohorts (METAL).
- Reported: 180 pQTLs (59 cis, 121 trans) across 70 proteins and 108 regions.
- Genome-wide/study-wide significance threshold p < 5×10⁻¹⁰ (from repo:
--clump-p1 5e-10,--cojo-p 5e-10, Manhattanabline(h=-log10(5e-10))). - cis = lead variant within ±1 Mb of the encoding gene (radius=1e6), else trans.
IN SCOPE (pipeline-derived, public data, low-hanging — attempted)
- Headline pQTL counts — reproduce 180 / 59 cis / 121 trans / 70 proteins
from the shipped sentinel table
INF1.merge.cis.vs.trans(the pipeline output) and re-classify cis/trans independently via the documented ±1 Mb rule. - Independent sentinel verification vs deposited primary data — for a set of showcase cis-pQTLs (OPG, CXCL11, TNFB/LTA, CD40, CXCL5, TRAIL), download the protein's deposited GWAS-Catalog summary statistics and re-derive −log10 p from β/SE at the reported sentinel; confirm it is the genome-wide-significant lead in the cis window. This is the non-circular 1:1 check (deposited data → reported value).
OUT OF SCOPE (hard last ~20% / restricted / not attempted — why)
- pQTL discovery meta-analysis from raw cohorts — individual-level genotype +
Olink protein data are in restricted biobanks (INTERVAL, EstBB, KORA, …);
not public → cannot rebuild the meta-analysis. (
data_restrictedfor primary data; but the derived summary statistics ARE public, which is what we reproduce on.) - Exact conditional/COJO signal count (227 signals after conditioning) — needs an LD reference panel + GCTA-COJO per region; multi-step, beyond 80/20.
- Colocalization (eQTLGen/GTEx) and Mendelian randomization / GSMR (Figs 3–6) — feasible in principle from public sumstats but multi-stage; not attempted here.
- Wet-lab / manual / external-database curation results — out of scope by definition.
No completeness claim
We reproduce the headline pQTL counts and independently verify a sample of cis sentinels against the deposited primary summary statistics. We do NOT reproduce the full discovery meta-analysis, conditional signal counting, colocalization, or MR.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.