Multi-INTACT: integrative analysis of the genome, transcriptome, and proteome identifies causal mechanisms of complex traits.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH; essentially 1:1 on the in-scope target. Ran the authors' own power_fdr.R + multi_intact.R (multi_intact_scan, INTACT R pkg) on the shipped per-replicate simulation summary tables to regenerate Fig. 2 (average Power & realized FDR for 7 PCG-implicating methods at the 5% target). All 7 methods' Power and FDR reproduce within ~0.01-0.02 of the published Fig.2 bar heights; Multi-INTACT reproduces as highest-power FDR-controlled method (power 0.486, FDR 0.040) and TWAS/PWAS reproduce as exceeding 5% FDR (the asterisks in Fig.2). Used 81/100 simulation replicates: a Google Drive per-file download quota blocked the last 19 .summary files; the n=81 averages already match within SE, so the remaining 19 would only tighten them. INTACT resolved to 1.2.0 (README pinned 1.0.2) but the used functions (fdr_rst==fdr_rst2, intact, linear prior) are algorithmically identical. NOT ATTEMPTED (the hard ~20%): the metabolite/METSIM real-data analysis (304 / ~70% KBA pairs, per-tissue GPPC) needs the large multi-GB GTEx-eQTL + UKB-pQTL Drive data, an in-house openmp_wrapper binary and an hours-long EM step, and the individual-level METSIM data is not distributable; the additional simulation study (Supp figs) was skipped per 80/20.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 86assessed: 2026-06-14 ⛓ 50dab839707b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether jointly modeling multiple 'gene products' (e.g., RNA transcript and protein levels) of a candidate gene, rather than relying on a single molecular phenotype, improves the identification of causal genes underlying complex trait GWAS signals.
- ★ Multi-INTACT achieves higher power than existing single-gene-product methods while maintaining calibrated false discovery rates in simulations. finding
- ★ Multi-INTACT correctly detects the true causal gene product(s) (expression, protein, or both) underlying a candidate gene's effect. finding
- ★ Applied to 1408 plasma metabolite GWAS, Multi-INTACT infers 52 to 109% more causal genes than protein-alone or expression-alone analyses. finding
- ★ Both gene products (expression and protein) are indicated as relevant for most gene nominations made by Multi-INTACT. finding
- ★ Multi-INTACT is a Bayesian method that aggregates colocalization and TWAS/canonical-correlation evidence across multiple gene products, using an EM algorithm to estimate empirical Bayes priors and posterior model probabilities for PCG and gene-product relevance. method
- TWAS methods suffer from severely inflated type I error due to failure to account for LD hitchhiking, while colocalization methods control type I error but are overly conservative. finding
- A candidate gene requires at least one molecular QTL type with gene-level colocalization probability greater than zero as a minimum necessary condition to be classified as a PCG. method
- The Multi-INTACT model is represented as a structural equation model allowing pleiotropic effects, making it robust to common violations of the exclusion restriction assumption. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| multi-SNP fine-mapping, colocalization analysis, TWAS (canonical correlation) | simulated data from real GTEx genotypes, chromosome 5, 500 samples, 1198 genes | simulated causal models (M_E, M_P, M_E+P, null) assigned per gene | power and false discovery rate at 5% FDR control level | — |
| TWAS/PWAS weight generation, colocalization analysis, Multi-INTACT integrative analysis | human plasma metabolite GWAS (1408 metabolites) integrated with GTEx expression QTL and UK Biobank protein QTL data | none (observational GWAS/QTL data) | gene probability of putative causality; gene product relevance probability | — |
- ▲ Multi-INTACT exhibits optimal power while properly controlling type I errors, outperforming TWAS, colocalization, and single-trait INTACT.
- ▲ TWAS methods show severely inflated type I error rates due to LD hitchhiking.
- ▼ Colocalization methods properly control type I error but are overly conservative (low power).
- ▲ Multi-INTACT infers more metabolite causal genes than protein-alone or expression-alone analyses. 52 to 109%
- – Simulated datasets contained approximately 80% non-PCGs and 20% PCGs. ~80%/~20%
- fold_change 52 to 109% more metabolite causal genes (Multi-INTACT vs. protein-alone or expression-alone analyses across 1408 metabolites)
- count 1198 genes (simulation region on chromosome 5 with 477K SNPs from 500 GTEx samples)
- count 100 simulated datasets (total datasets generated for simulation study)
- other 5% FDR control level (target FDR threshold used to evaluate power and FDR across methods)
- count 1408 metabolites (GWAS traits analyzed in the real-data application)
- other ~1500 common cis-SNPs per gene (minimum) (simulation gene inclusion criterion)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
Multi-INTACT is a Bayesian empirical-Bayes framework that aggregates colocalization and TWAS/PWAS evidence across multiple gene products (RNA expression and protein abundance) using canonical correlation between genetically predicted gene-product levels (composite instrumental variables) and GWAS summary statistics, with an EM algorithm estimating prior probabilities over four competing causal models (M0, ME, MP, ME+P). Each gene candidate receives a posterior probability of putative causality whose complement is a local false discovery rate (lfdr) enabling Bayesian FDR control. Performance was evaluated in 100 replicate simulations on real GTEx genotypes comparing average power and realized FDR at a 5% target level against single-gene-product comparators (TWAS, colocalization, single-trait INTACT), then applied to GWAS on 1408 metabolites integrating GTEx eQTL and UK Biobank pQTL datasets.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Canonical correlation of composite instrumental variables (genetically predicted Ê and P̂) with GWAS trait Y | Core Multi-INTACT gene-causality test applied to all gene candidates | — | not stated |
| Pairwise colocalization analysis (eQTL–GWAS and pQTL–GWAS separately) | Provides prior colocalization evidence integrated into Multi-INTACT; also run as standalone comparator | — | not stated |
| Multi-SNP fine-mapping analysis | Applied to eQTL, pQTL, and GWAS data to generate TWAS/PWAS weights used as inputs | — | not stated |
| Empirical Bayes EM algorithm over causal models (M0, ME, MP, ME+P) | Estimates prior distribution by pooling all gene candidates per dataset | 1198 genes per simulated dataset | not stated |
| Bayesian FDR control via local FDR (lfdr) at 5% target level | Genome-wide PCG implication in simulations and real-data application | 1198 genes per simulated dataset; 100 replicate datasets | not stated |
| Single-trait TWAS and single-trait INTACT (expression-only and protein-only) | Comparator methods in simulation power/FDR evaluation | 1198 genes per simulated dataset; 100 replicate datasets | not stated |
-
Multi-INTACT uses canonical correlation of composite genetically predicted gene-product levels as its core test statistic for gene-trait causality↳ Could also: Multivariable Mendelian randomization (MVMR) methods such as MVMR-IVW or MVMR-Egger could also jointly model multiple gene products as exposures and estimate their direct causal effects on the trait — MVMR provides per-exposure effect estimates with confidence intervals and supports formal tests of pleiotropy and heterogeneity; the paper explicitly notes the assumed DAG is similar to MVMR, so MVMR would offer a frequentist complement with interpretable, reportable effect sizes for each gene product
-
Colocalization evidence is incorporated pairwise (eQTL–GWAS and pQTL–GWAS) and combined into the Multi-INTACT prior separately for each gene product↳ Could also: Multi-trait colocalization methods such as HyPrColoc or moloc could simultaneously test colocalization across eQTL, pQTL, and GWAS signals in a single joint analysis — Joint multi-trait colocalization accounts for the correlation structure among all molecular signals simultaneously rather than aggregating pairwise results, which may be informative when the eQTL and pQTL signals are themselves correlated at a locus
-
The empirical Bayes prior over causal models is estimated via an EM algorithm that pools all gene candidates within each dataset↳ Could also: A fully Bayesian hierarchical model with prior distributions over the causal-model mixing weights could also be used to propagate uncertainty in the prior itself — Empirical Bayes treats the EM-estimated prior as fixed; a hierarchical Bayesian approach propagates prior uncertainty through to the posteriors, which may yield better-calibrated gene probabilities when the number of candidates per dataset is modest or the true prior is far from the EM point estimate
-
FDR is controlled using a Bayesian lfdr procedure that relies on the empirical Bayes posterior probabilities↳ Could also: A frequentist Benjamini-Hochberg FDR correction applied to p-values from TWAS or a permutation-based procedure could also control FDR at the gene level — Frequentist FDR methods are widely used and straightforward to compare across studies; Bayesian lfdr-based FDR additionally leverages the estimated proportion of true null genes to potentially increase power, but this advantage depends on accurate estimation of that proportion
-
Simulation performance is evaluated at a single fixed FDR threshold (5%) and summarized as average power and realized FDR over 100 replicates↳ Could also: Precision-recall curves or ROC curves across the full range of decision thresholds could also be reported alongside or instead of fixed-threshold summaries — Threshold-free curves characterize discrimination at all operating points simultaneously, allowing readers to assess method behavior across different FDR tolerances without the results being tied to any one chosen threshold
-
Simulations use a single genomic region (chromosome 5) from 500 GTEx samples with a fixed ~80/20 non-PCG/PCG split↳ Could also: Simulations spanning multiple chromosomes, varying sample sizes, or varying PCG proportions could also be used to characterize robustness — Evaluating across genomic regions with diverse LD architectures, allele frequencies, and QTL densities would characterize how method performance generalizes beyond the specific genetic structure of the chosen chromosome 5 region
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39901160 (Multi-INTACT, Okamoto et al., Genome Biology 2025)
Paper: Multi-INTACT: integrative analysis of the genome, transcriptome, and
proteome identifies causal mechanisms of complex traits.
DOI 10.1186/s13059-025-03480-2 · PMCID PMC11789355.
Code: https://github.com/jokamoto97/multi_intact_paper (Zenodo 10.5281/zenodo.13955303 = code snapshot).
Data: Google Drive "Multi-INTACT data" (folder id 18dGOJBgRtPMfo3LHUmKMzmOZt1HHYWMW),
linked from the repo README via https://tinyurl.com/mvbrzt2j →
https://public.websites.umich.edu/~xwen/share/multi_intact.html (Cloudflare-gated redirect).
Repo structure (3 analysis units)
- main_simulation_study/ — 100 simulated datasets (1198 genes × 1500 cis-SNPs
each), no E↔P effects. Pipeline:
power_fdr.Rreads shipped per-replicate summary tables (sim_data/*/sim.phi_0_6.*.summary) containing TWAS/PWAS z-scores, colocalization GLCPs, and gene-level chi-sq stats, and computes average Power and realized FDR at 5% target for 7 methods → Fig. 2. - additional_simulation_study/ — extra DAG classes (E→P effects); Supp. Figs.
- metabolite_data_analysis/ — METSIM/UKB-pQTL/GTEx-eQTL real-data integration;
per-tissue GPPC results (39 tissues × 2816 files each on Drive). Heavy; uses an
in-house
openmp_wrapper+ an HPC EM prior-estimation step (authors ship the prior estimates). Headline real-data claim: Multi-INTACT recovers 304 (~70%) of annotated KBA gene–metabolite pairs.
In scope (attempted)
- Fig. 2 (main simulation): Power + realized FDR for Multi-INTACT, INTACT
(expr/protein), Colocalization/GLCP (expr/protein), TWAS/PWAS (PTWAS expr/protein).
Pipeline = authors'
power_fdr.R+multi_intact.R(multi_intact_scan) on the INTACT R package. Fully specified, deterministic, low-hanging → primary target.
Out of scope / not attempted (and why)
- Additional simulation study (Supp figs) — same machinery as Fig. 2, lower marginal value; skipped per 80/20.
- Metabolite real-data analysis (304 KBA / 70%, GPPC tissue heatmaps) — requires
the large
data/Drive folder (multi-GB GTEx eQTL + UKB pQTL summary stats), the authors'openmp_wrapperbinary, and an hours-long EM step. Individual-level METSIM genotype/phenotype data is not distributable (privacy). This is the hard ~20%; not attempted. The shipped per-tissue GPPC outputs (multi_intact_results/) exist but are the authors' precomputed results, not a from-scratch rerun.
Pipelines named
- TWAS/PWAS: PTWAS z-scores (shipped). Colocalization: GLCP (shipped).
- INTACT: Bioconductor
INTACT::intact+fdr_rst. Multi-INTACT: repomulti_intact.R::multi_intact_scan(Wakefield BF over chi-sq + GLCP-linear prior). - Bayesian FDR control:
INTACT::fdr_rst(==fdr_rst2in the INTACT 1.0.2 the paper used).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.