Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Multi-INTACT: integrative analysis of the genome, transcriptome, and proteome identifies causal mechanisms of complex traits.

Genome Biol · 2025
L1 86/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
86/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 70% of all assessed papers rank 334 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH; essentially 1:1 on the in-scope target. Ran the authors' own power_fdr.R + multi_intact.R (multi_intact_scan, INTACT R pkg) on the shipped per-replicate simulation summary tables to regenerate Fig. 2 (average Power & realized FDR for 7 PCG-implicating methods at the 5% target). All 7 methods' Power and FDR reproduce within ~0.01-0.02 of the published Fig.2 bar heights; Multi-INTACT reproduces as highest-power FDR-controlled method (power 0.486, FDR 0.040) and TWAS/PWAS reproduce as exceeding 5% FDR (the asterisks in Fig.2). Used 81/100 simulation replicates: a Google Drive per-file download quota blocked the last 19 .summary files; the n=81 averages already match within SE, so the remaining 19 would only tighten them. INTACT resolved to 1.2.0 (README pinned 1.0.2) but the used functions (fdr_rst==fdr_rst2, intact, linear prior) are algorithmically identical. NOT ATTEMPTED (the hard ~20%): the metabolite/METSIM real-data analysis (304 / ~70% KBA pairs, per-tissue GPPC) needs the large multi-GB GTEx-eQTL + UKB-pQTL Drive data, an in-house openmp_wrapper binary and an hours-long EM step, and the individual-level METSIM data is not distributable; the additional simulation study (Supp figs) was skipped per 80/20.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.13955303

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 86
    assessed: 2026-06-14 ⛓ 50dab839707b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether jointly modeling multiple 'gene products' (e.g., RNA transcript and protein levels) of a candidate gene, rather than relying on a single molecular phenotype, improves the identification of causal genes underlying complex trait GWAS signals.

Core claims
  • Multi-INTACT achieves higher power than existing single-gene-product methods while maintaining calibrated false discovery rates in simulations. finding
  • Multi-INTACT correctly detects the true causal gene product(s) (expression, protein, or both) underlying a candidate gene's effect. finding
  • Applied to 1408 plasma metabolite GWAS, Multi-INTACT infers 52 to 109% more causal genes than protein-alone or expression-alone analyses. finding
  • Both gene products (expression and protein) are indicated as relevant for most gene nominations made by Multi-INTACT. finding
  • Multi-INTACT is a Bayesian method that aggregates colocalization and TWAS/canonical-correlation evidence across multiple gene products, using an EM algorithm to estimate empirical Bayes priors and posterior model probabilities for PCG and gene-product relevance. method
  • TWAS methods suffer from severely inflated type I error due to failure to account for LD hitchhiking, while colocalization methods control type I error but are overly conservative. finding
  • A candidate gene requires at least one molecular QTL type with gene-level colocalization probability greater than zero as a minimum necessary condition to be classified as a PCG. method
  • The Multi-INTACT model is represented as a structural equation model allowing pleiotropic effects, making it robust to common violations of the exclusion restriction assumption. mechanism
Experimental setups
Assay System Perturbation Readout Platform
multi-SNP fine-mapping, colocalization analysis, TWAS (canonical correlation) simulated data from real GTEx genotypes, chromosome 5, 500 samples, 1198 genes simulated causal models (M_E, M_P, M_E+P, null) assigned per gene power and false discovery rate at 5% FDR control level
TWAS/PWAS weight generation, colocalization analysis, Multi-INTACT integrative analysis human plasma metabolite GWAS (1408 metabolites) integrated with GTEx expression QTL and UK Biobank protein QTL data none (observational GWAS/QTL data) gene probability of putative causality; gene product relevance probability
Key results
  • Multi-INTACT exhibits optimal power while properly controlling type I errors, outperforming TWAS, colocalization, and single-trait INTACT.
  • TWAS methods show severely inflated type I error rates due to LD hitchhiking.
  • Colocalization methods properly control type I error but are overly conservative (low power).
  • Multi-INTACT infers more metabolite causal genes than protein-alone or expression-alone analyses. 52 to 109%
  • Simulated datasets contained approximately 80% non-PCGs and 20% PCGs. ~80%/~20%
Key statistics
  • fold_change 52 to 109% more metabolite causal genes (Multi-INTACT vs. protein-alone or expression-alone analyses across 1408 metabolites)
  • count 1198 genes (simulation region on chromosome 5 with 477K SNPs from 500 GTEx samples)
  • count 100 simulated datasets (total datasets generated for simulation study)
  • other 5% FDR control level (target FDR threshold used to evaluate power and FDR across methods)
  • count 1408 metabolites (GWAS traits analyzed in the real-data application)
  • other ~1500 common cis-SNPs per gene (minimum) (simulation gene inclusion criterion)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

Multi-INTACT is a Bayesian empirical-Bayes framework that aggregates colocalization and TWAS/PWAS evidence across multiple gene products (RNA expression and protein abundance) using canonical correlation between genetically predicted gene-product levels (composite instrumental variables) and GWAS summary statistics, with an EM algorithm estimating prior probabilities over four competing causal models (M0, ME, MP, ME+P). Each gene candidate receives a posterior probability of putative causality whose complement is a local false discovery rate (lfdr) enabling Bayesian FDR control. Performance was evaluated in 100 replicate simulations on real GTEx genotypes comparing average power and realized FDR at a 5% target level against single-gene-product comparators (TWAS, colocalization, single-trait INTACT), then applied to GWAS on 1408 metabolites integrating GTEx eQTL and UK Biobank pQTL datasets.

Replicationbiological Sample sizeSimulations: 500 GTEx samples, 477K SNPs on chromosome 5, 1198 genes, 100 replicate datasets, ~80% non-PCGs / ~20% PCGs; real-data application: 1408 metabolite GWAS traits integrated with GTEx eQTL and UK Biobank pQTL data; individual sample sizes for real data not stated in provided text GroupsMulti-INTACT vs. single-gene-product methods (TWAS, colocalization, single-trait INTACT); causal models ME vs. MP vs. ME+P vs. M0 Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionBayesian FDR control via local false discovery rate (lfdr) derived from posterior probability of putative causality
Statistical tests used
Test Applied to n Assumptions
Canonical correlation of composite instrumental variables (genetically predicted Ê and P̂) with GWAS trait Y Core Multi-INTACT gene-causality test applied to all gene candidates not stated
Pairwise colocalization analysis (eQTL–GWAS and pQTL–GWAS separately) Provides prior colocalization evidence integrated into Multi-INTACT; also run as standalone comparator not stated
Multi-SNP fine-mapping analysis Applied to eQTL, pQTL, and GWAS data to generate TWAS/PWAS weights used as inputs not stated
Empirical Bayes EM algorithm over causal models (M0, ME, MP, ME+P) Estimates prior distribution by pooling all gene candidates per dataset 1198 genes per simulated dataset not stated
Bayesian FDR control via local FDR (lfdr) at 5% target level Genome-wide PCG implication in simulations and real-data application 1198 genes per simulated dataset; 100 replicate datasets not stated
Single-trait TWAS and single-trait INTACT (expression-only and protein-only) Comparator methods in simulation power/FDR evaluation 1198 genes per simulated dataset; 100 replicate datasets not stated
Approaches that could also have been used
  • Multi-INTACT uses canonical correlation of composite genetically predicted gene-product levels as its core test statistic for gene-trait causality
    Could also: Multivariable Mendelian randomization (MVMR) methods such as MVMR-IVW or MVMR-Egger could also jointly model multiple gene products as exposures and estimate their direct causal effects on the trait — MVMR provides per-exposure effect estimates with confidence intervals and supports formal tests of pleiotropy and heterogeneity; the paper explicitly notes the assumed DAG is similar to MVMR, so MVMR would offer a frequentist complement with interpretable, reportable effect sizes for each gene product
  • Colocalization evidence is incorporated pairwise (eQTL–GWAS and pQTL–GWAS) and combined into the Multi-INTACT prior separately for each gene product
    Could also: Multi-trait colocalization methods such as HyPrColoc or moloc could simultaneously test colocalization across eQTL, pQTL, and GWAS signals in a single joint analysis — Joint multi-trait colocalization accounts for the correlation structure among all molecular signals simultaneously rather than aggregating pairwise results, which may be informative when the eQTL and pQTL signals are themselves correlated at a locus
  • The empirical Bayes prior over causal models is estimated via an EM algorithm that pools all gene candidates within each dataset
    Could also: A fully Bayesian hierarchical model with prior distributions over the causal-model mixing weights could also be used to propagate uncertainty in the prior itself — Empirical Bayes treats the EM-estimated prior as fixed; a hierarchical Bayesian approach propagates prior uncertainty through to the posteriors, which may yield better-calibrated gene probabilities when the number of candidates per dataset is modest or the true prior is far from the EM point estimate
  • FDR is controlled using a Bayesian lfdr procedure that relies on the empirical Bayes posterior probabilities
    Could also: A frequentist Benjamini-Hochberg FDR correction applied to p-values from TWAS or a permutation-based procedure could also control FDR at the gene level — Frequentist FDR methods are widely used and straightforward to compare across studies; Bayesian lfdr-based FDR additionally leverages the estimated proportion of true null genes to potentially increase power, but this advantage depends on accurate estimation of that proportion
  • Simulation performance is evaluated at a single fixed FDR threshold (5%) and summarized as average power and realized FDR over 100 replicates
    Could also: Precision-recall curves or ROC curves across the full range of decision thresholds could also be reported alongside or instead of fixed-threshold summaries — Threshold-free curves characterize discrimination at all operating points simultaneously, allowing readers to assess method behavior across different FDR tolerances without the results being tied to any one chosen threshold
  • Simulations use a single genomic region (chromosome 5) from 500 GTEx samples with a fixed ~80/20 non-PCG/PCG split
    Could also: Simulations spanning multiple chromosomes, varying sample sizes, or varying PCG proportions could also be used to characterize robustness — Evaluating across genomic regions with diverse LD architectures, allele frequencies, and QTL densities would characterize how method performance generalizes beyond the specific genetic structure of the chosen chromosome 5 region
Software: GTEx (source of real genotype data for simulations) · coloc (colocalization analysis, cited as ref 14) · TWAS method (cited as ref 15; specific package not named in provided text)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
9
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GO:0001889 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0006082 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0006629 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0006631 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0006644 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0006805 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0007034 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0007041 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0008202 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0008203 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0009636 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0016042 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0042594 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0042632 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0046395 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0055088 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39901160 (Multi-INTACT, Okamoto et al., Genome Biology 2025)

Paper: Multi-INTACT: integrative analysis of the genome, transcriptome, and proteome identifies causal mechanisms of complex traits. DOI 10.1186/s13059-025-03480-2 · PMCID PMC11789355. Code: https://github.com/jokamoto97/multi_intact_paper (Zenodo 10.5281/zenodo.13955303 = code snapshot). Data: Google Drive "Multi-INTACT data" (folder id 18dGOJBgRtPMfo3LHUmKMzmOZt1HHYWMW), linked from the repo README via https://tinyurl.com/mvbrzt2jhttps://public.websites.umich.edu/~xwen/share/multi_intact.html (Cloudflare-gated redirect).

Repo structure (3 analysis units)

  1. main_simulation_study/ — 100 simulated datasets (1198 genes × 1500 cis-SNPs each), no E↔P effects. Pipeline: power_fdr.R reads shipped per-replicate summary tables (sim_data/*/sim.phi_0_6.*.summary) containing TWAS/PWAS z-scores, colocalization GLCPs, and gene-level chi-sq stats, and computes average Power and realized FDR at 5% target for 7 methodsFig. 2.
  2. additional_simulation_study/ — extra DAG classes (E→P effects); Supp. Figs.
  3. metabolite_data_analysis/ — METSIM/UKB-pQTL/GTEx-eQTL real-data integration; per-tissue GPPC results (39 tissues × 2816 files each on Drive). Heavy; uses an in-house openmp_wrapper + an HPC EM prior-estimation step (authors ship the prior estimates). Headline real-data claim: Multi-INTACT recovers 304 (~70%) of annotated KBA gene–metabolite pairs.

In scope (attempted)

  • Fig. 2 (main simulation): Power + realized FDR for Multi-INTACT, INTACT (expr/protein), Colocalization/GLCP (expr/protein), TWAS/PWAS (PTWAS expr/protein). Pipeline = authors' power_fdr.R + multi_intact.R (multi_intact_scan) on the INTACT R package. Fully specified, deterministic, low-hanging → primary target.

Out of scope / not attempted (and why)

  • Additional simulation study (Supp figs) — same machinery as Fig. 2, lower marginal value; skipped per 80/20.
  • Metabolite real-data analysis (304 KBA / 70%, GPPC tissue heatmaps) — requires the large data/ Drive folder (multi-GB GTEx eQTL + UKB pQTL summary stats), the authors' openmp_wrapper binary, and an hours-long EM step. Individual-level METSIM genotype/phenotype data is not distributable (privacy). This is the hard ~20%; not attempted. The shipped per-tissue GPPC outputs (multi_intact_results/) exist but are the authors' precomputed results, not a from-scratch rerun.

Pipelines named

  • TWAS/PWAS: PTWAS z-scores (shipped). Colocalization: GLCP (shipped).
  • INTACT: Bioconductor INTACT::intact + fdr_rst. Multi-INTACT: repo multi_intact.R::multi_intact_scan (Wakefield BF over chi-sq + GLCP-linear prior).
  • Bayesian FDR control: INTACT::fdr_rst (== fdr_rst2 in the INTACT 1.0.2 the paper used).
Figures / tables: Fig. 2
fig2_mintact_power
Reported
~0.48 (Fig.2, dashed-line, highest)
Reproduced
0.486
within tolerance
fig2_mintact_fdr
Reported
~0.04 (Fig.2, <0.05)
Reproduced
0.040
within tolerance
fig2_twas_power
Reported
~0.455 (Fig.2)
Reproduced
0.451
within tolerance
fig2_twas_fdr
Reported
~0.155 (Fig.2, exceeds 5%, asterisk)
Reproduced
0.154
within tolerance
fig2_pwas_power
Reported
~0.455 (Fig.2)
Reproduced
0.469
within tolerance
fig2_pwas_fdr
Reported
~0.155 (Fig.2, exceeds 5%, asterisk)
Reproduced
0.149
within tolerance
fig2_coloc_expr_power
Reported
~0.245 (Fig.2)
Reproduced
0.245
exact
fig2_coloc_prot_power
Reported
~0.245 (Fig.2)
Reproduced
0.256
within tolerance
fig2_intact_expr_power
Reported
~0.32 (Fig.2)
Reproduced
0.319
within tolerance
fig2_intact_prot_power
Reported
~0.32 (Fig.2)
Reproduced
0.335
within tolerance
fig2_qualitative
Reported
Multi-INTACT optimal power with controlled type-I error (Results, Fig.2)
Reproduced
M-INTACT highest power (0.486) among FDR-controlled methods; TWAS/PWAS exceed 5% FDR (~0.15)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 86/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

301 k
tokens (I/O) · 34.5 M incl. cache
36 min
runtime · 0.01 CPU-h
2.6 GB
peak RAM
2
HPC jobs
hummel
machine