Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Multiomic analysis of malignant pleural mesothelioma identifies molecular axes and specialized tumor profiles driving intertumor heterogeneity.

Nat Genet · 2023
L1 80/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 56% of all assessed papers rank 484 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

IN PROGRESS. Paper is described well enough to reproduce the open downstream layer. Pinned accession GSE29354 is a text-mining FALSE POSITIVE (Bott 2011, a different 2011 MPM array study); the paper's real raw data is EGA EGAS00001004812 (controlled). Raw-read pipelines (alignment-nf/mutect/purple/IntOGen) are out of scope (controlled data + external cohorts). IN-SCOPE open reproduction from IARCbioinfo/MESOMICS_data: cohort/modality counts (C1-C4) reproduce EXACTLY (120/115/109/119); MOFA layer count (C5) exact (7); known drivers present (C11) exact; driver SNV-only frequencies (C12) reproduced deterministically but paper Fig4 combines SNV+SV+CNV with no per-gene % in text -> partial. MOFA factor-count rerun (C6/C7) running on «our HPC». Not attempted: IntOGen 30-driver discovery (C10, external).

💻 Code ↗ 🗄 Data: GSE29354

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 7d00a6099dee
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether the current WHO histopathological classification of malignant pleural mesothelioma (MPM) fully explains its interpatient molecular heterogeneity, or whether additional molecular dimensions (ploidy, tumor cell morphology, adaptive immune response, CpG island methylator profile) better capture this variation.

Core claims
  • The WHO histopathological classification of MPM accounts for only up to ~10% of interpatient molecular differences finding
  • Multi-Omics Factor Analysis (MOFA) of genomic, transcriptomic and epigenomic data identifies four independent, reproducible latent factors (ploidy, morphology, adaptive response, CIMP) collectively explaining up to 61% of interpatient molecular variation finding
  • Only latent factor 2 (morphology) is associated with existing histopathological/molecular classifications; LF1, LF3 and LF4 represent novel, statistically independent sources of variation finding
  • The four latent factors have independent, complementary prognostic value for overall survival in MPM finding
  • MOFA-derived latent factors are associated with orthogonal drug response profiles in MPM cell lines, indicating therapeutic relevance finding
  • CIMP-high phenotype is associated with epigenetic silencing of specific tumor suppressor genes via CpG island methylation mechanism
  • Whole-genome-doubling-positive tumors show upregulation of the E2F targets pathway relative to WGD-negative tumors finding
  • The interdependent morphology and adaptive response factors form a Pareto-optimal triangular structure delimited by extreme phenotypes, reflecting tumor task specialization mechanism
Experimental setups
Assay System Perturbation Readout Platform
whole-genome sequencing (WGS) 120 MPM tumor samples (MESOMICS cohort) none somatic mutations, copy number, structural variants, ploidy
bulk RNA-seq 120 MPM tumor samples (MESOMICS cohort) none gene expression levels
DNA methylation profiling 120 MPM tumor samples (MESOMICS cohort) none methylation at gene body, promoter and enhancer regions
Multi-Omics Factor Analysis (MOFA, integrative statistical modeling) 120 MPM tumors (WGS + RNA-seq + methylation) none latent factors explaining interpatient molecular variance
drug response profiling 59 MPM cell lines (from Iorio et al., de Reyniès et al., Blum et al.) candidate drugs drug response associated with ploidy, morphology and CIMP factor positions
survival analysis / Cox proportional hazards modeling MESOMICS cohort and validation cohorts none hazard ratios for overall survival by latent factor
Key results
  • WHO histopathological type explains only 2-10% (average 6%) of interpatient molecular variance across omics layers 2-10% (avg 6%)
  • Four MOFA latent factors collectively explain 19-61% (average 33%) of interpatient molecular variance 19-61% (avg 33%)
  • Only LF2 is significantly associated with histopathological classification and prior molecular scores median q=6.94e-10 to 6.94e-11
  • LF1 (ploidy factor) strongly correlates with tumor ploidy r=0.87
  • LF4 (CIMP factor) strongly correlates with CIMP index r=0.92
  • Combining all four latent factors improves survival prediction (AUC) over individual factors or prior prognostic markers
  • E2F targets pathway is the most upregulated pathway in WGD-positive vs WGD-negative tumors q=0.048
  • Five COSMIC tumor suppressor genes (CBFA2T3, FBLN2, PRF1, SLC34A2, WT1) show expression negatively correlated with CIMP index and CpG island methylation median q=2.6e-3
Key statistics
  • correlation r=0.87 (LF1 (ploidy factor) vs tumor ploidy)
  • correlation r=0.92 (LF4 (CIMP factor) vs CIMP index)
  • pvalue median q value = 6.94 x 10^-11 (LF2 association with histopathological classification and prior molecular scores)
  • pvalue q value = 0.048 (E2F targets pathway upregulation in WGD+ vs WGD- samples)
  • pvalue median q value = 2.6 x 10^-3 (5 tumor suppressor genes negatively correlated with CIMP index/methylation)
  • fold_change up to 61% (19-61%, avg 33%) (interpatient molecular variance explained by MOFA latent factors vs 2-10% (avg 6%) by WHO type)
  • count n=120 (MESOMICS cohort of MPM tumors profiled by WGS, RNA-seq and methylation)
  • count n=59 (MPM cell lines used for drug response/latent factor analysis)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper describes an observational multiomic cohort study of 120 malignant pleural mesothelioma (MPM) tumors profiled by whole-genome sequencing, transcriptomics and epigenomics (methylation). The main analytical approach was unsupervised integration via Multi-Omics Factor Analysis (MOFA) to derive latent factors, followed by two-sided Pearson correlation tests relating factors to histopathology and molecular scores, differential expression/pathway enrichment comparisons between subgroups (e.g., WGD+ versus WGD-), and survival analyses reporting hazard ratios with 95% confidence intervals. Results were reported using q values (indicating some form of multiple-testing adjustment), exact p/q values, correlation coefficients (r) and hazard ratios with confidence intervals.

Replicationbiological Sample sizeCohort of n=120 MPM tumors profiled by WGS, transcriptomics and epigenomics (Supplementary Tables 1-3); some subgroup comparisons specify n (e.g., n=11 WGD+ samples), with authors noting limited power/sample size for that subgroup and for MME-only survival analyses; no formal power calculation described in this excerpt GroupsHistopathological types (epithelioid/biphasic/sarcomatoid); WGD-positive vs WGD-negative tumors; extremes along MOFA latent factors (ploidy, morphology, adaptive immune response, CIMP) Pairingunpaired Randomization/blindingna DispersionCI Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionyes
Statistical tests used
Test Applied to n Assumptions
Two-sided Pearson correlation test Correlations between MOFA latent factors, histopathological classification, ploidy, CIMP index and previously published molecular scores (Fig. 1b, d, e) n=120 (MESOMICS cohort) not stated
Differential gene expression / pathway enrichment analysis WGD-positive versus WGD-negative tumors, e.g. E2F targets pathway (Supplementary Tables 9-10) n=11 WGD+ samples versus remaining WGD- samples of the n=120 cohort not stated
Survival analysis with hazard ratios and forest plot (implied Cox-type model) Overall survival by MOFA latent factors (Fig. 1f, Extended Data Fig. 5, Supplementary Tables 12-20) not explicitly stated beyond overall cohort (n=120); subgroup analysis restricted to MME samples only noted as lower powered not stated
Correlation between gene expression and CIMP index/CpG methylation Five COSMIC tumor suppressor genes (CBFA2T3, FBLN2, PRF1, SLC34A2, WT1) (Supplementary Table 11) not explicitly stated not stated
AUROC-based comparison of survival model performance Combined latent-factor survival models versus previously proposed prognostic factors (Extended Data Fig. 5) not explicitly stated not stated
Approaches that could also have been used
  • Multi-omic integration was performed using MOFA to derive unsupervised latent factors.
    Could also: Other multi-omic integration frameworks such as Similarity Network Fusion (SNF), iCluster/iClusterPlus, or multi-block partial least squares (e.g., DIABLO/mixOmics) — Applying an alternative integration method to the same data can serve as a cross-validation of factor structure and stability, since different frameworks make different assumptions about how omic layers relate to one another.
  • Associations between latent factors, histopathology, ploidy and CIMP index were assessed with two-sided Pearson correlation tests.
    Could also: Spearman rank correlation — A rank-based correlation is less sensitive to non-linear monotonic relationships and outliers, which can be informative when checking whether a linear correlation coefficient adequately captures the association.
  • Multiple comparisons across correlations and enrichment tests were summarized using q values, without the specific correction method named in the excerpt.
    Could also: Explicitly reporting the method (e.g., Benjamini-Hochberg FDR or Storey's q-value approach) and the exact family/scope of tests corrected together — Naming the method and scope helps readers interpret exactly which comparisons the adjusted significance threshold applies to.
  • The E2F-targets pathway enrichment finding in WGD+ (n=11) versus WGD- tumors could not be replicated in the TCGA cohort, which the authors attribute to low sample size.
    Could also: A meta-analytic or mixed-effects model combining the MESOMICS and TCGA cohorts, or a permutation-based enrichment test such as GSEA with empirical permutation p values — Pooling cohorts or using permutation-based null distributions can increase effective power and provide an alternative way to assess robustness of enrichment signals in small subgroups.
  • Prognostic value of latent factors was assessed via hazard ratios and forest plots, with performance also compared using an AUC-of-ROC-like metric.
    Could also: Multivariable Cox proportional-hazards models that jointly adjust for clinical covariates (e.g., stage, age, treatment), or time-dependent ROC/C-index metrics — Joint adjustment for clinical covariates can help characterize whether latent factors add independent prognostic information beyond established clinical variables, and time-dependent discrimination metrics are tailored to survival outcomes.
  • In the MME-only subgroup, ploidy and CIMP factors were not statistically significant for prognosis, which the authors attribute to limited power, while noting the effect sizes were similar to the full cohort.
    Could also: Reporting a formal post hoc or a priori power/sample-size calculation, or presenting a Bayesian analysis with credible intervals — These approaches can help distinguish an underpowered non-significant result from a true null effect when working with a reduced subgroup size.
Software: Multi-Omics Factor Analysis (MOFA)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36928603 (MESOMICS, Mangiante/Alcala et al., Nat Genet 2023)

Paper: Multiomic analysis of malignant pleural mesothelioma identifies molecular axes and specialized tumor profiles driving intertumor heterogeneity. DOI 10.1038/s41588-023-01321-1 · PMCID PMC10101853.

Accession correction (IMPORTANT — text-mining false positive)

The RU was seeded with GSE29354, but that GEO series is NOT this paper's data. GSE29354 = "Mesothelioma tumor gene expression profiles" (Bott et al. 2011, PMID 21642991, 53 Affymetrix arrays) — an older, different MPM study that MESOMICS merely cites. Verified via NCBI: Series_status = Public on May 19 2011, Series_pubmed_id = 21642991. It is not used for any reproduction here.

The MESOMICS paper's actual data:

  • Raw WGS + RNA-seq + DNA methylation: EGA EGAS00001004812controlled access (request from MESOMICS data access committee). The raw-read pipelines (alignment-nf, mutect-nf, mutect2-nf, purple-nf) require these reads.
  • Open processed "minimum datasets": GitHub IARCbioinfo/MESOMICS_data (phenotypic_map/MESOMICS/) — MAF of all somatic small variants, expression count/FPKM matrices, CNV/LOH/SV/methylation MOFA inputs (D_*_MOFA.RData), sample-overview supplementary tables, and the R scripts that produce the paper's central figures, with a rendered PhenotypicMap_MESOMICS.md (authors' expected output).

Pipeline-derived results: in scope vs out of scope

OUT OF SCOPE — raw-read pipelines on controlled-access data

Result Pipeline(s) Why out of scope
BAM alignment of WGS/WES IARCbioinfo/alignment-nf (BWA+GATK) needs raw FASTQ from EGA (controlled) → data_restricted
Somatic SNV/indel calling mutect-nf, mutect2-nf needs raw BAMs from EGA (controlled)
Copy-number / purity / ploidy purple-nf needs raw BAMs from EGA (controlled)
IntOGen driver discovery (30 genes) IntOGen needs full per-patient mutation set + external runs; not packaged to rerun

The brief's pinned code (alignment-nf) sits entirely in this out-of-scope tier: it is a real, public, third-party-style Nextflow alignment tool, but it cannot be run because its required inputs (raw human sequencing reads) are EGA-controlled.

IN SCOPE — downstream integration on the OPEN processed data (MESOMICS_data repo)

These are genuinely reproducible from public files + the authors' shipped R scripts, and cover the paper's CENTRAL claim ("molecular axes / three specialized tumor profiles"):

  1. Phenotypic map (MOFA → Pareto archetypes)PhenotypicMap_MESOMICS.R

    • MOFA2 on 7 molecular layers, num_factors = 10 (claim C5, C6)
    • Select 4 survival-associated factors, order by R², name them Ploidy/Morphology/Adaptive-response/CIMP (Fig 1a R² matrix; claim C7)
    • ParetoTI k_fit_pch (bootstrap, ks 2:6, 2..n LFs, t_ratio) → 3 archetypes optimal on LF2+LF3 (claim C8); 3 phenotypes = cell-division / tumour-immune / acinar (claim C9)
    • Compute = MODERATELY HEAVY (MOFA 10k iters + 200-bootstrap Pareto) → «our HPC» SLURM.
  2. Deterministic driver-gene mutation frequencies — from shipped MAF var_annovar_maf_corr.txt (claim C11, C12)

    • per-gene mutated-sample fraction for BAP1/NF2/SETD2/TP53/LATS2; compare to the repo's own Supplementary Table (S44-46 SNVs) and the paper's Fig 4.
    • Fully deterministic, MOFA-independent → strongest 1:1 check; light compute.
  3. Cohort N and per-modality counts (claims C1-C4) — count rows/cols of the shipped matrices / sample tables vs reported 120 (115 WGS / 109 RNA-seq / 119 meth).

Compute placement

All cloning + env build on «our HPC» front1 (internet); all data under «infra» reproductions/pmid-36928603/; MOFA+Pareto run as a SLURM job (--partition=std --nodes=1 --cpus-per-task=N, no --mem). «host» holds only small results + pointers.

No-completeness statement

We attempt the open downstream integration + deterministic MAF che

Figures / tables: TablesFig 1Fig 2Fig 2cFig 4
C1
Reported
120 MPM tumors (discovery cohort)
Reproduced
120 (120 rows in TableS2 / 120 unique ID_MESOMICS / 120 per-sample somatic VCFs)
exact
C2
Reported
115 samples with WGS
Reproduced
115 (TableS2 WGS column non-empty)
exact
C3
Reported
109 samples with RNA-seq
Reproduced
109 (TableS2 RNA column non-empty; 128 multiregion libraries in count matrix)
exact
C4
Reported
119 samples with DNA methylation
Reproduced
119 (TableS2 850K column non-empty)
exact
C5
Reported
7 molecular layers integrated by MOFA
Reproduced
7 (7 shipped D_*_MOFA.RData; create_mofa builds 7 views)
exact
C6
Reported
10 MOFA latent factors trained
Reproduced
PENDING (MOFA rerun «job»)
partial
C7
Reported
4 factors named ploidy/morphology/adaptive-response/CIMP
Reproduced
PENDING (top-4 factors by R2 from MOFA rerun)
partial
C8
Reported
3 archetypes, optimal at 2 LFs (LF2,LF3)
Reproduced
3 archetypes (best t_ratio at 2 LFs) per authors' rendered .md + script ks=3 final fit
partial
C11
Reported
MPM drivers BAP1/NF2/SETD2/TP53/LATS2 present
Reproduced
all 5 present in shipped S44 SNVs MAF
exact
C12
Reported
per-gene small-variant mutation frequencies (Fig 4, combined alterations)
Reproduced
SNV-only %: BAP1 16.7, NF2 8.3, SETD2 5.8, TP53 7.5, LATS2 PENDING
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

105.1 k
tokens (I/O) · 4 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.