Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genomics Define Malignant Transformation in Myeloma Precursor Conditions.

J Clin Oncol · 2025
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for the shipped slice, only PARTIALLY reproducible 1:1. The paper (JCO 2025; 374-patient SMM/MGUS WES+WGS study) reports its headline numbers (cohort sizes, IMWG vs genomic c-index 0.74/0.79, SMM/MGUS progression 36%/5%, genomic-MM classification %) from the PROPRIETARY 'Myeloma Genome Project pipeline' + clinical survival modeling, which is NOT in the shipped repo and partly uses controlled-access data (EGA/dbGaP) -> out of scope, not attempted. The shipped code github.com/bachisiozic/CNV_mmsig (commit 9992eb02) is a focused CNV-signature fitting tool implementing the paper's 'five distinct CNV signatures'. On «our HPC» (R 4.3.3 + GenomicRanges 1.54.1, conda-on-«infra») we reproduced: C1 the reference encodes exactly 5 normalized signatures over 28 features (matches the paper's claim), and C2 the shipped EM + shipped reference correctly recover KNOWN signature mixtures (each pure signature self-recovers at >=99.6%; a 50/50 mixture is recovered as ~0.50/0.50) -- a deterministic, fabrication-free validation that the released algorithm and signature definitions are self-consistent and functional. We did NOT reproduce: the paper's per-patient/per-cohort signature contributions (the tool's segmentation input is not shipped, and producing it needs the unshipped pipeline on ~3.67 TB raw WGS, PRJNA1301307 = the >20%); and C3 the feature extractor errored on synthetic segmentation (not debugged, the 20%). No fabrication indicators found.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 78
    assessed: 2026-06-14 ⛓ f9c94088d723
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can comprehensive genomic profiling redefine myeloma precursor conditions (MGUS and SMM) by distinguishing clones that have already undergone malignant transformation (genomic MM) from truly premalignant clones (genomic MGUS), thereby improving prediction of progression to multiple myeloma?

Core claims
  • Genomics can identify malignant transformation in MGUS and SMM, defining biologically distinct subsets termed genomic MM and genomic MGUS that are indistinguishable from or distinct from MM at the genomic level. finding
  • Most SMM has genomic features of malignant transformation indistinguishable from MM, whereas ~60% of MGUS and ~10% of SMM show no evidence of transformation (genomic MGUS) and do not progress. finding
  • A workflow based on 28 myeloma genomic defining events associated with progression classifies MGUS/SMM into genomic MM versus genomic MGUS. method
  • Integrating genomic features with the IMWG 2/20/20 model significantly improves prediction of progression among genomic MM patients. finding
  • APOBEC mutagenesis, t(4;14) (NSD2;IGH), RAS mutations, hyperdiploidy, and MYC translocations are among the genomic events associated with progression to MM. mechanism
  • Genomic MGUS clones have mutational and CNV profiles resembling normal B cells, suggesting malignant transformation has not occurred and may never occur. finding
  • The MM genomic background and malignant transformation can be acquired very early, even before the clinical SMM phase as currently defined. finding
  • A catalog of 134 myeloma genomic defining events (derived from NDMM) provides the reference framework applied to precursor conditions. resource
Experimental setups
Assay System Perturbation Readout Platform
Whole-exome sequencing (WES) CD138+ sorted clonal plasma cells from MGUS/SMM patients (n=190 WES) none single-nucleotide variants, indels, copy number variants, oncogenic rearrangements, V(D)J Myeloma Genome Project pipeline; aligned to GRCh38
Whole-genome sequencing (WGS) CD138+ sorted clonal plasma cells from MGUS/SMM patients (n=184 WGS) none CNVs, single base substitutions/indels, structural variants/chromothripsis, mutational signatures Myeloma Genome Project pipeline; aligned to GRCh38
CNV signature analysis (de novo) MGUS and SMM tumor genomes none five distinct CNV signatures (SMM CNV SIG 1-5)
Mutational signature analysis (SBS) MGUS/SMM patients with mutational burden >50 SBS none APOBEC presence/absence (SBS2 and SBS13)
Plasma cell sorting / sample QC Bone marrow tumor samples; PBMC or CD138-negative BM as germline match none clonal B-cell population confirmation and normal-cell contamination assessment CD138+ magnetic beads or flow cytometry
Clinical risk modeling (Cox/Kaplan-Meier, c-index) MGUS/SMM patient cohorts (301 with follow-up) none time to progression to MM, IMWG 2/20/20 prognostic performance
Key results
  • Patients with SMM progressed more and earlier than MGUS
  • Progression to MM after median follow-up of 46 months: 36% of SMM vs 5% of MGUS 80/220 (36%) SMM vs 4/81 (5%) MGUS
  • Genomic evidence of malignant transformation (genomic MM) in training set: 39% of MGUS and 94% of SMM 39% MGUS; 94% SMM
  • In validation set, genomic MM in 47% of MGUS and 85% of SMM 47% MGUS; 85% SMM
  • No genomic MGUS patient progressed to MM in either training or validation set
  • Among genomic MM (not on trials, n=158), 35% progressed and 17% remained stable >5 years 55/158 (35%) progressed; 29/158 (17%) stable >5y
  • All IMWG 2/20/20 high-risk SMM were classified as genomic MM; 8% of early-intervention trial patients were genomic MGUS 5/62 (8%) trial patients = genomic MGUS
  • IMWG 2/20/20 c-index for SMM progression was 0.67 (training) and 0.72 (validation) c-index 0.67 (0.63-0.71); 0.72 (0.68-0.76)
Key statistics
  • count 374 patients (290 SMM, 84 MGUS); 277 training, 97 validation (study cohort with WES (n=190) or WGS (n=184))
  • pvalue P < .0001 (SMM vs MGUS difference in progression)
  • count 80 of 220 (36%) SMM and 4 of 81 (5%) MGUS progressed (progression after 46-month median follow-up)
  • correlation c-index 0.67 (0.63-0.71) training; 0.72 (0.68-0.76) validation (IMWG 2/20/20 prediction of SMM progression)
  • count 28 myeloma genomic defining events associated with progression (identified in training set from 134-event catalog)
  • count 55 of 158 (35%) progressed; 29 of 158 (17%) stable >5 years (genomic MM patients not on early-intervention trials)
  • count 5 of 62 (8%) high-risk SMM trial patients classified genomic MGUS (early intervention trial enrollees)
  • count 62 of 374 (16.5%) enrolled in SMM early intervention trials (cohort composition)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a pre-specified training/validation cohort design (n = 277 training, n = 97 validation) of patients with MGUS or SMM who underwent whole-exome (n = 190) or whole-genome sequencing (n = 184). Progression to MM was analyzed with Kaplan-Meier curves and multivariate Cox proportional hazards models; model discrimination was quantified by the concordance index (c-index) with 95% confidence intervals. A genomic classification workflow distinguishing 'genomic MM' from 'genomic MGUS' was built on the training set and validated in the independent cohort, then integrated with the IMWG 2/20/20 clinical risk model to improve progression prediction.

Replicationbiological Sample size374 total patients (277 training, 97 validation) with available WES or WGS; 301 with outcome follow-up; institution-based split (Mayo Clinic and MDACC as validation); no formal power calculation mentioned in provided text GroupsMGUS vs SMM; IMWG 2/20/20 low/intermediate/high-risk strata; genomic MM vs genomic MGUS; progressors vs non-progressors Pairingunpaired Randomization/blindingnot stated DispersionCI Exact p-valuesno Effect sizesyes Confidence intervalsyes
Statistical tests used
Test Applied to n Assumptions
Kaplan-Meier survival analysis with log-rank test (implied by P value reporting for time-to-event comparison) Time-to-progression comparison between MGUS and SMM (Fig 1A); IMWG 2/20/20 risk strata in training and validation cohorts (Figs 1B, 1C); genomic MM vs genomic MGUS (Figs 3B, 3C) 301 patients with outcome follow-up (81 MGUS, 220 SMM); training n = 277; validation n = 97 not stated
Multivariate Cox proportional hazards model Association between baseline clinical parameters (serum creatinine, hemoglobin, IMWG 2/20/20 risk) and risk of SMM progression in the training set (Fig 1E) Training set patients with SMM and outcome follow-up; exact n for this model not stated in provided text not stated
Concordance index (c-index) with 95% confidence interval Discriminative performance of the IMWG 2/20/20 model in training (c-index 0.67 [0.63–0.71]) and validation (c-index 0.72 [0.68–0.76]) cohorts (Fig 1D) 161 training-set and 74 validation-set patients with complete IMWG 2/20/20 data na
Association analysis for genomic events with progression (specific test not named in provided text) Screening of 134 candidate myeloma genomic defining events to identify 28 associated with progression in the training set, considering all patients and SMM only (Fig 2; Data Supplement Tables S3–S5) Training set; exact per-test n not stated not stated
Approaches that could also have been used
  • 134 candidate genomic events were screened individually for association with progression in the training set to select 28 significant events; the specific statistical test and any multiplicity adjustment were not named
    Could also: A regularized Cox regression (e.g., LASSO or elastic-net penalized Cox) applied jointly to all 134 candidate events could also be used for simultaneous variable selection — Penalized regression handles a large number of correlated binary predictors in a single model, intrinsically addresses multiple-testing inflation without a separate correction step, and produces a continuous genomic risk score rather than a binary classification
  • Model discrimination was summarized using the concordance index (c-index) alone
    Could also: Calibration metrics such as calibration plots, Brier score, or integrated Brier score could also be reported alongside the c-index — The c-index measures rank discrimination but does not assess whether predicted event probabilities match observed rates; calibration metrics quantify absolute predictive accuracy, which is complementary for evaluating clinical utility of a risk tool
  • The training/validation split was institution-based (geographic/center allocation), with a roughly 75/25 ratio
    Could also: Repeated k-fold cross-validation or bootstrap internal validation could also be applied to estimate generalization performance — A single fixed split can yield optimistic or pessimistic performance estimates depending on cohort similarity; resampling-based validation provides a more stable estimate of out-of-sample error when total sample size is moderate (n = 374)
  • Progression was analyzed with a standard Cox model, with non-progression events (death before MM, treatment on trial) handled as censored observations
    Could also: A competing-risks regression framework (e.g., Fine-Gray subdistribution hazard model) could also be applied, treating treatment initiation or death before progression as competing events — Standard Cox censoring of competing events can overestimate the cumulative incidence of progression; competing-risks models directly model the cumulative incidence in the presence of events that preclude MM progression, which may be clinically relevant given that 62 patients were enrolled in early-intervention trials
  • Patients were assigned to one of two discrete genomic classes (genomic MM vs genomic MGUS) via a threshold-based workflow on 28 binary events
    Could also: A continuous genomic complexity score (e.g., a weighted count or logistic/Cox-derived score over the 28 events) could also be derived and combined with the IMWG 2/20/20 model in a joint regression — A continuous score preserves gradations of genomic burden, avoids threshold-dependent classification decisions, and facilitates formal interaction testing and integration with clinical variables in a single model
  • Progression-free survival curves were compared with log-rank tests and results reported as P < .0001
    Could also: Restricted mean survival time (RMST) at a prespecified time horizon could also be reported as a complement to log-rank p-values — RMST provides an absolute, clinically interpretable measure of survival benefit that does not require the proportional-hazards assumption and is directly comparable across studies and follow-up windows
Software: Myeloma Genome Project pipeline (custom; available on GitHub)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
7
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41061199

Paper: Maura F, et al. "Genomics Define Malignant Transformation in Myeloma Precursor Conditions." J Clin Oncol 2025. PMID 41061199 · PMCID PMC12614327 · DOI 10.1200/jco-25-01733.

Shipped code: https://github.com/bachisiozic/CNV_mmsig (latest commit 9992eb025ec32c8d35e7d2b201419aa747b9bb3b, 2026-03-18). Per BRIEF rule P16, applying this shipped tool to the described inputs is a valid reproduction target even though it is a focused slice of the paper.

Data accession: SRA/BioProject PRJNA1301307 ("Smoldering Multiple Myeloma", Mayo Clinic) = 57 WGS samples, ~3.67 TB raw reads (open SRA). Other cohorts in the paper are controlled-access (EGA: EGAD00001006363/…3309/EGAS00001004467; dbGaP phs001323.v3.p1).


What the paper reports (high level)

  • 374 patients (290 SMM + 84 MGUS), WES (n=190) + WGS (n=184), aligned to GRCh38 and analyzed with the "Myeloma Genome Project pipeline" (proprietary, not shipped).
  • Downstream: 28 "myeloma genomic defining events", a genomic classifier, and survival modeling (IMWG vs genomic IMWG c-index 0.74/0.79 train/validation; SMM progression 80/220 = 36%, MGUS 4/81 = 5%).
  • De novo identification of five distinct CNV signatures — this is the slice the shipped repo implements/fits.

In scope (attempted — pipeline-derived & shipped)

# Result Pipeline Feasible?
C1 "Five distinct CNV signatures" — the reference profile matrix shipped with the tool CNV_mmsig.R reference Ref/CNV_SIGNATURES_PROFILES.txt YES — verify it encodes exactly 5 signatures over the documented feature set
C2 The CNV-signature fitting method reproduces: generate_cn_feature_matrix()em_signatures_CNV() yields a normalized per-sample contribution vector over the 5 signatures CNV_mmsig.R YES — run end-to-end on a constructed allele-specific segmentation input + shipped references

Out of scope (NOT attempted — the hard 20%, with reasons)

  • Upstream MGP pipeline (alignment + allele-specific copy-number calling that produces the sample/Chrom/start/end/major/minor segments the tool needs): proprietary, not in the repo; would require aligning + CN-calling 3.67 TB of raw WGS — explicitly the >20% we do not chase. The tool's segmentation input is not shipped, so the paper's per-patient signature contributions cannot be reproduced 1:1.
  • Survival / c-index modeling, genomic classifier, 28 defining events, driver frequencies: not in the shipped repo; depend on the full pipeline + clinical data; several cohorts are controlled-access.

Honest consequence

We can reproduce the shipped tool's structure and mechanics (5 signatures; the EM fitting runs and returns a valid decomposition), but not the paper's numeric per-patient/per-cohort signature contributions, because the segmentation inputs are not shipped and regenerating them needs the unshipped proprietary pipeline on multi-TB raw data. Expected status: partial.

C1
Reported
five distinct CNV signatures
Reproduced
shipped reference = 5 signatures x 28 features, each signature column sums to 1.0
exact
C2
Reported
CNV-signature fitting method (em_signatures_CNV) decomposes a profile into the 5 signatures
Reproduced
pure-signature self-recovery min 0.9961 (diag [0.9966,0.9986,0.9961,0.9989,0.9995]); 50/50 mix(sig0,sig2) -> [0.4975,0.0022,0.4946,0.0041,0.0016]; deterministic (seed 42)
within tolerance
C3
Reported
generate_cn_feature_matrix() produces a per-sample CNV feature matrix from segmentation
Reproduced
ERROR 'subscript out of bounds' on synthetic whole-chromosome segmentation; not debugged
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

Reproduction validated only the shipped slice: the 5×28 CNV-signature reference is internally consistent (columns sum to 1.0) and the EM fitter deterministically recovers known mixtures (diag ≥0.9961), with no fabrication indicators. The paper's actual headline results (cohort sizes, c-index 0.74/0.79, SMM/MGUS progression 36%/5%, genomic-MM %) were never tested because the proprietary Myeloma Genome Project pipeline and segmentation inputs are unshipped and several cohorts are controlled-access — a data-availability/scope limitation, not an authors' defect. C3's released feature extractor additionally errored ('subscript out of bounds') on synthetic input and was not debugged. Net: solid-but-limited — the shipped tool is functional and self-consistent, but the central clinical conclusion stands only in limited form for lack of reproducible inputs.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

176.5 k
tokens (I/O) · 13.4 M incl. cache
20 min
runtime · 0.01 CPU-h
1.3 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine