Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Enforced MYC expression directs a distinct transcriptional state during plasma cell differentiation.

Life Sci Alliance · 2025
L1 74/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
74/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 43% of all assessed papers rank 644 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

STRONG PARTIAL / headline reproduced. Third-party headline tool PGCNA2 (medmaca/PGCNA, commit 3aed133) applied to the paper's own data (P16-valid). Full faithful pipeline rebuilt on «our HPC» exactly per Methods: STAR 2.6.0c + RSEM 1.3.1 on all 98 SRA runs (PRJNA1093409) against GENCODE v35/GRCh38.p13 -> tximport 1.22.0 -> DESeq2 1.34.0 LRT (FDR<0.01, union of LRT and pairwise contrasts) -> deposited-VST subset -> PGCNA2 -n 1000 -b 100. RESULTS: sample dims 32/48/18 EXACT (C1-C3); PGCNA modules for rnaseq_1 = 16, EXACT match to reported M1-M16 (C4) - the central computational claim, robust (modal=16 across 1000 runs, highest-modularity clustering=16). DE-transcript counts reproduced in the correct range: rnaseq_2 within 5% (within-tol, C6), rnaseq_1 -17% and rnaseq_3 -28% (partial, C5/C7); the residual gap is explained by the paper not specifying the exact set of 'every contrast' unioned nor the apeglm DEG thresholds. C8 (2582 MYC.WT-vs-control DEGs) mismatched (our FDR-only count 6836 is far higher), consistent with an unstated |LFC| cutoff on apeglm-shrunk estimates - not exactly checkable from the text. NOT attempted (out of scope): wet-lab flow/ELISpot/Western panels, the interactive network website, and gene-signature enrichment against the proprietary 43,572-signature DB (not shipped).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 74
    assessed: 2026-06-21 ⛓ 4da566c38999
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests how enforced/deregulated MYC expression (combined with BCL2) impacts the ability of human B cells to complete plasma cell differentiation, and whether this impact depends on specific MYC transactivation domain elements (MYC boxes, notably MBII and residue W135).

Core claims
  • Acute MYC (T58I) and BCL2 overexpression drives an aberrant B-cell differentiation phenotype with altered surface marker expression while functional antibody secretion is retained finding
  • MYC deregulation has little impact on the core regulatory circuitry controlling B-cell identity; induction of BLIMP1 and IRF4 remains largely intact finding
  • Enforced MYC expression dampens expression of secretory programmes (XBP1 targets, immunoglobulin genes) associated with plasma cell differentiation finding
  • MYC overexpression drives diverse changes in gene expression related to translation and metabolism, including induction of classical MYC target genes finding
  • Establishment of the aberrant differentiated state depends on MYC homology box II (MBII) mechanism
  • The MBII dependence resolves to a single conserved amino acid residue, W135 mechanism
  • An in vitro B-cell activation/differentiation model permissive for long-lived plasma cell generation, with removal of CD40/NFκB signalling at day 3, was used to test oncogene impact independent of sustained CD40 signalling resource
  • Parsimonious Gene Correlation Network Analysis (PGCNA) was used to identify modular patterns of coordinated gene expression change method
Experimental setups
Assay System Perturbation Readout Platform
Flow cytometry (surface phenotyping) Primary human peripheral blood memory B cells, retrovirally transduced MYC T58I-t2A-BCL2 overexpression vs MSCV control vs untransduced CD2, CD19, CD20, CD27, CD38, CD138 surface expression, cell size (FSC-A)
EdU/Ki67 proliferation assay Primary human B cells undergoing PC differentiation MYC T58I-t2A-BCL2 overexpression vs untransduced Percentage EdU+Ki67+ cells (cell cycle activity) 1-h EdU pulse labelling, flow cytometry
Antibody quantification Primary human B cells undergoing PC differentiation (culture supernatant) MYC T58I-t2A-BCL2 vs control conditions Total IgM and IgG antibody secretion
Western blot Transduced primary human B cells, day 6 post-transduction MYC T58I-t2A-BCL2 vs control virus MYC and BCL2 protein levels normalized to β-actin
Bulk RNA-seq (time course) Primary human B cells across differentiation days 0, 3, 6, 13, 20 MYC T58I-t2A-BCL2 vs MSCV vs untransduced Genome-wide gene expression, UMAP clustering, individual gene expression (MYC, CD2, BCL2, surface markers, transcription factors, MYC targets, XBP1 targets, immunoglobulin genes)
Parsimonious Gene Correlation Network Analysis (PGCNA) RNA-seq data from same B-cell differentiation time course MYC T58I-t2A-BCL2 vs control Coordinated modular patterns of gene expression change
Cell counting/viability Transduced primary human B cells at days 13 and 20 MYC T58I-t2A-BCL2 vs MSCV vs untransduced Absolute cell number, geometric mean FSC-A (cell size) Counting beads (eBeads) with flow cytometry
Key results
  • T58I-t2A-BCL2 cells show decreased CD27, CD138, and CD19 but increased CD20 expression relative to controls
  • MYC expression is maintained at supra-physiological levels throughout differentiation in T58I-t2A-BCL2 conditions, unlike progressive repression in controls
  • T58I-t2A-BCL2 cells show increased cell size and increased cell number at day 13 and day 20
  • Functional IgM and IgG antibody secretion is established by day 6 and sustained at day 13 despite aberrant phenotype
  • IRF4 and PRDM1 (BLIMP1) induction remains intact with only modest reductions in maximal expression modest
  • XBP1 induction is suppressed at day 6 and all subsequent time points but remains elevated relative to days 0 and 3
  • Immunoglobulin genes (IGHG1, IGHG2, IGHG3, IGHM) show significantly dampened expression at later time points
  • Classical MYC target genes (TERT, JAG2, TRAP1, FABP5) and a wide range of other MYC targets are profoundly increased
Key statistics
  • count n = 2 replicates (Western blot quantification of MYC and BCL2 normalized to β-actin)
  • count n = 1–4 samples per time point and condition (RNA-seq time course sampling across differentiation days 0, 3, 6, 13, 20)
  • other statistical significance denoted by asterisks (*P<0.05; **P<0.01; ***P<0.001; ****P<0.0001) via unpaired two-tailed t test (Comparisons of flow cytometry, cell count, and antibody quantification data across conditions)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper compares untransduced, MSCV-vector, and MYC(T58I)-BCL2-overexpressing human B cells across a differentiation time course using flow cytometry, antibody quantification, and RNA-seq. Group comparisons of phenotypic/functional readouts were assessed with unpaired two-tailed t-tests (and one-way ANOVA in some supplementary panels), with significance reported as threshold categories (ns, *, **, ***, ****) rather than exact p-values. Time-course RNA-seq differential expression across conditions was summarized with FDR-corrected pairwise comparisons at each time point, alongside UMAP for visualization and a correlation-network method (PGCNA) for global/modular expression pattern analysis.

Replicationbiological Sample sizeDescribed qualitatively as 'data representative of at least two independent experiments'; RNA-seq stated as n=1-4 samples per time point/condition; Western blot quantification stated as n=2 replicates. No power calculation or formal sample-size justification described. GroupsUntransduced, MSCV-vector control, and MYC(T58I)-t2A-BCL2-transduced B cells, compared across a differentiation time course (day 0, 3, 6, 13, 20) Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionFDR correction (Benjamini-Hochberg-type not explicitly named)
Statistical tests used
Test Applied to n Assumptions
Unpaired two-tailed t-test Fig 1B, F, G, H (flow cytometry percentages, EdU/Ki67, antibody quantification) not specified beyond 'at least two independent experiments' not stated
Unpaired two-tailed t-test Fig S1B, D, E (Western blot quantification, FSC-A geometric mean, cell counts) n=2 replicates stated for Western blot quantification (Fig S1B); not otherwise specified not stated
One-way ANOVA Fig S1D, E (FSC-A geometric mean, absolute cell counts) not specified not stated
FDR-corrected pairwise comparisons (method/model not specified) RNA-seq gene expression, all pairwise conditions at each time point (Fig 2, Fig S2; Table S2) n=1-4 samples per time point/condition, stated as representative of two independent experiments not stated
Approaches that could also have been used
  • Multiple unpaired two-tailed t-tests were used across several phenotypic/functional panels spanning several time points and pairwise condition comparisons, without a stated multiplicity correction for these tests.
    Could also: A two-way ANOVA (condition x time) with a post-hoc test such as Tukey's or Sidak's, or applying an FDR/Bonferroni correction across the full set of t-tests — This would jointly model time and condition effects and control the family-wise error rate across the many comparisons, which can be informative when the same dataset is tested repeatedly across time points.
  • Significance is reported as threshold categories (ns, *, **, ***, ****) rather than exact p-values.
    Could also: Reporting exact p-values alongside the significance thresholds — Exact p-values let readers gauge the strength of evidence continuously rather than relying on a categorical cutoff.
  • Error bars are described as mean±SD in some panels and mean±SEM in others.
    Could also: Consistently reporting SD (or a 95% confidence interval) throughout — SD conveys the spread of the data itself, while SEM reflects precision of the mean estimate; a single consistent choice, or a 95% CI, can make it easier to compare variability across panels, particularly for small n.
  • RNA-seq differential expression across pairwise conditions/time points was corrected with FDR, but the underlying statistical/test model (e.g., a specific count-based framework) is not named in the text provided.
    Could also: A negative-binomial generalized linear model as implemented in tools such as DESeq2, edgeR, or limma-voom — These frameworks explicitly model the mean-variance relationship of RNA-seq count data and are a standard complementary approach for differential expression testing with FDR control.
  • Some comparisons rely on small sample sizes (e.g., n=2 for Western blot quantification, n=1-4 for RNA-seq per time point/condition) analyzed with parametric tests (t-test/ANOVA).
    Could also: Non-parametric approaches (e.g., Mann-Whitney U) or exact/permutation-based tests — With very small n, non-parametric or exact methods do not rely on assumptions of normality and can be a useful complementary check alongside parametric results.
  • Global transcriptional relationships were explored using UMAP for visualization and PGCNA for modular/network analysis.
    Could also: Complementary approaches such as principal component analysis (PCA) for visualization or WGCNA for network/module detection — Using more than one dimensionality-reduction or network method can provide a cross-check on cluster structure and gene-module assignments derived from a single method.
Software: PGCNA (Parsimonious Gene Correlation Network Analysis)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40721291

Paper: Vardaka et al. 2025, Enforced MYC expression directs a distinct transcriptional state during plasma cell differentiation. Life Sci Alliance. DOI 10.26508/lsa.202402814. Code: https://github.com/medmaca/PGCNA (PGCNA2; commit 3aed133, 2025-04-22) Data: GEO GSE262809 (SuperSeries) = SubSeries GSE262804 (rnaseq_1), GSE262805 (rnaseq_2), GSE262807 (rnaseq_3). Human RNA-seq, NovaSeq 6000, 150bp PE. PRJNA1093409.

Pipeline (from Methods)

FastQC v0.11.8 → TrimGalore v0.6.10 → STAR v2.6.0c (GRCh38.p13) → RSEM v1.3.1 → tximport v1.22.0 → DESeq2 v1.34.0 (LRT, apeglm shrinkage, VST). DESeq2 FDR<0.01 retains transcripts → PGCNA2 (-n 1000 -b 100) → modules.

In scope (pipeline-derived, reproducible)

id result reported location pipeline feasibility
C1 rnaseq_1 sample count 32 Fig3 "15,941 × 32" data deposit EXACT (deposited VST)
C2 rnaseq_2 sample count 48 Fig4C "14,360 × 48" data deposit EXACT (deposited VST)
C3 rnaseq_3 sample count 18 Fig7E "7,148 × 18" data deposit EXACT (deposited VST)
C4 PGCNA modules (rnaseq_1) 16 (M1–M16) Fig3 PGCNA2 run PGCNA2 on VST matrix
C5 DE genes rnaseq_1 (FDR<0.01) 15,941 Fig3 / Methods DESeq2 LRT needs raw counts (realign)
C6 DE genes rnaseq_2 14,360 Fig4C DESeq2 LRT needs raw counts (realign)
C7 DE genes rnaseq_3 7,148 Fig7E DESeq2 LRT needs raw counts (realign)
C8 DEGs MYCwt vs ctrl D13 2,582 Table S2 DESeq2 needs raw counts (realign)

Out of scope (wet-lab / manual / not pipeline-derived)

  • Flow cytometry, ELISpot/ASC functional assays, protein/Western (Fig 1,2,5,6 wet-lab panels).
  • Interactive network website (visualization only).
  • Gene-signature enrichment vs proprietary 43,572-signature DB (DB not shipped) — out of scope.

Strategy

  1. Tier 1 (exact, from deposit): C1–C3 sample dims — confirmed directly from VST headers.
  2. Tier 2 (headline method): C4 — run PGCNA2 -n 1000 -b 100 on rnaseq_1 VST matrix → module count.
  3. Tier 3 (harder, full pipeline): C5–C8 — realign FASTQ (SRA/PRJNA1093409) → STAR→RSEM→DESeq2 LRT to recover exact DE-gene filters + 2,582 DEG. All heavy compute on «our HPC» SLURM.
Figures / tables: Fig3Fig4CFig7ETable
C1
Reported
rnaseq_1 = 32 samples (Fig3 '15,941 x 32')
Reproduced
32 sample columns in GSE262804 VST
exact
C2
Reported
rnaseq_2 = 48 samples (Fig4C '14,360 x 48')
Reproduced
48 sample columns in GSE262805 VST
exact
C3
Reported
rnaseq_3 = 18 samples (Fig7E '7,148 x 18')
Reproduced
18 sample columns in GSE262807 VST
exact
C4
Reported
16 PGCNA modules (M1-M16) from rnaseq_1 (Fig3)
Reproduced
16 modules (best/highest-modularity clustering ModuleNum=16; modal across 1000 Leidenalg runs=16, 370/1000)
exact
C5
Reported
15,941 DE transcripts FDR<0.01 rnaseq_1
Reproduced
13,178 (DESeq2 LRT + 66 pairwise contrasts union, FDR<0.01)
partial
C6
Reported
14,360 DE transcripts rnaseq_2
Reproduced
15,123 (+5.3%)
within tolerance
C7
Reported
7,148 DE transcripts rnaseq_3
Reproduced
5,151 (-27.9%)
partial
C8
Reported
2,582 DEGs MYC.WT vs control day13
Reproduced
6,836 (FDR<0.01) Wald D13_MYC.WT-BCL2 vs D13_MSCV-backbone
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 74/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

Strong partial, headline confirmed. The central computational claim (16 PGCNA modules from rnaseq_1, plus all 32/48/18 sample dimensions) reproduced exactly and robustly (modal 16/1000 runs), so the core conclusion holds. The only deviations are in DE-transcript counts (C5 -17.3%, C7 -27.9%, C6 +5.3%, and C8 6,836 vs 2,582), all explained by the paper's failure to specify the exact unioned contrast set and the apeglm |LFC| threshold, compounded by the deposit shipping only VST rather than raw counts. This is an authors'-side underspecification / reduced-auditability issue, not a fabrication signal — magnitude and direction hold throughout.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

665.1 k
tokens (I/O) · 72.5 M incl. cache
268 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.