Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Aberration in DNA methylation in B-cell lymphomas has a complex origin and increases with disease severity.

PLoS Genet · 2013
L1 49/100 3/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
49/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 7% of all assessed papers rank 1081 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the computational core of the paper's HELP-array M-score phylogeny/heterogeneity pipeline (github.com/lima1/maphylogeny) for the 3 of 6 original sample groups available on GEO (CD34+ progenitors n=8 from GSE18700; GCB-DLBCL n=40 and ABC-DLBCL n=20 from GSE23967, 25626 common HpaII/MspI probes on GPL6604). The paper's NBC/NGC/FL methylation cohorts have no GEO accession ('pending GEO accession number' per the paper itself, confirmed absent from GEO as of this run) and are excluded from scope, not silently dropped. Core heterogeneity finding (M-score IQR higher in lymphoma than in normal CD34+ progenitors, Mann-Whitney) reproduced with matching significance (p<2.2e-16 in both paper and this run) and correct direction (median IQR lymphoma=0.755 vs CD34=0.402). Group-level and sample-level bootstrap phylogenies (1000 and 200 bootstraps respectively, ape::fastme.bal, Pearson correlation distance) were built; PHYLIP fconsense (used by the paper via Dendroscope) was substituted with ape::consensus() since PHYLIP is unavailable on this cluster -- a documented, transparent deviation. Paper-reported subgroup sample counts (GCB=39, ABC=18) differ slightly from GEO-deposited counts observed here (GCB=40, ABC=20, plus 9 non-classifiable) -- an unexplained but minor dataset-provenance discrepancy, not a pipeline error. The paper's gene-density- and CTCF-site-stratified heterogeneity sub-analyses were not reproduced (only pooled, unstratified comparisons were run) -- documented as a scope gap. All numeric results below are provisional and intended for human audit.

💻 Code ↗ 🗄 Data: GSE23967

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The authors hypothesized that direct comparison of genome-wide DNA methylation patterning in normal B-cells, follicular lymphomas, and diffuse large B-cell lymphomas would reveal how gene deregulation arises during lymphomagenesis and explain the different clinical behavior of these lymphoma subtypes, which share many of the same mutant alleles.

Core claims
  • B-cell non-Hodgkin lymphomas display striking intra-tumor (intra-sample) and inter-patient (inter-sample) cytosine methylation heterogeneity that increases progressively with disease aggressiveness (NBC<NGC<FL<GCB<ABC). finding
  • Epigenetic heterogeneity is initiated already in normal germinal center B-cells, which are more heterogeneous than naive B-cells, and may cooperate with somatic mutations to predispose NGC to malignant transformation. finding
  • The extent of aberrant methylation (distance of a tumor's methylation pattern from that of normal B-cells) is a significant predictor of survival and improves prognostic prediction beyond the International Prognostic Index in DLBCL. finding
  • Patterns of aberrant methylation are non-random and depend on chromosomal region and local gene density: centromeric and gene-poor regions progressively lose methylation, while gene-rich and intermediate regions gain intra-sample variation. finding
  • Aberrant methylation states spread locally to neighboring promoters in the same direction, with the effect decaying with genomic distance and being stronger for hypo-methylation; spreading is limited by CTCF insulator binding sites. mechanism
  • DNA methylation abnormalities arise via two distinct processes: lymphomagenic transcriptional regulators perturbing promoter methylation in a target gene-specific manner, and spreading of aberrant epigenetic states to neighboring promoters in the absence of CTCF binding sites. mechanism
  • Two quantitative parameters were derived to measure epigenetic heterogeneity: the M-score (intra-sample methylation heterogeneity, intermediate values near zero indicating mixed methylation within a sample) and the inter-quartile range of M-scores across samples (inter-sample heterogeneity). method
  • Increased intra-sample variation is an inherent feature of neoplastic transformation, not an artifact of copy number alteration, sample purity, probe signal-to-noise, proliferation/mitotic rate, or patient age. finding
Experimental setups
Assay System Perturbation Readout Platform
HELP assay (genome-wide DNA methylation profiling on microarray) Human primary cells: normal naive B-cells (NBC, n=8), normal germinal center B-cells (NGC, n=10), follicular lymphoma (FL, n=8), GCB-DLBCL (n=39), ABC-DLBCL (n=18) none Normalized array signal intensity per probeset (M-score) reflecting degree of CpG methylation; IQR of M-scores across samples Custom-designed NimbleGen microarrays with probesets representing >50,000 CpGs corresponding to regulatory regions of roughly 14,000 human genes
ERRBS (enhanced reduced representation bisulfite sequencing; base-pair resolution quantitative bisulfite sequencing) DLBCL patient samples (six DLBCL samples validated by orthogonal assays) none Base-pair resolution cytosine methylation levels; used to validate HELP methylation profiles, intra-sample heterogeneity, CpG-density relationships, and chromosomal-region patterns
MassARRAY quantitative bisulfite-based methylation assay DLBCL patient samples (among the six DLBCL samples used for orthogonal validation) none Quantitative CpG methylation levels validating HELP profiles and increasing intra-sample heterogeneity MassARRAY
Flow cytometry Normal B-cell (NBC, NGC) and lymphoma patient samples none Sample purity (percentage of target cell population)
SNP array genotyping (copy number analysis) Same lymphoma patients profiled for methylation none Copy number variations, used as a covariate/control for methylation differences
Gene expression profiling DLBCL patient tumors none Transcriptional signatures used to sub-classify DLBCL into GCB and ABC subtypes
HELP-based DNA methylation profiling of cell lines with known doubling times Lymphoma/cancer cell lines none (natural variation in proliferation rate) Relationship between mitotic rate/doubling time and DNA methylation heterogeneity
Computational analyses: phylogenetic clustering of methylation patterns, Kaplan-Meier and Cox proportional hazards survival modeling, neighboring-promoter (i±1 to i±5) methylation spreading analysis, CTCF binding site annotation, chromosomal-region and gene-density partitioning CD34+ bone marrow hematopoietic progenitor cells, NBC, NGC, FL, GCB and ABC DLBCL methylation datasets with clinical outcome and IPI data none Methylation distance/heterogeneity score, phylogenetic tree topology, hazard ratios and concordance index, ΔM-score of neighboring promoters, differentially methylated probeset counts
Key results
  • DNA methylation distributions in lymphoma samples differ significantly from normal B-cells, with an increased proportion of probes at intermediate M-scores (high intra-sample variation) rising progressively from FL to GCB to ABC DLBCL Kolmogorov-Smirnov test, FDR-corrected p<2.2×10^-16
  • Inter-sample variation (IQR of M-scores) is small in normal B-cell controls but progressively increases in FL, GCB, and ABC DLBCL Mann-Whitney test, FDR-corrected p<2.2×10^-16
  • Adding the methylation heterogeneity score to the IPI improved prognostic concordance and yielded significant risk stratification in combined GCB and ABC DLBCL samples Concordance 0.64 to 0.7 (ΔC 0.06; 95% CI -0.08-0.20); HR = 3.85, p<0.03
  • Phylogenetic clustering shows progressive departure of genome-wide methylation from bone marrow CD34+ progenitors to NBC and NGC, then FL, then DLBCL, correlating with disease severity
  • Centromeric regions are hyper-methylated in normal cells but show gradual loss of methylation in lymphomas, while intermediate chromosomal regions show increasing intra-sample variation with disease severity (NBC<NGC<FL<GCB<ABC) Kolmogorov-Smirnov test p<2.2×10^-16 for NBC-FL, NBC-GCB and NBC-ABC pairs (intermediate regions)
  • Gene-rich regions show increased intra-sample variation in lymphomas while gene-poor regions become progressively hypo-methylated; inter-sample variation increases in lymphoma subtypes in both gene-poor and gene-dense regions Mann-Whitney test: FL p<1×10^-3; GCB and ABC p<1×10^-10
  • 3,414 probesets were significantly hyper-methylated and 2,044 significantly hypo-methylated in ABC DLBCL versus normal germinal center B-cells 3,414 hyper- and 2,044 hypo-methylated probesets; FDR-corrected p<5.0×10^-3
  • Promoters neighboring an aberrantly methylated promoter show methylation change in the same direction, with the effect decaying from i±1 to i±5 and remaining significant at i±5 only for hypo-methylation; aberrantly hypo-methylated (but not hyper-methylated) promoters also show greater inter-sample variation in ABC lymphomas i±1 hypo-methylation p=4.56×10^-5; i±1 hyper-methylation p=3.11×10^-3; i±5 hypo-methylation p=3.01×10^-3; i±5 hyper-methylation p>0.05
Key statistics
  • pvalue FDR corrected p-value<2.2×10^-16 (Kolmogorov-Smirnov test comparing M-score distributions between pairs of normal and lymphoma samples (intra-sample variation))
  • pvalue FDR corrected p-value<2.2×10^-16 (Mann-Whitney test comparing IQR values between pairs of normal and lymphoma tissues (inter-sample variation))
  • other concordance improved from 0.64 to 0.7 (ΔC 0.06; 95% CI -0.08-0.20) (Cox model concordance for IPI alone vs. IPI plus methylation heterogeneity score in GCB and ABC DLBCL analyzed together)
  • other HR = 3.85, p<0.03 (Kaplan-Meier/Cox risk stratification of DLBCL patients into high- vs low-risk groups by median risk score using IPI plus methylation heterogeneity score)
  • count 3,414 hyper-methylated and 2,044 hypo-methylated probesets (FDR-corrected p-value<5.0×10^-3) (Differentially methylated promoter probesets in ABC DLBCL (n=18) vs NGC (n=10))
  • pvalue p-value: 4.56×10^-5 (hypo, i±1); 3.11×10^-3 (hyper, i±1); 3.01×10^-3 (hypo, i±5); >0.05 (hyper, i±5) (Neighboring-promoter concordant aberrant methylation (spreading) analysis in ABC vs NGC)
  • pvalue FL: p-value<1×10^-3; GCB and ABC: p-value<1×10^-10 (Mann Whitney test) (Increase in inter-sample variation in lymphoma subtypes vs normal cells in gene-poor and gene-dense regions)
  • count NBC 8, NGC 10, FL 8, GCB 39, ABC 18 samples; >50,000 CpGs covering roughly 14,000 genes; sample purity >90% (Cohort composition, array coverage, and flow-cytometry-confirmed purity of profiled samples)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study compared genome-wide DNA methylation profiles (HELP assay/microarray, with orthogonal validation by ERRBS and MassARRAY bisulfite sequencing) across normal B-cell subsets and B-cell lymphoma subtypes using biological samples (n=8–39 per group). Distributional differences in methylation scores (M-score) and their inter-sample spread (IQR) were assessed using the Kolmogorov-Smirnov test and the Mann-Whitney U test, with FDR correction applied across comparisons. Kaplan-Meier curves and multivariate Cox proportional-hazards models (incorporating the International Prognostic Index and a methylation heterogeneity score) were used to relate methylation patterning to patient survival, with the added prognostic value quantified via a concordance index (C-statistic) reported with a 95% confidence interval.

Replicationbiological Sample sizeGroup sample sizes stated explicitly (NBC=8, NGC=10, FL=8, GCB=39, ABC=18); no formal power/sample-size calculation described Groupsnormal B-cell subsets (NBC, NGC) vs. lymphoma subtypes (FL, GCB DLBCL, ABC DLBCL) Pairingunpaired Randomization/blindingnot stated DispersionIQR Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionFDR correction (specific procedure, e.g. Benjamini-Hochberg, not named)
Statistical tests used
Test Applied to n Assumptions
Kolmogorov-Smirnov test comparing M-score distributions between pairs of normal and lymphoma samples (Figure 1B) and by chromosomal region across disease severity (Figure 3B) NBC=8, NGC=10, FL=8, GCB=39, ABC=18 (as given in Figure 1A) not stated
Mann-Whitney U test comparing IQR (inter-sample variation) between pairs of normal and lymphoma tissues (Figure 1C) and between gene-poor vs. gene-dense regions (Figure 3C) same group sizes as above; not separately restated for regional analysis not stated
Kaplan-Meier analysis with multivariate Cox proportional-hazards model survival stratification using IPI and methylation heterogeneity score in GCB and ABC DLBCL samples (Figure 2B–2C) GCB and ABC samples analyzed together (39+18=57 as listed in Figure 1A) not stated
Concordance index (C-statistic) comparison comparing predictive concordance of IPI alone vs. IPI plus methylation heterogeneity score not stated
Differential methylation testing (FDR-corrected, method not further specified) identifying hyper- and hypo-methylated probesets in ABC vs. NGC (Figure 4A–4C) not stated
Approaches that could also have been used
  • Differences in methylation-score distributions between normal and lymphoma samples were assessed with the Kolmogorov-Smirnov test.
    Could also: A permutation test or a mixture-model-based comparison (e.g. modeling the bimodal vs. intermediate components directly) — These approaches can characterize which part of the distribution (e.g. the intermediate/heterogeneous component) is driving the difference, complementing the omnibus KS statistic.
  • Inter-sample variation (IQR) was compared between groups using the Mann-Whitney U test.
    Could also: Reporting an accompanying effect-size measure such as the Hodges-Lehmann estimator or rank-biserial correlation alongside the p-value — This would convey the magnitude of the difference in variation between groups, not just its statistical significance.
  • Multiple comparisons across probesets and sample pairs were corrected using an FDR method.
    Could also: A Bonferroni correction, or explicit reporting of q-values per test — Bonferroni offers a more conservative family-wise error control, while explicit q-value reporting is a common convention in large-scale genomic multiple-testing settings and can aid comparison across studies.
  • A Cox proportional-hazards model with IPI and methylation heterogeneity score was used to stratify patients and estimate a concordance-index improvement.
    Could also: A likelihood-ratio test comparing the nested Cox models (with vs. without the methylation score), or bootstrap validation of the concordance index — A likelihood-ratio test would directly assess whether adding the methylation score significantly improves model fit, and bootstrap resampling could provide a validated estimate of the concordance-index confidence interval, which here spans zero (−0.08–0.20).
  • Group sizes varied considerably across cell/tumor types (e.g. FL n=8 vs. GCB n=39).
    Could also: A formal power analysis or a resampling/bootstrap approach to characterize the stability of estimates for the smaller groups — This would help convey the precision of estimates from smaller groups relative to larger ones when comparing effect magnitudes across subtypes.
  • Spread of methylation heterogeneity was summarized using the inter-quartile range (IQR).
    Could also: Reporting the standard deviation or a bootstrap-based confidence interval for variability alongside the IQR — This can offer a complementary, parametric view of dispersion and facilitate comparison with studies that use SD-based variability metrics.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

C1_group_distance_topology
Reported
Group-level HELP M-score Pearson-distance shows CD34+ progenitors most distant from DLBCL subtypes, with GCB and ABC closer to each other than either is to CD34+ (paper's broader claim: progressive divergence CD34->NBC/NGC->FL->DLBCL).
Reproduced
{'pearson_distance_CD34_GCB': 0.07502135, 'pearson_distance_CD34_ABC': 0.11138332, 'pearson_distance_GCB_ABC': 0.02300536, 'group_consensus_tree_newick': '(CD34,GCB,ABC)1;'}
partial
C2_sample_clade_separation
Reported
Sample-level bootstrap phylogeny shows GCB and ABC DLBCL subtypes forming largely separate clades.
Reproduced
{'within_GCB_cophenetic_mean': 0.2236, 'within_ABC_cophenetic_mean': 0.165, 'between_cophenetic_mean': 0.1936}
partial
C3_subtype_distribution_difference
Reported
M-score distributions differ significantly between GCB and ABC DLBCL subtypes.
Reproduced
{'D': 0.0799, 'p_value': '~0 (below double-precision resolution)'}
partial
C4_heterogeneity_lymphoma_vs_normal
Reported
Paper's central claim: M-score inter-sample heterogeneity (IQR) is significantly higher in lymphoma than in normal (CD34+) cells, increasing with disease severity.
Reproduced
{'W': 1114399439, 'p_value': '0 (R output; below double-precision resolution, consistent with <2.2e-16)', 'median_IQR_lymphoma': 0.7552, 'median_IQR_CD34': 0.4021}
within tolerance
C5_dataset_sample_counts_vs_paper
Reported
GEO-deposited sample counts for the DLBCL HELP cohort (GSE23967) compared to the paper's reported group sizes used in its final statistics.
Reproduced
{'GCB': 40, 'ABC': 20, 'NC_not_classifiable': 9, 'total': 69}
did not match
C6_NBC_NGC_FL_scope
Reported
Full 6-group (NBC/NGC/FL/GCB/ABC vs CD34+) progressive-methylation-divergence trajectory as described in the paper.
Reproduced
N/A -- scope check only.
m.public.grade.error

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 49/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

The paper's headline finding reproduces cleanly: per-probe M-score IQR is far higher in DLBCL than in normal CD34+ progenitors (median 0.7552 vs 0.4021, Mann-Whitney p<2.2e-16), matching the paper's reported significance and direction. The binding limitation is on the authors' side: the NBC/NGC/FL HELP data was 'pending GEO accession' in 2013 and still does not exist, so the 'increases with disease severity' gradient (3 of 6 groups) was never testable, and the paper gives no numeric distances for its phylogeny claims. Smaller, well-documented deviations are ours (PHYLIP fconsense → ape::consensus, pooled KS instead of stratified Mann-Whitney, GCB=40/ABC=20 instead of 39/18 with no published exclusion list). Overall: an honest, competently executed partial reproduction whose coverage — not whose correctness — is compromised, hence yellow throughout rather than red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.