Detecting aberrant DNA methylation in Illumina DNA methylation arrays: a toolbox and recommendations for its use.
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce, and largely 1:1. Ran the authors' own R package OutlierMeth (the paper's deliverable; bundled reference tables shipped in the 92MB OutlierMeth_1.2.1.tar.gz on the master branch) on the paper's named dataset GSE40279 (656 whole-blood 450K), reference=all (TCGA-GEO), p=0.01 (paper default). The package's own flagMeth() was used (no re-implementation). Headline blood claim C1 reproduced at 2606 vs 2546 reported (2.4% off, fully explained by the documented removal of GSE127857 from the v1.2.1 reference vs the 2023 paper reference). Age-marker availability C2 reproduced exactly (56 markers). C3 (24.4 vs 28) and C4 (90.5 vs 92) within a few points. C5 (active-marker fraction) did NOT reproduce (62.5 vs 44.6, non-significant Fisher) because the paper does not give a precise operational definition of an 'active' probe and our marker source lists 69 vs 71 cg's. No fabrication: all checked numbers are derivable from the shipped package + public GSE40279. NOT attempted: Fig 3 tumor/normal outlier rates and Fig 4 pan-cancer classifier (need TCGA tumor cohorts + a trained model outside this RU's GSE40279 scope), Fig 1 qualitative threshold patterns (no pinnable number), and rebuilding the reference from >2000 TCGA-GEO normals (reference is shipped pre-computed and was used as-is).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 73assessed: 2026-06-15 ⛓ d46854b4be15
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15👤 1 human curator(s) · Level L2 2026-06-15
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan probe-specific thresholds be derived to identify aberrant (outlier) DNA methylation on the Illumina 450K array from a reference of normal tissues, and how do outlier-based methylation calls compare to continuous methylation data across common analyses of increasing complexity?
- ★ Probe-specific upper and lower thresholds for flagging aberrant DNA methylation can be derived from a reference database of >2,000 normal and tumour-adjacent normal samples spanning 25 tissue types. resource
- ★ An R package, OutlierMeth, provides the thresholds and functions to flag DNA methylation outliers. resource
- ★ Outlier (threshold) methylation data is as effective as, or slightly better than, continuous data for simple tasks like distinguishing tumour from normal, but becomes less useful as problem complexity increases. finding
- ★ Threshold values reflect known genomic patterns: CpG Islands/TSS-proximal probes have lower thresholds near zero and upper thresholds near 0.20, while Shelves/OpenSeas probes have upper thresholds near 1. finding
- ★ Individual studies (especially GEO datasets) are a larger source of threshold variation than tissue type, with brain being a notable influential tissue exception. finding
- ★ Blood was excluded from the reference because it has distinctive methylation patterns, with 2,546 probes flagged as outliers in every blood sample. method
- DNA methylation age markers are relatively inactive (produce few outliers) and are significantly less active than non-age markers, so outliers are too crude to capture subtle age-related changes. finding
- Reference thresholds were computed using the R quantile function at significance levels between 0.01 and 0.00001 for four reference collections (TCGA, GEO, TCGA-GEO combined, blood). method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Illumina Human 450K DNA methylation array (reference database construction) | 2,015–2,000+ normal and tumour-adjacent normal human samples, 25 tissue types, from TCGA and GEO | none | probe-specific upper/lower methylation outlier thresholds (quantile-based) | Illumina Infinium HumanMethylation450 (450K) array; R quantile function |
| Illumina 450K methylation array (outlier characterization in blood) | whole blood samples | none | proportion of probes flagged as outliers per sample; commonly outlying probes | Illumina 450K |
| Illumina 450K methylation array (aberrant methylation in cancer) | TCGA tumour tissues: BRCA, COADREAD, LUAD, LUSC, PRAD, THCA vs normal tissues | tumour vs normal (disease state) | proportion and direction of flagged outliers per tumour type | Illumina 450K |
| Supervised classification (LASSO logistic regression) on 450K data | TCGA pan-tumour (N=8,378) vs pan-normal (N=747) | tumour vs normal | classification AUC using continuous vs flagged outlier probe panels | Illumina 450K; LASSO logistic regression in R |
| Unsupervised clustering (PCA) on 450K data | TCGA lung (LUAD, LUSC + normal adjacent) and kidney (KIRP, KIRC + normal adjacent) tissues | none | cluster separation of tumour subtypes/normal using top 5% SD probes (continuous vs outlier) | Illumina 450K; PCA, GEO reference used to flag outliers |
| Epigenetic age marker outlier analysis on 450K data | GSE40279, 656 healthy blood samples, ages 19–101 | none | proportion of Hannum age markers flagged as outliers; active vs inactive markers | Illumina 450K; Hannum 71-marker age clock |
| Gene set / pathway analysis | most commonly outlying probes in whole blood | none | enriched biological pathways (blood cell/immune biology) | — |
- ▲ Flagged outlier panel (45 probes) distinguished pan-tumour from pan-normal better than continuous panel (194 probes) AUC 0.978 (CI 0.971–0.984) vs 0.950 (CI 0.936–0.964)
- – Continuous probes retrained after flagging outliers showed only minor accuracy change, all 194 sites remained active AUC 0.927 (CI 0.913–0.942)
- ▲ 2,546 probes were flagged as outliers in every single blood sample, associated with blood/immune biology 2,546 probes
- – COADREAD had the highest proportion of outlier-bearing samples and LUAD the lowest among 6 tumour types >95% (COADREAD) vs 69.2% (LUAD) of samples exceeding normal level
- – Relative influence of tissue types was not well correlated between GEO and TCGA Spearman rho = 0.248; P = 0.492
- ▼ Age markers were significantly less active than non-age markers 44.6% vs 64.1% active; Fisher's exact P = 0.003
- – Most age markers produced few/no outliers; majority of samples showed non-normal methylation in fewer than 4 of 56 markers 28% of samples had all 56 markers normal; 92% had <4 abnormal markers
- – Continuous data clustered tumour subtypes and normal adjacent tissues better than outlier data, though difference was minimal
- other AUC: 0.978; CI: 0.971–0.984; Mann-Whitney P < 0.001 (flagged outlier panel (45 probes) pan-tumour vs pan-normal)
- other AUC: 0.950; CI: 0.936–0.964; Mann-Whitney P < 0.001 (continuous panel (194 probes) pan-tumour vs pan-normal)
- other AUC: 0.927; CI: 0.913–0.942; Mann-Whitney P < 0.001 (continuous probes flagged for outliers after selection)
- correlation Spearman rho = 0.248; P = 0.492 (tissue-type influence correlation between GEO and TCGA)
- pvalue Fisher's exact P = 0.003 (age markers (44.6% active) vs non-age markers (64.1% active))
- count 2,546 probes (probes flagged as outliers in every blood sample)
- count 8,378 tumour / 747 normal (TCGA pan-tumour vs pan-normal samples for classification)
- count 2,015 samples, 25 tissue types (normal and tumour-adjacent normal reference database from TCGA and GEO)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This methods-development paper constructs a large reference database (2,015 normal/tumour-adjacent tissue samples from TCGA and GEO) and derives probe-specific quantile thresholds for flagging aberrant DNA methylation on Illumina 450K arrays. The statistical work spans descriptive characterization of the reference distribution, supervised classification via logistic regression with LASSO penalization evaluated by ROC/AUC and Mann-Whitney tests, unsupervised clustering via PCA, a Spearman correlation between tissue-type influence rankings, and a Fisher's exact test comparing activity rates of age-associated vs. non-age probes. Results are reported with AUCs and 95% CIs for classification tasks and exact P-values for discrete comparisons, with no formal multiple-testing correction stated for the battery of inferential tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Quantile-based threshold derivation (R quantile function, 0.01–0.00001 significance levels) | Per-probe upper and lower outlier thresholds across the full reference database; applied separately to TCGA-only, GEO-only, combined, and blood collections | 2,015 normal/tumour-adjacent samples across 25 tissue types | not stated |
| Logistic regression with LASSO variable selection | Pan-tumour vs. pan-normal classification (Figure 4), applied separately to continuous methylation data, flagged-outlier data, and continuous-selected probes subsequently flagged | 8,378 tumour + 747 normal samples from TCGA (partitioned into training and test sets; split proportions not stated) | not stated |
| Mann-Whitney U test (used to compare AUC model performance) | Significance of all three ROC models distinguishing pan-tumour from pan-normal (Figure 4a); P < 0.001 for all three panels | 8,378 tumour + 747 normal | not stated |
| Principal component analysis (PCA; unsupervised) | Clustering of lung (LUAD, LUSC, normal-adjacent) and kidney (KIRP, KIRC, normal-adjacent) tissue datasets, using top 5% most variable probes (Figure 5) | Sample counts for lung and kidney subsets not explicitly stated | not stated |
| Spearman rank correlation | Comparison of tissue-type influence ranks between GEO and TCGA datasets (Figure 2c); rho = 0.248, P = 0.492 | ~25 tissue types (exact number of data points in correlation not stated) | na |
| Fisher's exact test | Comparison of proportion of 'active' probes among 56 age-associated markers vs. remaining non-age probes (Figure 6d); P = 0.003 | 56 age markers vs. remainder of 450K probes available in reference; sample base of 656 blood samples | not stated |
| Gene set analysis (method not specified) | Characterization of the 2,546 probes flagged as outliers in every blood sample | 2,546 probes | not stated |
-
Outlier thresholds were derived as empirical quantiles of the reference distribution at fixed significance levels (0.01–0.00001) using R's quantile function↳ Could also: Robust outlier-detection methods such as Median Absolute Deviation (MAD)-based Z-scores or the Tukey fence (1.5×IQR rule) could also define per-probe thresholds — MAD-based approaches are less sensitive to the very outliers they seek to detect and can be preferable when the reference distribution itself contains a small proportion of extreme values; reporting which method was chosen and why would help readers calibrate expectations about threshold conservatism
-
A single combined multi-tissue reference distribution was used for all tissue types rather than tissue-specific references↳ Could also: Tissue-stratified or leave-one-tissue-out reference distributions could also be constructed and evaluated — A stratified approach can reduce false positives driven by tissue-specific methylation patterns; the paper discusses this trade-off qualitatively but a quantitative comparison (e.g., sensitivity/specificity by approach) would allow readers to select the appropriate reference for their own tissue context
-
Classification performance for distinguishing tumour from normal was evaluated on the same TCGA data used to derive the reference thresholds (with a train/test partition, but no independent external cohort)↳ Could also: External validation on a held-out GEO cohort or cross-dataset cross-validation could also be used to assess generalizability — Even with an internal train/test split, thresholds derived from overlapping samples can optimistically bias performance estimates; external validation provides a stronger estimate of how the tool would perform on new data
-
PCA was used for unsupervised clustering of lung and kidney tissue datasets, and cluster quality was assessed qualitatively by visual inspection of PCA plots↳ Could also: Quantitative cluster quality metrics such as the adjusted Rand index, silhouette score, or cophenetic correlation from hierarchical clustering could also be reported — Numerical cluster-quality metrics allow objective comparison between continuous and outlier-flagged data and across tissue types, complementing the visual PCA summaries shown in Figure 5
-
Three separate Mann-Whitney tests were used to assess significance of each of the three ROC models (continuous, outlier, and continuous-then-flagged panels)↳ Could also: DeLong's method for comparing correlated AUC curves could also be used to directly test whether the AUC difference between models is statistically significant — Because the three models were evaluated on the same test samples, their AUCs are correlated; DeLong's test accounts for this correlation and provides a direct P-value for the pairwise model comparison rather than individual significance statements
-
A Fisher's exact test was used to compare the proportion of 'active' probes in age-associated markers vs. all remaining probes↳ Could also: A permutation test drawing random probe sets of the same size (n=56) from the full array could also quantify whether the observed activity deficit is larger than expected by chance — A permutation approach would account for any probe-level correlation structure (e.g., probes in the same genomic region tend to co-vary) that the independence assumption of Fisher's exact test does not capture
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
99 downstream papers · 1 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Genome-wide methylation profiles reveal quantitative... 2013 · 3,350 cites
- Identification of tissue-specific cell death using m... 2016 · 501 cites
- Improved precision of epigenetic clock estimates acr... 2019 · 373 cites
- Methylome-wide Analysis of Chronic HIV Infection Rev... 2016 · 248 cites
- Urine DNA methylation assay enables early detection... 2020 · 179 cites
- Improving cell mixture deconvolution by identifying... 2016 · 171 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37218167
Paper: Downs BM, Thursby SJ, Cope L. Detecting aberrant DNA methylation in Illumina DNA methylation arrays: a toolbox and recommendations for its use. Epigenetics 2023;18(1):2213874. PMID 37218167 / PMC10208159.
Code: https://github.com/bdowns4/OutlierMeth (R package OutlierMeth, GPL-2).
Functions: referenceMeth (build per-probe quantile thresholds from a normal-tissue
beta matrix), flagMeth (binarize: +1 hyper / −1 hypo / 0 in-range / NA not-in-ref,
vs a reference at significance p), deltMeth, relMeth. The full package WITH the
bundled reference threshold datasets (all=TCGA-GEO, tcga, geo, blood) ships
as OutlierMeth_1.2.1.tar.gz (92 MB) on the master branch.
Data: GSE40279 (Hannum et al. 2013) — 656 healthy whole-blood 450K samples.
Processed beta matrix is public: GSE40279_average_beta.txt.gz (GEO suppl).
Default analysis (paper, verbatim): "Unless otherwise specified, the reference
dataset used in this study was calculated from the TCGA-GEO dataset at the 0.01
significance level." → reference = all, p = 0.01.
In scope (pipeline-derived, attempted)
| # | Reported result | Pipeline | Inputs | Feasible? |
|---|---|---|---|---|
| C1 | "2,546 probes were flagged as outliers in every single blood sample" | flagMeth(beta, all, p=0.01) on GSE40279, count probes flagged ( |
flag | =1) in all 656 samples |
| C2 | "56 of the 71 age markers were available in our reference set" | intersect Hannum-71 CpG list with rownames(all) |
Hannum-71 list + all reference |
YES if marker list obtainable |
| C3 | "in 28% of samples, all 56 age markers fell into the normal range" | flagMeth on 56 markers, fraction of samples with all flags=0 | GSE40279 + all + marker list |
secondary |
| C4 | "92% of samples showed non-normal methylation in fewer than 4 markers" | per-sample count of | flag | =1 among 56 markers < 4 |
| C5 | "percent of active age markers … 44.6 and 64.1" (age vs other probes) | "active" = flagged in ≥1 sample; fraction among 56 markers vs other probes | same | secondary |
Out of scope (not attempted, and why)
- Fig 3 tumor/normal outlier rates (COADREAD >95%, LUAD 69.2%, LUSC bias): needs per-cancer TCGA tumor beta matrices (many GB, dbGaP-adjacent), not the single GSE40279 accession named for this RU — out of the named data scope.
- Fig 4 pan-cancer classifier (194-probe AUC 0.950 / 45-probe AUC 0.978): needs a trained classifier + tumor cohorts not shipped; modelling step, not in the package.
- Fig 1 threshold-by-genomic-context patterns: qualitative ("near zero", "near 0.20"); no pinnable number.
- Reference construction from >2,000 TCGA-GEO normals: the reference is shipped pre-computed in the tarball; rebuilding it needs raw TCGA+GEO normals (out of scope). We instead USE the shipped reference (the package's own deliverable), which is the faithful way to reproduce the downstream flagging numbers.
Caveat (possible version drift)
The shipped reference is package v1.2.1 (2025); the README notes geo/all "no
longer include dataset GSE127857". The paper's 2,546 was computed with the 2023
reference. A small deviation in the count is therefore expected and is itself an
honest reproducibility finding, not necessarily a discrepancy in method.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Run with the authors' own OutlierMeth_1.2.1 package on the paper's named GSE40279, the headline blood result (C1: 2606 vs 2546) and age-marker availability/normality (C2–C4) reproduce well, with the small deviations fully explained by a documented reference-version change. The one real failure is C5: the age-marker activity comparison flips from Fisher P=0.003 (44.6% vs 64.1%) to P=0.481 (62.5% vs 66.9%), driven by the paper never defining an 'active' probe — an authors-side specification gap, so those exact values are not derivable from the shared data. No fabrication signal: all checked numbers are regenerable except where the method is underspecified. Overall a solid, largely 1:1 reproduction with one explainable but substantive significance flip on a secondary claim → yellow.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.