Detecting aberrant DNA methylation in Illumina DNA methylation arrays: a toolbox and recommendations for its use.
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce, and largely 1:1. Ran the authors' own R package OutlierMeth (the paper's deliverable; bundled reference tables shipped in the 92MB OutlierMeth_1.2.1.tar.gz on the master branch) on the paper's named dataset GSE40279 (656 whole-blood 450K), reference=all (TCGA-GEO), p=0.01 (paper default). The package's own flagMeth() was used (no re-implementation). Headline blood claim C1 reproduced at 2606 vs 2546 reported (2.4% off, fully explained by the documented removal of GSE127857 from the v1.2.1 reference vs the 2023 paper reference). Age-marker availability C2 reproduced exactly (56 markers). C3 (24.4 vs 28) and C4 (90.5 vs 92) within a few points. C5 (active-marker fraction) did NOT reproduce (62.5 vs 44.6, non-significant Fisher) because the paper does not give a precise operational definition of an 'active' probe and our marker source lists 69 vs 71 cg's. No fabrication: all checked numbers are derivable from the shipped package + public GSE40279. NOT attempted: Fig 3 tumor/normal outlier rates and Fig 4 pan-cancer classifier (need TCGA tumor cohorts + a trained model outside this RU's GSE40279 scope), Fig 1 qualitative threshold patterns (no pinnable number), and rebuilding the reference from >2000 TCGA-GEO normals (reference is shipped pre-computed and was used as-is).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 73assessed: 2026-06-15 ⛓ d46854b4be15
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15👤 1 human curator(s) · Level L2 2026-06-15
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper aims to determine probe-specific thresholds for identifying aberrant (outlier) DNA methylation on Illumina methylation arrays, and to evaluate whether outlier-based methylation data can substitute for continuous methylation data in common analyses.
- ★ Developed OutlierMeth, an R package that computes probe-specific thresholds and flags DNA methylation outliers from array data resource
- ★ Probe-specific outlier thresholds were derived from a reference database of >2,000 normal and tumor-adjacent normal samples spanning 25 tissue types on the Illumina 450K array method
- ★ Outlier calls perform as well as, or better than, continuous methylation data for simple tasks like tumor-vs-normal classification, but become less useful as analytical complexity increases (clustering, age prediction) finding
- ★ CpG Islands/Shores show lower thresholds near zero and upper thresholds near 0.20 in normal tissue, while Shelves/OpenSeas show upper thresholds near 1, reflecting known cancer-associated methylation patterns finding
- ★ Whole blood was excluded from the reference set because it has a distinctive methylation pattern; 2,546 probes were flagged as outliers in every blood sample, enriched for blood/immune-related genes finding
- Brain tissue samples were unusually influential on threshold calculation in both TCGA and GEO, likely due to cellular heterogeneity in brain finding
- ★ Tumor types show variable rates of outlier methylation (COADREAD highest, LUAD lowest), with some types (e.g., LUSC) showing a directional bias toward hypomethylation finding
- ★ Age-associated methylation markers are less likely than average probes to be flagged as outliers, since aging-related changes are gradual and subtle, though a weak association with age persists finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Illumina 450K DNA methylation array | normal and tumor-adjacent normal tissues (TCGA + GEO, 25 tissue types, >2,000 samples) | none | probe-level methylation beta-value distribution / thresholds | Illumina Human 450K methylation array |
| Illumina 450K DNA methylation array | whole blood samples | none | proportion of outlier probes per sample | Illumina Human 450K methylation array |
| Illumina 450K DNA methylation array | TCGA tumor tissues (BRCA, COADREAD, LUAD, LUSC, PRAD, THCA) | disease state (tumor vs normal) | proportion of samples with flagged outlier probes | Illumina Human 450K methylation array |
| Logistic regression with LASSO variable selection on methylation array data | TCGA pan-tumor (N=8,378) vs pan-normal (N=747) samples | none | classification accuracy (AUC) distinguishing tumor from normal | — |
| Unsupervised clustering / PCA of methylation array data | TCGA lung tissue (LUAD, LUSC, and normal adjacent) | none | cluster separation of tumor types and normal tissue | — |
| Unsupervised clustering / PCA of methylation array data | TCGA kidney tissue (KIRP, KIRC, and normal adjacent) | none | cluster separation of tumor subtypes and normal tissue | — |
| Illumina 450K DNA methylation array (Hannum age marker analysis) | GSE40279 whole blood samples (n=656, ages 19-101) | none | proportion of 56 age markers flagged as outliers per sample; active probe rate | Illumina Human 450K methylation array |
| Gene set analysis | genes associated with 2,546 commonly outlying blood probes | none | pathway/biology enrichment | — |
- ▲ Outlier-derived 45-probe panel classified pan-tumor vs pan-normal with higher accuracy than the continuous 194-probe panel AUC 0.978 (CI 0.971-0.984) vs AUC 0.950 (CI 0.936-0.964)
- – Reflagging the 194 continuous-selected probes as outliers caused only a minor accuracy change and all 194 probes remained active AUC 0.927 (CI 0.913-0.942)
- – COADREAD tumors had the highest proportion of samples exceeding the outlier threshold; LUAD had the lowest 95% (COADREAD) vs 69.2% (LUAD)
- – 2,546 probes were flagged as outliers in every single whole blood sample, linked to blood/immune biology genes 2,546 probes
- – Relative influence of tissue types on thresholds was not well correlated between GEO and TCGA Spearman rho = 0.248; P = 0.492
- – Continuous data outperformed outlier data for finer-grained clustering (tumor subtypes in kidney; tumor vs normal in lung), though differences were minimal overall
- – Most samples showed few or no outlier age markers despite known age-related methylation drift 28% of samples had 0/56 markers abnormal; 92% had <4 markers abnormal
- ▼ Age markers were significantly less likely to be 'active' (informative outliers) than non-age probes 44.6% vs 64.1% active; Fisher's exact P = 0.003
- other AUC 0.978, CI 0.971-0.984, Mann-Whitney P<0.001 (45-probe outlier panel classifying pan-tumor vs pan-normal)
- other AUC 0.950, CI 0.936-0.964, Mann-Whitney P<0.001 (194-probe continuous panel classifying pan-tumor vs pan-normal)
- other AUC 0.927, CI 0.913-0.942, Mann-Whitney P<0.001 (continuous-selected probes reflagged as outliers)
- correlation Spearman rho = 0.248, P = 0.492 (correlation of tissue-type outlier influence between GEO and TCGA)
- pvalue Fisher's exact P = 0.003 (difference in percent active probes between age markers (44.6%) and non-age markers (64.1%))
- count 2,546 probes flagged as outliers in every blood sample (whole blood outlier analysis)
- other 95% of COADREAD samples and 69.2% of LUAD samples exceeded the 2% outlier threshold (tumor-type outlier rate comparison)
- count >2,000 samples (2,015 per Discussion) across 25 normal tissue types (size of normal reference database from TCGA and GEO)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This methods-development paper constructs a large reference database (2,015 normal/tumour-adjacent tissue samples from TCGA and GEO) and derives probe-specific quantile thresholds for flagging aberrant DNA methylation on Illumina 450K arrays. The statistical work spans descriptive characterization of the reference distribution, supervised classification via logistic regression with LASSO penalization evaluated by ROC/AUC and Mann-Whitney tests, unsupervised clustering via PCA, a Spearman correlation between tissue-type influence rankings, and a Fisher's exact test comparing activity rates of age-associated vs. non-age probes. Results are reported with AUCs and 95% CIs for classification tasks and exact P-values for discrete comparisons, with no formal multiple-testing correction stated for the battery of inferential tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Quantile-based threshold derivation (R quantile function, 0.01–0.00001 significance levels) | Per-probe upper and lower outlier thresholds across the full reference database; applied separately to TCGA-only, GEO-only, combined, and blood collections | 2,015 normal/tumour-adjacent samples across 25 tissue types | not stated |
| Logistic regression with LASSO variable selection | Pan-tumour vs. pan-normal classification (Figure 4), applied separately to continuous methylation data, flagged-outlier data, and continuous-selected probes subsequently flagged | 8,378 tumour + 747 normal samples from TCGA (partitioned into training and test sets; split proportions not stated) | not stated |
| Mann-Whitney U test (used to compare AUC model performance) | Significance of all three ROC models distinguishing pan-tumour from pan-normal (Figure 4a); P < 0.001 for all three panels | 8,378 tumour + 747 normal | not stated |
| Principal component analysis (PCA; unsupervised) | Clustering of lung (LUAD, LUSC, normal-adjacent) and kidney (KIRP, KIRC, normal-adjacent) tissue datasets, using top 5% most variable probes (Figure 5) | Sample counts for lung and kidney subsets not explicitly stated | not stated |
| Spearman rank correlation | Comparison of tissue-type influence ranks between GEO and TCGA datasets (Figure 2c); rho = 0.248, P = 0.492 | ~25 tissue types (exact number of data points in correlation not stated) | na |
| Fisher's exact test | Comparison of proportion of 'active' probes among 56 age-associated markers vs. remaining non-age probes (Figure 6d); P = 0.003 | 56 age markers vs. remainder of 450K probes available in reference; sample base of 656 blood samples | not stated |
| Gene set analysis (method not specified) | Characterization of the 2,546 probes flagged as outliers in every blood sample | 2,546 probes | not stated |
-
Outlier thresholds were derived as empirical quantiles of the reference distribution at fixed significance levels (0.01–0.00001) using R's quantile function↳ Could also: Robust outlier-detection methods such as Median Absolute Deviation (MAD)-based Z-scores or the Tukey fence (1.5×IQR rule) could also define per-probe thresholds — MAD-based approaches are less sensitive to the very outliers they seek to detect and can be preferable when the reference distribution itself contains a small proportion of extreme values; reporting which method was chosen and why would help readers calibrate expectations about threshold conservatism
-
A single combined multi-tissue reference distribution was used for all tissue types rather than tissue-specific references↳ Could also: Tissue-stratified or leave-one-tissue-out reference distributions could also be constructed and evaluated — A stratified approach can reduce false positives driven by tissue-specific methylation patterns; the paper discusses this trade-off qualitatively but a quantitative comparison (e.g., sensitivity/specificity by approach) would allow readers to select the appropriate reference for their own tissue context
-
Classification performance for distinguishing tumour from normal was evaluated on the same TCGA data used to derive the reference thresholds (with a train/test partition, but no independent external cohort)↳ Could also: External validation on a held-out GEO cohort or cross-dataset cross-validation could also be used to assess generalizability — Even with an internal train/test split, thresholds derived from overlapping samples can optimistically bias performance estimates; external validation provides a stronger estimate of how the tool would perform on new data
-
PCA was used for unsupervised clustering of lung and kidney tissue datasets, and cluster quality was assessed qualitatively by visual inspection of PCA plots↳ Could also: Quantitative cluster quality metrics such as the adjusted Rand index, silhouette score, or cophenetic correlation from hierarchical clustering could also be reported — Numerical cluster-quality metrics allow objective comparison between continuous and outlier-flagged data and across tissue types, complementing the visual PCA summaries shown in Figure 5
-
Three separate Mann-Whitney tests were used to assess significance of each of the three ROC models (continuous, outlier, and continuous-then-flagged panels)↳ Could also: DeLong's method for comparing correlated AUC curves could also be used to directly test whether the AUC difference between models is statistically significant — Because the three models were evaluated on the same test samples, their AUCs are correlated; DeLong's test accounts for this correlation and provides a direct P-value for the pairwise model comparison rather than individual significance statements
-
A Fisher's exact test was used to compare the proportion of 'active' probes in age-associated markers vs. all remaining probes↳ Could also: A permutation test drawing random probe sets of the same size (n=56) from the full array could also quantify whether the observed activity deficit is larger than expected by chance — A permutation approach would account for any probe-level correlation structure (e.g., probes in the same genomic region tend to co-vary) that the independence assumption of Fisher's exact test does not capture
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
99 downstream papers · 1 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Genome-wide methylation profiles reveal quantitative... 2013 · 3,350 cites
- Identification of tissue-specific cell death using m... 2016 · 501 cites
- Improved precision of epigenetic clock estimates acr... 2019 · 373 cites
- Methylome-wide Analysis of Chronic HIV Infection Rev... 2016 · 248 cites
- Urine DNA methylation assay enables early detection... 2020 · 179 cites
- Improving cell mixture deconvolution by identifying... 2016 · 171 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37218167
Paper: Downs BM, Thursby SJ, Cope L. Detecting aberrant DNA methylation in Illumina DNA methylation arrays: a toolbox and recommendations for its use. Epigenetics 2023;18(1):2213874. PMID 37218167 / PMC10208159.
Code: https://github.com/bdowns4/OutlierMeth (R package OutlierMeth, GPL-2).
Functions: referenceMeth (build per-probe quantile thresholds from a normal-tissue
beta matrix), flagMeth (binarize: +1 hyper / −1 hypo / 0 in-range / NA not-in-ref,
vs a reference at significance p), deltMeth, relMeth. The full package WITH the
bundled reference threshold datasets (all=TCGA-GEO, tcga, geo, blood) ships
as OutlierMeth_1.2.1.tar.gz (92 MB) on the master branch.
Data: GSE40279 (Hannum et al. 2013) — 656 healthy whole-blood 450K samples.
Processed beta matrix is public: GSE40279_average_beta.txt.gz (GEO suppl).
Default analysis (paper, verbatim): "Unless otherwise specified, the reference
dataset used in this study was calculated from the TCGA-GEO dataset at the 0.01
significance level." → reference = all, p = 0.01.
In scope (pipeline-derived, attempted)
| # | Reported result | Pipeline | Inputs | Feasible? |
|---|---|---|---|---|
| C1 | "2,546 probes were flagged as outliers in every single blood sample" | flagMeth(beta, all, p=0.01) on GSE40279, count probes flagged ( |
flag | =1) in all 656 samples |
| C2 | "56 of the 71 age markers were available in our reference set" | intersect Hannum-71 CpG list with rownames(all) |
Hannum-71 list + all reference |
YES if marker list obtainable |
| C3 | "in 28% of samples, all 56 age markers fell into the normal range" | flagMeth on 56 markers, fraction of samples with all flags=0 | GSE40279 + all + marker list |
secondary |
| C4 | "92% of samples showed non-normal methylation in fewer than 4 markers" | per-sample count of | flag | =1 among 56 markers < 4 |
| C5 | "percent of active age markers … 44.6 and 64.1" (age vs other probes) | "active" = flagged in ≥1 sample; fraction among 56 markers vs other probes | same | secondary |
Out of scope (not attempted, and why)
- Fig 3 tumor/normal outlier rates (COADREAD >95%, LUAD 69.2%, LUSC bias): needs per-cancer TCGA tumor beta matrices (many GB, dbGaP-adjacent), not the single GSE40279 accession named for this RU — out of the named data scope.
- Fig 4 pan-cancer classifier (194-probe AUC 0.950 / 45-probe AUC 0.978): needs a trained classifier + tumor cohorts not shipped; modelling step, not in the package.
- Fig 1 threshold-by-genomic-context patterns: qualitative ("near zero", "near 0.20"); no pinnable number.
- Reference construction from >2,000 TCGA-GEO normals: the reference is shipped pre-computed in the tarball; rebuilding it needs raw TCGA+GEO normals (out of scope). We instead USE the shipped reference (the package's own deliverable), which is the faithful way to reproduce the downstream flagging numbers.
Caveat (possible version drift)
The shipped reference is package v1.2.1 (2025); the README notes geo/all "no
longer include dataset GSE127857". The paper's 2,546 was computed with the 2023
reference. A small deviation in the count is therefore expected and is itself an
honest reproducibility finding, not necessarily a discrepancy in method.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Run with the authors' own OutlierMeth_1.2.1 package on the paper's named GSE40279, the headline blood result (C1: 2606 vs 2546) and age-marker availability/normality (C2–C4) reproduce well, with the small deviations fully explained by a documented reference-version change. The one real failure is C5: the age-marker activity comparison flips from Fisher P=0.003 (44.6% vs 64.1%) to P=0.481 (62.5% vs 66.9%), driven by the paper never defining an 'active' probe — an authors-side specification gap, so those exact values are not derivable from the shared data. No fabrication signal: all checked numbers are regenerable except where the method is underspecified. Overall a solid, largely 1:1 reproduction with one explainable but substantive significance flip on a secondary claim → yellow.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.