Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

pwrEWAS: a user-friendly tool for comprehensive power estimation for epigenome wide association studies (EWAS).

BMC Bioinformatics · 2019
L1 88/100 PQI 94
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
88/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 74% of all assessed papers rank 276 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1. pwrEWAS is a Bioconductor power-estimation tool; reproduced by running the published pwrEWAS() function (bioconda build 1.14.0 + pwrEWAS.data 1.14.0) on the paper's exact, fully-specified e-cigarette example (Blood adult ref, N=20-260 step40, J=100000, 2500 target DM CpGs, targetDelta=0.02/0.10/0.15/0.20, limma, BH-FDR=0.05, sims=50) on «our HPC» SLURM («job», 10 min, 16 cpus). The three headline numbers of Fig.2 -- subjects needed for 80% power to detect 10/15/20% methylation differences = 220/180/140 -- reproduced EXACTLY (marginal power crosses 0.80 at precisely those grid sample sizes). The secondary '~36%' sentence (P of detecting at least one DM CpG at dBeta=0.02, N=20) reproduced as 50%: this maps to the pwrEWAS 'probTP' metric (P(>=1 true positive)), NOT the return field misleadingly named 'classicalPower'; the value sits right at the detection threshold and is averaged over only 50 sims, so 0.50 vs 0.36 is Monte-Carlo/seed noise (same metric, same conclusion) -> graded partial, not mismatch. Independent supporting check: empirical FDR = 0.047-0.050 at dBeta>=0.10, exactly at the target BH 0.05. No fabrication concern: every reported value is derivable from the shipped tool+data. NOT attempted (per 80/20): Fig.3 (variance vs #sims, a tuning experiment), Table 3 (hardware-dependent runtime benchmark), and tool-vs-tool comparisons in the Discussion (external). GSE92767 is the bundled saliva reference dataset and is not exercised by this blood-tissue example.

💻 Code ↗ 🗄 Data: GSE92767

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 88
    assessed: 2026-06-14 ⛓ 0fa9ec91ee71
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

There is an outstanding need for a user-friendly, publicly available tool to comprehensively estimate statistical power for epigenome-wide association studies (EWAS); pwrEWAS addresses this by using a semi-parametric simulation approach to estimate power for two-group DNA methylation comparisons.

Core claims
  • pwrEWAS is a user-friendly tool (R package and Shiny web interface) for comprehensive power estimation in two-group EWAS using Illumina HumanMethylation BeadChip technology. resource
  • Power is estimated via a semi-parametric simulation approach that generates DNAm beta-values using CpG-specific means and variances estimated from curated tissue-type-specific reference datasets. method
  • Effect sizes (Δβ) are drawn from a truncated normal distribution N(0,τ²), reflecting that CpG-specific methylation differences come from a continuous rather than fixed discrete distribution. method
  • pwrEWAS reports marginal power, marginal type I error rate, marginal FDR, false discovery cost (FDC), the distribution of simulated Δβ, and the probability of identifying at least one true positive. method
  • pwrEWAS comprises three major steps: data generation, differential methylation analysis, and power evaluation, and allows users to select from several tissue types and statistical methods. method
  • Existing power-evaluation methods for EWAS rely on a limited number of single-locus distributions that may yield unrealistic data and lacked accompanying software, a limitation pwrEWAS overcomes. finding
  • pwrEWAS allows up- or down-scaling to any number of CpGs (default P=100,000; e.g. P=866,836 for the EPIC array) that the investigator plans to measure. method
Experimental setups
Assay System Perturbation Readout Platform
Illumina Infinium HumanMethylation450 BeadChip DNAm profiling (reference datasets) Saliva (GSE92767) none CpG-specific methylation beta-values; estimated means and variances Illumina Infinium HumanMethylation450
Illumina Infinium HumanMethylation450 BeadChip DNAm profiling (reference datasets) Whole blood — Adults (GSE42861), Children (GSE83334), Newborns (GSE82273) none CpG-specific methylation beta-values; estimated means and variances Illumina Infinium HumanMethylation450
Illumina Infinium HumanMethylation450 BeadChip DNAm profiling (reference datasets) Cord-blood whole blood (GSE69176) and cord-blood PBMC (GSE110128) none CpG-specific methylation beta-values; estimated means and variances Illumina Infinium HumanMethylation450
Illumina Infinium HumanMethylation450 BeadChip DNAm profiling (reference datasets) Adult PBMC (GSE67170) none CpG-specific methylation beta-values; estimated means and variances Illumina Infinium HumanMethylation450
Illumina Infinium HumanMethylation450 BeadChip DNAm profiling (reference datasets) Placenta (GSE62733) none CpG-specific methylation beta-values; estimated means and variances Illumina Infinium HumanMethylation450
Illumina Infinium HumanMethylation450 BeadChip DNAm profiling (reference datasets) Liver (GSE61258) none CpG-specific methylation beta-values; estimated means and variances Illumina Infinium HumanMethylation450
Illumina Infinium HumanMethylation450 BeadChip DNAm profiling (reference datasets) Colon (GSE77718) and Lymphoma (GSE42372) none CpG-specific methylation beta-values; estimated means and variances Illumina Infinium HumanMethylation450
In silico simulation / power estimation (semi-parametric, beta-distribution) Simulated two-group EWAS DNAm data derived from tissue-specific reference datasets imposed Δβ on K of P CpGs (truncated normal N(0,τ²)) marginal power, type I error, FDR, FDC, probability of ≥1 true positive
Key results
  • pwrEWAS estimates statistical power as a function of assumed sample size and effect size(s) for two-group DNAm comparisons.
  • Marginal power is calculated as the proportion of true positives among all truly differentially methylated CpGs.
  • A CpG is considered truly differentially methylated if its absolute difference in mean methylation exceeds the default detection limit (0.01), and 'detected' if its FDR is below the default threshold (0.05).
Key statistics
  • count > 450,000 CpGs (HumanMethylation450) and > 850,000 CpGs (EPIC) (CpG dinucleotides interrogated by the two Illumina arrays)
  • count P = 100,000 (default); P = 866,836 for EPIC array (number of CpGs sampled/simulated)
  • other detection limit default 0.01 (±0.005) (threshold defining truly differentially methylated CpGs)
  • other FDR threshold default 0.05 (threshold for a CpG to be 'detected')
  • other 99.99th percentile of |Δβ,k| used to tune τ (calibrating truncated normal SD to target maximal methylation difference)
  • other differences in methylation ranged 1.25–14.4% (Rakyan et al.), 1–60% (Tsai et al.) (effect-size ranges in prior EWAS power studies cited)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a software/methods paper presenting pwrEWAS, an R/Shiny tool for statistical power estimation in two-group epigenome-wide association studies (EWAS) of DNA methylation. Rather than reporting an experimental study with hypothesis tests, it describes a semi-parametric, simulation-based framework in which DNAm beta-values are generated from beta-distributions using CpG-specific means and variances drawn from curated tissue-specific reference data sets, with differences (Δβ) imposed on a subset of CpGs from a truncated normal distribution. Simulated data sets are then subjected to differential methylation analysis and power is summarized as marginal power, marginal type I error rate, marginal FDR, and false discovery cost (FDC).

Replicationunclear Sample sizeSample size is a user-supplied input parameter (can be specified as a single value or a range); the tool's purpose is to help determine required sample size rather than reporting a study's own n. Effect size (Δβ), number of CpGs (default P=100,000), expected number of differentially methylated CpGs (K), target FDR (default 0.05), and number of simulated data sets are also user-specified. GroupsTwo groups (e.g. case vs control, exposed vs unexposed) Pairingunpaired Randomization/blindingna Dispersionunclear Effect sizesyes Multiplicity correctionFalse discovery rate (FDR) control; a target FDR threshold (default 0.05) is used to define 'detected' CpGs. Specific FDR procedure not named in this text.
Statistical tests used
Test Applied to n Assumptions
User-selectable differential methylation analysis methods (the tool offers a choice among several statistical methods for the two-group comparison; specific methods not enumerated in this text) Step 2 'differential methylation analysis' within each simulated data set, comparing mean methylation between the two comparator groups per CpG User-specified total sample size split into two groups (N1 and N2); no fixed n stated stated
t-tests (referenced as methods used in prior EWAS power work, e.g. Wang et al., Tsai et al.) Background discussion of related approaches, not pwrEWAS's own analysis na
Wilcoxon rank-sum test (referenced from prior work, Tsai et al.) Background discussion of related approaches na
Logistic regression (referenced from prior work, Rakyan et al.) Background discussion of related approaches na
Approaches that could also have been used
  • Methylation β-values are simulated from beta-distributions using CpG-specific means and variances estimated from reference data sets.
    Could also: Simulation on the M-value (logit-transformed) scale using normal/Gaussian models, as is common in limma-based EWAS pipelines. — M-value modeling can stabilize variance and aligns with linear-model frameworks frequently used in practice; offering both scales would let users match the simulation to their planned analysis pipeline.
  • Imposed differences in mean methylation (Δβ) are drawn from a truncated normal distribution, justified by previously observed EWAS differences.
    Could also: Alternative effect-size distributions (e.g. mixtures, empirical/bootstrapped distributions of observed Δβ, or fixed discrete effect sizes). — Different distributional choices capture different biological scenarios; an empirical or mixture distribution could more directly reflect a specific study's anticipated effects.
  • Power is summarized as marginal (average) power, type I error, and FDR across simulated CpGs.
    Could also: Reporting Monte Carlo uncertainty for these estimates (e.g. standard errors or 95% confidence intervals across the simulated data sets). — Quantifying simulation variability would convey how precisely power is estimated for a given number of simulated data sets and help users choose an adequate number of simulations.
  • The tool focuses on two-group comparisons of DNA methylation.
    Could also: Extension to continuous exposures, multi-group, or covariate-adjusted/regression designs with confounder control. — Many EWAS involve continuous phenotypes or require adjustment for cell composition and batch; supporting regression-based designs would broaden applicability to common study settings.
  • Comparator groups are assumed to share identical CpG-specific variances, differing only in mean for differentially methylated CpGs.
    Could also: Allowing group-specific (heteroscedastic) variances in the simulation. — Permitting unequal variances between groups could better represent settings where disease or exposure alters methylation variability, not just the mean.
  • Multiplicity is handled via FDR control with a user-specified target threshold.
    Could also: Offering family-wise error rate control (e.g. Bonferroni) or alternative FDR estimators as selectable options. — Some confirmatory EWAS designs favor stricter FWER control; providing multiple correction options would let power estimates match the error-control strategy planned for the final analysis.
Software: R statistical programming language · Shiny (RStudio Inc.) web interface 2016 · pwrEWAS (R/Bioconductor package, available on GitHub)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
78
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE42861 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
GSE110128 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE42372 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE61258 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE62733 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE67170 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE69176 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE77718 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE82273 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE83334 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE92767 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31035919 (pwrEWAS)

Paper: pwrEWAS: a user-friendly tool for comprehensive power estimation for EWAS. BMC Bioinformatics 2019. PMID 31035919 · PMCID PMC6489300 · DOI 10.1186/s12859-019-2804-7. Code: https://github.com/stefangraw/pwrEWAS (Bioconductor pkg pwrEWAS + data pkg pwrEWAS.data). Data: GSE92767 is one of 12 bundled reference datasets (saliva); reference data ships inside the pwrEWAS.data Bioconductor package — no separate GEO download needed.

What pwrEWAS does

Semi-parametric, simulation-based power estimation for two-group EWAS. CpG-specific means/variances are estimated from a chosen tissue reference dataset; beta-distributed DNAm is simulated; a target number of DM CpGs gets a delta drawn from a truncated normal; differential methylation is detected (default limma) with BH-FDR control. Reports marginal power = TP/(TP+FN) averaged over sims datasets, plus classical power, marginal type-I error, empirical FDR, FDC.

IN SCOPE (pipeline-derived, reproduced)

The single fully-specified worked example ("e-cigarette study", Results / Fig 2):

  • tissueType = "Blood adult"; total sample size 20–260 step 40; NcntPer=0.5
  • J = 100000 CpGs; targetDmCpGs = 2500; targetDelta = c(0.02, 0.10, 0.15, 0.20)
  • detectionLimit = 0.01; DMmethod = "limma"; FDRcritVal = 0.05; sims = 50
  • Pipeline = the pwrEWAS() function itself (the published tool).
  • Claims C1–C4 (see claims.tsv): sample size at ~80% marginal power for Δβ=0.10/0.15/0.20, and classical power at n=20 for Δβ=0.02.

OUT OF SCOPE / not attempted (with reason)

  • Fig 3 (variance vs # simulated datasets, 5–100 sims × 100 repeats): a tuning/ justification experiment, not a headline result; cheap-ish but 100× repeats — skip per 80/20.
  • Table 3 (runtime benchmarking on 6 threads): hardware-dependent wall-clock, not a scientific result; not comparable across machines → out of scope.
  • Comparisons to other power tools in the discussion: external, not pipeline output.
  • Wet-lab / dataset curation (Table 1 selection): manual, out of scope.

Reproducibility note

Simulation is stochastic (truncated-normal deltas, beta sampling). A fixed set.seed before the call makes a single run reproducible, but the paper's exact seed is unknown, so grading is within-tol (do the power curves land on the same grid points / ~0.8 crossings, and is classical power ~0.36). Not byte-exact by construction.

Figures / tables: Fig 2
C1
Reported
~220 subjects for 80% power, dBeta=0.10
Reproduced
220 (marginal power 0.806)
exact
C2
Reported
~180 subjects for 80% power, dBeta=0.15
Reproduced
180 (marginal power 0.833)
exact
C3
Reported
~140 subjects for 80% power, dBeta=0.20
Reproduced
140 (marginal power 0.844)
exact
C4
Reported
~36% prob. of detecting >=1 of 2500 DM CpGs, dBeta=0.02, N=20
Reproduced
50% (probTP = 25/50 sims)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 88/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

The headline Fig.2 numbers (220/180/140 subjects for 80% power at Δβ=0.10/0.15/0.20) reproduced exactly by re-running the published pwrEWAS() function with the paper's fully-specified e-cigarette parameters on bundled reference data. The only deviation is the secondary ~36%→50% probability (P(≥1 true positive) at Δβ=0.02≈detectionLimit, N=20), a high-variance binary quantity averaged over 50 sims with the paper's seed unknown — pure Monte-Carlo/seed noise on our stochastic-simulation side, same metric and same scientific conclusion. No authors'-side defect and no fabrication concern: every reported value is derivable from the shipped tool+data, and independent FDR control (0.047–0.050 at target 0.05) corroborates the pipeline.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

126.7 k
tokens (I/O) · 9.5 M incl. cache
24 min
runtime · 1.43 CPU-h
32.3 GB
peak RAM
1
HPC jobs
hummel
machine