Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Predicting enhancers in mammalian genomes using supervised hidden Markov models.

BMC Bioinformatics · 2019
L1 70/100 PQI 92
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
70/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 37% of all assessed papers rank 732 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH: YES. ehmm is the authors' own R package shipping a self-contained, fully-specified genome-wide demo (pretrained mESC supervised HMM + quantile normalization -> enhancer/promoter prediction = the paper's Fig-2b cross-sample ESC->liver scenario). This room was a REQUEUE: the prior run built+launched but was cut off mid-decode. On resume I found (a) the «infra» work dir reclaimed by the janitor and (b) the «user» «infra» quota EXCEEDED (blocking ALL rooms; FS itself only ~18% full). Operator approved reclaiming ~17 GB of regenerable conda caches, which I deleted via an srun on a compute node (never front1). I then rebuilt the pipeline clean (clone @ f5f108086e50f61bcdac430ee2ad5d611d2141d2, R4.1.3 conda env to dodge the R>=4.2 USE_FC_LEN_T BLAS break, compiled ehmm 1.0) and ran the shipped genome-wide demo to COMPLETION on the correct ENCODE liver data (verified marks) -> 4591 enhancers / 5331 promoters (applyModel 516.9 s; «job» elapsed 00:59:44, DONE_OK). The reproduced liver counts are the cross-sample regime, NOT bit-comparable to the paper's within-sample ESC 5357/8040 (different cell type) but the same algorithm and order of magnitude. Datasets profiled (GSE120376: 26 samples, all ehmm input marks present, delivers-promised yes, grade A; ENCODE demo BAMs grade B). NOT ATTEMPTED (the hard ~20%): exact ESC integers (need GSE120376 raw reads -> BWA mm10) and Fig-2a CV AUPRC + REPTILE/ChromHMM/EpiCSeg benchmarks (no evaluate/CV subprogram shipped). No fabrication flags. All grades are provisional pending human review.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-14 ⛓ 90e0527bba2b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Enhancers and promoters share a common physical topology—a central nucleosome-free accessible DNA stretch flanked by nucleosomes with distinct histone modification patterns—and encoding this structural prior into a supervised hidden Markov model can predict active enhancers genome-wide with higher precision and spatial resolution than existing unsupervised or 'black box' supervised methods.

Core claims
  • eHMM predicts enhancers with high precision and recall comparable to state-of-the-art methods and consistently outperforms them in accuracy and resolution finding
  • Enhancers and promoters are modeled as a central accessible DNA stretch flanked by two nucleosomes, imposed via constrained HMM state transitions mechanism
  • eHMM requires only a minimal set of four features (ATAC-seq plus ChIP-seq for H3K27ac, H3K4me1, H3K4me3) to predict enhancers method
  • eHMM's parameters are easy to interpret compared to other 'black box' enhancer prediction methods finding
  • eHMM predictions are spatially far more precise than REPTILE, with prediction centers much closer to ATAC-seq accessibility peaks finding
  • eHMM produces far fewer false enhancer predictions overlapping annotated TSS/promoters than REPTILE finding
  • Quantile normalization improves cross-sample prediction accuracy when applying a pre-trained model to a different sample finding
  • eHMM can be used as a stand-alone, pre-trained tool for enhancer prediction without additional training or parameter tuning resource
Experimental setups
Assay System Perturbation Readout Platform
ATAC-seq mouse ESC E14, liver E12.5, lung E16.5 none chromatin accessibility signal used as central feature of eHMM
ChIP-seq (H3K27ac, H3K4me1, H3K4me3) mouse ESC E14, liver E12.5, lung E16.5 none histone modification signal defining flanking nucleosome states
CAGE (FANTOM5) mouse ESC, liver, lung tissue-stages none bidirectional transcription initiation used to define enhancer training/validation sets orthogonal to histone ChIP-seq
5-fold cross-validation and cross-sample validation of eHMM predictions mouse ESC E14, liver E12.5, lung E16.5 none area under precision-recall curve (AUPRC) against FANTOM5 enhancers
Benchmark comparison against ChromHMM, EpiCSeg, and REPTILE mouse ESC E14, liver E12.5/E14.5, lung E16.5/E14.5; validated on FANTOM5 and EnhancerAtlas regions none AUPRC comparison across methods and state numbers
Transcription factor and chromatin remodeler ChIP-seq (Nanog, Oct4, Sox2, CTCF, p300) mouse ESC none binding enrichment at predicted enhancers vs promoters
MeDIP-seq mouse ESC none DNA methylation levels at predicted enhancers and promoters
phastCons sequence conservation analysis mouse ESC (cross-species comparison) none sequence conservation score at predicted enhancers and promoters
Key results
  • eHMM within-sample (5-fold CV) precision-recall performance on FANTOM5 enhancers across ESC, liver, and lung AUPRC 0.947-0.971
  • eHMM cross-sample validation (ESC-trained model applied to liver and lung) still performs well though slightly lower than within-sample AUPRC 0.928 (liver E12.5), 0.865 (lung E16.5)
  • Quantile normalization improves cross-sample prediction quality relative to no normalization AUPRC increase of 0.041 (liver), 0.025 (lung)
  • eHMM and REPTILE (supervised methods) outperform ChromHMM and EpiCSeg (unsupervised) within and across cell types
  • eHMM-predicted enhancer centers are much closer to ATAC-seq accessibility peaks than REPTILE predictions median distance 42 bp (eHMM) vs 343 bp (REPTILE), ~8-fold closer
  • eHMM predicted enhancers overlap annotated TSS far less often than REPTILE predicted enhancers 3.2% (eHMM) vs 17.8%-35.0% (REPTILE, threshold-dependent)
  • Genome-wide eHMM prediction in mouse ESC yields a defined number of enhancers and promoters without threshold selection 5357 enhancers, 8040 promoters
  • Run time comparison shows REPTILE uses least total CPU time but longest real time; eHMM has lowest total real time among compared methods eHMM 46.597 s real / 171.157 s CPU; REPTILE 90.917 s real / 145.550 s CPU
Key statistics
  • other AUPRC 0.947-0.971 (eHMM within-sample 5-fold CV on FANTOM5 in mouse ESC, liver, lung)
  • other AUPRC 0.928 (liver), 0.865 (lung) (eHMM cross-sample validation using ESC-trained model)
  • other AUPRC increase of 0.041 (liver) and 0.025 (lung) (effect of quantile normalization on cross-sample validation)
  • count 5357 enhancers, 8040 promoters (eHMM genome-wide predictions in mouse ESC)
  • count 2604 (c=0.9) to 12,830 (c=0.1) enhancers (REPTILE genome-wide predictions in mouse ESC by threshold)
  • count ChromHMM 19,643 (n=12) to 88,716 (n=6); EpiCSeg 37,911 (n=12) to 103,293 (n=6) enhancers (unsupervised method genome-wide enhancer counts by state number)
  • other median distance 42 bp (eHMM) vs 343 bp (REPTILE) (spatial accuracy of predicted enhancer center to nearest ATAC-seq peak)
  • other 3.2% (eHMM) vs 17.8%-35.0% (REPTILE) (fraction of predicted enhancers overlapping an annotated TSS)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper presents eHMM, a supervised hidden Markov model for enhancer prediction, and evaluates it using area under the precision-recall curve (AUPRC) as the primary performance metric. Evaluation follows a 5-fold cross-validation scheme within three mouse cell-type samples, supplemented by cross-sample validation using a model trained on ESC data and applied to liver and lung samples. No formal statistical hypothesis tests or p-values are reported; performance comparisons between methods are made by visual inspection of precision-recall curves and by directly comparing scalar AUPRC values and spatial accuracy metrics (median distances, IQR).

Replicationunclear Sample sizeThree publicly available mouse samples from ENCODE/FANTOM5 selected; no power analysis or sample-size justification reported GroupseHMM vs. ChromHMM (n=6,8,10,12 states), EpiCSeg (n=6,8,10,12 states), and REPTILE; within-sample vs. cross-sample validation settings Pairingna Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
5-fold cross-validation with area under the precision-recall curve (AUPRC) as metric Within-sample enhancer prediction performance in mouse ESC E14, liver E12.5, and lung E16.5 (Figs. 2a–b) 3 mouse samples; test set = 1/5 of enhancer training set per fold not stated
Cross-sample validation (AUPRC) ESC-trained model applied to liver E12.5 and lung E16.5 (FANTOM5) and to ESC E14, liver E14.5, lung E14.5 (EnhancerAtlas) (Figs. 2b–c) not stated
Median and IQR of distance distributions Spatial accuracy of predicted enhancers: distance to closest ATAC-seq peak and to closest TSS (Fig. 3c) Genome-wide predictions in mouse ESC (5357 enhancers for eHMM) not stated
Fraction/proportion calculation Fraction of predicted enhancers overlapping an annotated TSS, across methods and thresholds Genome-wide predicted enhancers per method per threshold na
Empirical runtime measurement CPU time and real time for training and prediction (Table 1) na
Approaches that could also have been used
  • AUPRC values are reported as single point estimates from cross-validation folds without uncertainty quantification
    Could also: Bootstrap resampling or reporting the mean and standard deviation of AUPRC across the 5 folds could also have been used to convey estimate variability — Uncertainty intervals around AUPRC allow readers to judge whether differences between methods exceed sampling noise, especially when fold-level values are not shown
  • Method comparisons are made by visually inspecting precision-recall curves and directly comparing scalar AUPRC values, without a formal significance test
    Could also: A permutation test or bootstrap-based comparison of AUPRC differences could also have been used to formally assess whether one method's AUPRC is reliably higher than another's — Formal comparison of AUC estimates (e.g., via DeLong-style tests adapted to PR curves, or bootstrap confidence intervals on the difference) provides a principled basis for claiming one method outperforms another beyond chance
  • Spatial accuracy is summarized using the median and IQR of center-to-peak distances for each method
    Could also: A Mann-Whitney U test (Wilcoxon rank-sum test) comparing the full distance distributions between eHMM and REPTILE could also have been applied — A formal non-parametric test on the distance distributions would supplement the descriptive summary by quantifying whether the observed difference in medians (42 bp vs. 343 bp) is statistically reliable given the distribution shapes
  • Three mouse samples (ESC, liver, lung) were used for evaluation without reported biological or technical replicates within each condition
    Could also: Repeated k-fold cross-validation (e.g., 5×5 or 10×5) or evaluation across additional independent samples from the same tissue-stage could also have been used — Repeated cross-validation reduces the variance of the AUPRC estimate attributable to the specific random fold partitioning, and additional independent samples would strengthen generalizability claims
  • Training and test sets are constructed to be unbalanced, reflecting genomic proportions of enhancers vs. background
    Could also: Stratified k-fold cross-validation explicitly preserving the positive/negative ratio in each fold could also have been used — Explicitly stratified splits ensure that each fold has the same class imbalance as the full dataset, which can reduce fold-to-fold variance in AUPRC when positive examples are rare
  • The performance metric is exclusively AUPRC; the F1-score or other threshold-specific metrics are not reported alongside it
    Could also: Reporting the F1-score or precision/recall at a fixed operating threshold (e.g., the Viterbi path used for whole-genome prediction) could also have been used alongside AUPRC — A threshold-specific metric such as F1 makes the operational trade-off between precision and recall concrete for practitioners choosing a prediction cutoff, complementing the threshold-free AUPRC summary
Software: eHMM (novel method, built on EpiCSeg framework) · ChromHMM · EpiCSeg · REPTILE · MACS2 (peak calling for ATAC-seq)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
17
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE11431 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE120376 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE29184 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE3859 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig 2bTable
C3
Reported
authors' shipped genome-wide demo = Fig-2b cross-sample scenario: pretrained mESC supervised HMM + quantile-norm applied genome-wide (mm10) to ENCODE mouse fetal-liver ATAC+H3K27ac/H3K4me1/H3K4me3 (ESC->liver E12.5, reported AUPRC 0.928); yields enhancerRegions.bed / promoterRegions.bed
Reproduced
REPRODUCED end-to-end: 4591 enhancers / 5331 promoters genome-wide; applyModel 516.9 s (ehmm 1.0 @ f5f1080, R4.1.3, «our HPC» «job», DONE_OK)
exact
C1
Reported
5357 genome-wide enhancers (mouse ESC, Results)
Reproduced
4591 in the LIVER cross-sample regime via the shipped demo; same order of magnitude, not the same cell type. Exact ESC integer needs GSE120376 raw reads -> BWA mm10 (not attempted, heavy)
partial
C2
Reported
8040 genome-wide promoters (mouse ESC, Results)
Reproduced
5331 in the liver cross-sample regime
partial
C4
Reported
runtime 46.597 s (Table 1, ESC train+predict)
Reproduced
516.9 s applyModel-only, genome-wide liver, 32 cores, prediction-only — not comparable (different data scope + hardware)
partial
C5
Reported
HMM inputs = ATAC + H3K27ac + H3K4me1 + H3K4me3
Reproduced
bamtab = ATAC-seq/H3K27AC/H3K4ME1/H3K4ME3, mark-name check passed — exact match
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 70/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

An incomplete run of the authors' own ehmm R package. The pipeline cloned at the pinned commit, built an R4.1/Bioconductor env (deliberately R4.1 to dodge the R>=4.2 BLAS break in ehmm's 2019 dgemm_ calls), compiled, and launched the shipped demo genome-wide on the correct ENCODE mouse fetal-liver data (verified via the ENCODE API) — but the genome-wide Viterbi decode was cut off at finalization before emitting enhancer/promoter counts. The exact paper counts (5357/8040) are ESC within-sample and out of scope, and the in-scope demo has no tabulated paper value (Fig-2b AUPRC 0.928 scenario), so the target was always a regime check. No number captured, no anomaly, no fabrication signal — a partial limited by our finalization timing.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

355.3 k
tokens (I/O) · 23.5 M incl. cache
117 min
runtime · 0.9 CPU-h
49 GB
peak RAM
1
HPC jobs
hummel
machine