Predicting enhancers in mammalian genomes using supervised hidden Markov models.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No relevant deviation in data/preprocessing
- 🟡Reported values were only indirectly comparable
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH: YES. ehmm is the authors' own R package shipping a self-contained, fully-specified genome-wide demo (pretrained mESC supervised HMM + quantile normalization -> enhancer/promoter prediction = the paper's Fig-2b cross-sample ESC->liver scenario). This room was a REQUEUE: the prior run built+launched but was cut off mid-decode. On resume I found (a) the «infra» work dir reclaimed by the janitor and (b) the «user» «infra» quota EXCEEDED (blocking ALL rooms; FS itself only ~18% full). Operator approved reclaiming ~17 GB of regenerable conda caches, which I deleted via an srun on a compute node (never front1). I then rebuilt the pipeline clean (clone @ f5f108086e50f61bcdac430ee2ad5d611d2141d2, R4.1.3 conda env to dodge the R>=4.2 USE_FC_LEN_T BLAS break, compiled ehmm 1.0) and ran the shipped genome-wide demo to COMPLETION on the correct ENCODE liver data (verified marks) -> 4591 enhancers / 5331 promoters (applyModel 516.9 s; «job» elapsed 00:59:44, DONE_OK). The reproduced liver counts are the cross-sample regime, NOT bit-comparable to the paper's within-sample ESC 5357/8040 (different cell type) but the same algorithm and order of magnitude. Datasets profiled (GSE120376: 26 samples, all ehmm input marks present, delivers-promised yes, grade A; ENCODE demo BAMs grade B). NOT ATTEMPTED (the hard ~20%): exact ESC integers (need GSE120376 raw reads -> BWA mm10) and Fig-2a CV AUPRC + REPTILE/ChromHMM/EpiCSeg benchmarks (no evaluate/CV subprogram shipped). No fabrication flags. All grades are provisional pending human review.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-14 ⛓ 90e0527bba2b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetEnhancers and promoters share a common physical topology—a central nucleosome-free accessible DNA stretch flanked by nucleosomes with distinct histone modification patterns—and encoding this structural prior into a supervised hidden Markov model can predict active enhancers genome-wide with higher precision and spatial resolution than existing unsupervised or 'black box' supervised methods.
- ★ eHMM predicts enhancers with high precision and recall comparable to state-of-the-art methods and consistently outperforms them in accuracy and resolution finding
- ★ Enhancers and promoters are modeled as a central accessible DNA stretch flanked by two nucleosomes, imposed via constrained HMM state transitions mechanism
- ★ eHMM requires only a minimal set of four features (ATAC-seq plus ChIP-seq for H3K27ac, H3K4me1, H3K4me3) to predict enhancers method
- eHMM's parameters are easy to interpret compared to other 'black box' enhancer prediction methods finding
- ★ eHMM predictions are spatially far more precise than REPTILE, with prediction centers much closer to ATAC-seq accessibility peaks finding
- ★ eHMM produces far fewer false enhancer predictions overlapping annotated TSS/promoters than REPTILE finding
- Quantile normalization improves cross-sample prediction accuracy when applying a pre-trained model to a different sample finding
- ★ eHMM can be used as a stand-alone, pre-trained tool for enhancer prediction without additional training or parameter tuning resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ATAC-seq | mouse ESC E14, liver E12.5, lung E16.5 | none | chromatin accessibility signal used as central feature of eHMM | — |
| ChIP-seq (H3K27ac, H3K4me1, H3K4me3) | mouse ESC E14, liver E12.5, lung E16.5 | none | histone modification signal defining flanking nucleosome states | — |
| CAGE (FANTOM5) | mouse ESC, liver, lung tissue-stages | none | bidirectional transcription initiation used to define enhancer training/validation sets orthogonal to histone ChIP-seq | — |
| 5-fold cross-validation and cross-sample validation of eHMM predictions | mouse ESC E14, liver E12.5, lung E16.5 | none | area under precision-recall curve (AUPRC) against FANTOM5 enhancers | — |
| Benchmark comparison against ChromHMM, EpiCSeg, and REPTILE | mouse ESC E14, liver E12.5/E14.5, lung E16.5/E14.5; validated on FANTOM5 and EnhancerAtlas regions | none | AUPRC comparison across methods and state numbers | — |
| Transcription factor and chromatin remodeler ChIP-seq (Nanog, Oct4, Sox2, CTCF, p300) | mouse ESC | none | binding enrichment at predicted enhancers vs promoters | — |
| MeDIP-seq | mouse ESC | none | DNA methylation levels at predicted enhancers and promoters | — |
| phastCons sequence conservation analysis | mouse ESC (cross-species comparison) | none | sequence conservation score at predicted enhancers and promoters | — |
- ▲ eHMM within-sample (5-fold CV) precision-recall performance on FANTOM5 enhancers across ESC, liver, and lung AUPRC 0.947-0.971
- ▼ eHMM cross-sample validation (ESC-trained model applied to liver and lung) still performs well though slightly lower than within-sample AUPRC 0.928 (liver E12.5), 0.865 (lung E16.5)
- ▲ Quantile normalization improves cross-sample prediction quality relative to no normalization AUPRC increase of 0.041 (liver), 0.025 (lung)
- ▲ eHMM and REPTILE (supervised methods) outperform ChromHMM and EpiCSeg (unsupervised) within and across cell types
- ▼ eHMM-predicted enhancer centers are much closer to ATAC-seq accessibility peaks than REPTILE predictions median distance 42 bp (eHMM) vs 343 bp (REPTILE), ~8-fold closer
- ▼ eHMM predicted enhancers overlap annotated TSS far less often than REPTILE predicted enhancers 3.2% (eHMM) vs 17.8%-35.0% (REPTILE, threshold-dependent)
- – Genome-wide eHMM prediction in mouse ESC yields a defined number of enhancers and promoters without threshold selection 5357 enhancers, 8040 promoters
- – Run time comparison shows REPTILE uses least total CPU time but longest real time; eHMM has lowest total real time among compared methods eHMM 46.597 s real / 171.157 s CPU; REPTILE 90.917 s real / 145.550 s CPU
- other AUPRC 0.947-0.971 (eHMM within-sample 5-fold CV on FANTOM5 in mouse ESC, liver, lung)
- other AUPRC 0.928 (liver), 0.865 (lung) (eHMM cross-sample validation using ESC-trained model)
- other AUPRC increase of 0.041 (liver) and 0.025 (lung) (effect of quantile normalization on cross-sample validation)
- count 5357 enhancers, 8040 promoters (eHMM genome-wide predictions in mouse ESC)
- count 2604 (c=0.9) to 12,830 (c=0.1) enhancers (REPTILE genome-wide predictions in mouse ESC by threshold)
- count ChromHMM 19,643 (n=12) to 88,716 (n=6); EpiCSeg 37,911 (n=12) to 103,293 (n=6) enhancers (unsupervised method genome-wide enhancer counts by state number)
- other median distance 42 bp (eHMM) vs 343 bp (REPTILE) (spatial accuracy of predicted enhancer center to nearest ATAC-seq peak)
- other 3.2% (eHMM) vs 17.8%-35.0% (REPTILE) (fraction of predicted enhancers overlapping an annotated TSS)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper presents eHMM, a supervised hidden Markov model for enhancer prediction, and evaluates it using area under the precision-recall curve (AUPRC) as the primary performance metric. Evaluation follows a 5-fold cross-validation scheme within three mouse cell-type samples, supplemented by cross-sample validation using a model trained on ESC data and applied to liver and lung samples. No formal statistical hypothesis tests or p-values are reported; performance comparisons between methods are made by visual inspection of precision-recall curves and by directly comparing scalar AUPRC values and spatial accuracy metrics (median distances, IQR).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| 5-fold cross-validation with area under the precision-recall curve (AUPRC) as metric | Within-sample enhancer prediction performance in mouse ESC E14, liver E12.5, and lung E16.5 (Figs. 2a–b) | 3 mouse samples; test set = 1/5 of enhancer training set per fold | not stated |
| Cross-sample validation (AUPRC) | ESC-trained model applied to liver E12.5 and lung E16.5 (FANTOM5) and to ESC E14, liver E14.5, lung E14.5 (EnhancerAtlas) (Figs. 2b–c) | — | not stated |
| Median and IQR of distance distributions | Spatial accuracy of predicted enhancers: distance to closest ATAC-seq peak and to closest TSS (Fig. 3c) | Genome-wide predictions in mouse ESC (5357 enhancers for eHMM) | not stated |
| Fraction/proportion calculation | Fraction of predicted enhancers overlapping an annotated TSS, across methods and thresholds | Genome-wide predicted enhancers per method per threshold | na |
| Empirical runtime measurement | CPU time and real time for training and prediction (Table 1) | — | na |
-
AUPRC values are reported as single point estimates from cross-validation folds without uncertainty quantification↳ Could also: Bootstrap resampling or reporting the mean and standard deviation of AUPRC across the 5 folds could also have been used to convey estimate variability — Uncertainty intervals around AUPRC allow readers to judge whether differences between methods exceed sampling noise, especially when fold-level values are not shown
-
Method comparisons are made by visually inspecting precision-recall curves and directly comparing scalar AUPRC values, without a formal significance test↳ Could also: A permutation test or bootstrap-based comparison of AUPRC differences could also have been used to formally assess whether one method's AUPRC is reliably higher than another's — Formal comparison of AUC estimates (e.g., via DeLong-style tests adapted to PR curves, or bootstrap confidence intervals on the difference) provides a principled basis for claiming one method outperforms another beyond chance
-
Spatial accuracy is summarized using the median and IQR of center-to-peak distances for each method↳ Could also: A Mann-Whitney U test (Wilcoxon rank-sum test) comparing the full distance distributions between eHMM and REPTILE could also have been applied — A formal non-parametric test on the distance distributions would supplement the descriptive summary by quantifying whether the observed difference in medians (42 bp vs. 343 bp) is statistically reliable given the distribution shapes
-
Three mouse samples (ESC, liver, lung) were used for evaluation without reported biological or technical replicates within each condition↳ Could also: Repeated k-fold cross-validation (e.g., 5×5 or 10×5) or evaluation across additional independent samples from the same tissue-stage could also have been used — Repeated cross-validation reduces the variance of the AUPRC estimate attributable to the specific random fold partitioning, and additional independent samples would strengthen generalizability claims
-
Training and test sets are constructed to be unbalanced, reflecting genomic proportions of enhancers vs. background↳ Could also: Stratified k-fold cross-validation explicitly preserving the positive/negative ratio in each fold could also have been used — Explicitly stratified splits ensure that each fold has the same class imbalance as the full dataset, which can reduce fold-to-fold variance in AUPRC when positive examples are rare
-
The performance metric is exclusively AUPRC; the F1-score or other threshold-specific metrics are not reported alongside it↳ Could also: Reporting the F1-score or precision/recall at a fixed operating threshold (e.g., the Viterbi path used for whole-genome prediction) could also have been used alongside AUPRC — A threshold-specific metric such as F1 makes the operational trade-off between precision and recall concrete for practitioners choosing a prediction cutoff, complementing the threshold-free AUPRC summary
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
eHMM enhancer centers are a median 42 bp from the nearest ATAC-seq peak versus 343 bp for REPTILE, indicating ~8-fold better spatial accuracyatac-seq mouse-esc 2019×1papers★ This paper is the founder (earliest)
-
eHMM achieves AUPRC 0.947–0.971 in within-sample 5-fold cross-validation for FANTOM5 enhancer prediction in mouse ESC, liver, and lungother mouse-esc-liver-lung 2019×1papers★ This paper is the founder (earliest)
-
Supervised HMMs outperform unsupervised HMMs (ChromHMM, EpiCSeg) for enhancer prediction; best unsupervised performance requires 10–12 statesother mouse-esc-liver-lung up 2019×1papers★ This paper is the founder (earliest)
-
eHMM-predicted enhancers in mouse ESC are unimodally distributed in distance to the nearest annotated TSS with an IQR of 11–85 kbother mouse-esc 2019×1papers★ This paper is the founder (earliest)
-
Only 3.2% of eHMM-predicted enhancers overlap an annotated TSS compared to 17.8–35.0% for REPTILE, demonstrating higher specificityother mouse-esc down 2019×1papers★ This paper is the founder (earliest)
-
Genome-wide eHMM prediction in mouse ESC identifies 5357 enhancers and 8040 promotersother mouse-esc 2019×1papers★ This paper is the founder (earliest)
-
Cross-sample application of an ESC-trained eHMM to mouse liver E12.5 and lung E16.5 yields AUPRC 0.928 and 0.865, reduced versus within-sample performanceother mouse-liver-lung down 2019×1papers★ This paper is the founder (earliest)
-
Quantile normalization of ChIP-seq and ATAC-seq features improves cross-sample enhancer prediction AUPRC by 0.041 in mouse liver and 0.025 in mouse lungother mouse-liver-lung up 2019×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
An incomplete run of the authors' own ehmm R package. The pipeline cloned at the pinned commit, built an R4.1/Bioconductor env (deliberately R4.1 to dodge the R>=4.2 BLAS break in ehmm's 2019 dgemm_ calls), compiled, and launched the shipped demo genome-wide on the correct ENCODE mouse fetal-liver data (verified via the ENCODE API) — but the genome-wide Viterbi decode was cut off at finalization before emitting enhancer/promoter counts. The exact paper counts (5357/8040) are ESC within-sample and out of scope, and the in-scope demo has no tabulated paper value (Fig-2b AUPRC 0.928 scenario), so the target was always a regime check. No number captured, no anomaly, no fabrication signal — a partial limited by our finalization timing.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.