Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comprehensive enhancer-target gene assignments improve gene set level interpretation of genome-wide regulatory data.

Genome Biol · 2022
not yet assessed 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

DROP (no_code). The only resolvable code artifact for PMID 35473573 is TrimGalore, a generic adapter/quality-trimming tool, which does not implement the paper's reported computational claims (comprehensive enhancer-target gene assignments improving gene-set-level interpretation of genome-wide regulatory data). The headline results are not reproducible from this link: TrimGalore only trims reads and the paper provides no pinnable trimmed-read output to grade against. Public data (GSE180260) appears to exist, but without the actual analysis pipeline there is no in-scope, comparable pipeline-derived result. No scope.md/claims.tsv were produced and no SLURM/«our HPC» jobs were run in this room; nothing was fabricated. NOT ATTEMPTED: locating the paper's true analysis code (possibly a separate Babraham/author repo or supplementary), downloading/processing GSE180260, and any enhancer-target or enrichment comparison. Recommendation: re-screen the code-link extraction (likely a text-mining false positive that captured a preprocessing dependency instead of the analysis repo); if the real analysis pipeline is found, this RU can be upgraded from drop to a genuine reproduction attempt.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-14 ⛓ 9239c8b39f50
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper aims to determine the best genome-wide definitions of human enhancer locations and their distal target genes, testing whether integrating multiple spatial and in silico enhancer-gene linking approaches outperforms the naïve nearest-gene assignment for gene set level interpretation of regulatory data.

Core claims
  • Integrating multiple enhancer-defining and enhancer-target gene linking methods yields 1860 genome-wide enhancer-to-target gene definitions (EnTDefs), of which the top 741 (~40%) significantly outperform the naïve nearest-gene assignment of distal regions. finding
  • A handful of top-ranked general EnTDefs perform well across cell types and outperform seven independent computational and experiment-based enhancer-gene pair datasets. finding
  • GSE-based ranking of EnTDefs is highly concordant with ranking based on overlap with curated benchmarks of enhancer-gene interactions. finding
  • Using top EnTDefs for GSE with DNA methylation or ATAC-seq data better recapitulates biological processes changed in parallel gene expression data than lower-ranked EnTDefs. finding
  • General (non-cell-type-specific) EnTDefs are more favorable than cell-type-specific EnTDefs (CT-EnTDefs). finding
  • EnTDefs were ranked by concordance of GO biological process GSE results from 87 ENCODE TF ChIP-seq datasets with curated GO annotations using F1 scores. method
  • Adding FANTOM5 and ChIA enhancer-gene assignment methods, and omitting enhancer extension, significantly improve EnTDef performance, whereas ChromHMM and the loop (L2/L3) methods contribute little. finding
  • 1860 EnTDefs constitute a reusable resource of genome-wide distal enhancer-to-target gene definitions aggregated across >500 cell types. resource
Experimental setups
Assay System Perturbation Readout Platform
TF ChIP-seq (used for GSE evaluation) ENCODE samples across multiple cell types none distal ChIP-seq peaks assigned to genes; GO BP enrichment F1 score concordance with curated TF GO annotations
ChIA-PET >500 cell types (ENCODE) none enhancer-target gene interaction links and CTCF convergent-motif loop boundaries
DNase-seq (DHS) multiple human cell types none enhancer locations and DNase-signal correlation-based enhancer-promoter links (Thurman)
CAGE (FANTOM5) >500 cell types none enhancer locations and expression-correlation-based enhancer-gene links
ChromHMM chromatin state annotation ENCODE UCSC tracks none enhancer region definitions
DNA methylation (Bisulfite-seq) GSE experiment with parallel gene expression other recapitulation of biological processes changed in expression data
ATAC-seq GSE experiment with parallel gene expression other recapitulation of biological processes changed in expression data
Key results
  • Top 741 (~40%) EnTDefs significantly outperform the >5 kb nearest-gene LocDef 741 of 1860 (~40%), Wilcoxon FDR<0.05
  • EnTDefs ranked 2-19 not significantly worse than the best-performing EnTDef ranks 2-19, p>0.01
  • All ten EnTDef_plus5kb significantly outperform the nearest TSS method ~0.05 increase in average F1, p<0.0001
  • Top 10 EnTDefs (distal only) outperform GREAT, FET, and Poly-Enrich using 5 kb LocDef average F1 0.47 vs 0.45, p<0.007
  • Top 10 EnTDefs and 5 kb LocDef significantly outperform >5 kb LocDef F1 0.47 and 0.45 vs 0.27
  • Adding FANTOM5 enhancer definition significantly improves >50% of EnTDefs; ChromHMM only ~5% FANTOM5 >50%, DNase ~24%, Thurman ~16%, ChromHMM ~5%
  • FANTOM5 and ChIA enhancer-gene assignment methods improve ~70% of EnTDefs; L and Thurman improve ~1.7% and ~9% ~70% vs ~1.7% and ~9%
  • EnTDefs without enhancer extension improve ~60% of EnTDefs vs ~7% with 1 kb extension ~60% vs ~7%
Key statistics
  • count 1860 EnTDefs (total genome-wide enhancer-to-target gene definitions generated)
  • count 1,768,201 individual enhancer-target links (from 685,921 enhancers and 21,094 target genes across >500 cell types)
  • count 87 ENCODE ChIP-seq datasets, 34 TFs (evaluation datasets for GSE F1 scoring)
  • fold_change average F1 = 0.47 vs 0.45 (top 10 EnTDefs vs 5 kb LocDef GSE methods, Wilcoxon p<0.007)
  • pvalue p = 2.37 × 10^-14 (top 10 EnTDefs vs >5 kb LocDef (F1 0.47 vs 0.27))
  • pvalue p = 1.32 × 10^-8 (5 kb LocDef vs >5 kb LocDef (F1 0.45 vs 0.27))
  • pvalue p = 0.91 (Friedman test: Poly-Enrich, GREAT, FET using 5 kb LocDef performed equally well)
  • count median 2 genes per enhancer (range 1-2); median 20 enhancers per gene (range 2-98) (characteristics of top 741 EnTDefs)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This computational benchmarking study generated 1,860 genome-wide enhancer-to-target gene definitions (EnTDefs) by combining four enhancer-location sources with four linking methods across >500 cell types. Performance was evaluated using gene set enrichment (GSE) testing on 87 ENCODE ChIP-seq datasets for 34 transcription factors, with F1 scores (comparing significantly enriched GO biological process terms to curated TF GO annotations) as the primary metric. EnTDefs were ranked by average F1 score across TFs, and pairwise Wilcoxon tests — with FDR or p-value thresholds — were used to identify the set of significantly top-performing definitions. Results were reported as average F1 scores and p-values comparing key approaches.

Replicationunclear Sample size87 ENCODE ChIP-seq datasets for 34 distinct TFs used as evaluation corpus; 1,860 EnTDefs benchmarked; no formal power calculation described Groups1,860 EnTDefs vs. nearest-gene baseline (>5 kb LocDef); top EnTDefs vs. nearest TSS, GREAT, FET, and cell-type-specific definitions Pairingmixed Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionFDR (specific procedure not named) for the 1,860 EnTDef vs. baseline comparisons; uncorrected p < 0.01 threshold for sequential pairwise top-ranked comparisons
Statistical tests used
Test Applied to n Assumptions
Wilcoxon signed-rank test Identifying the 741 EnTDefs that significantly outperform the >5 kb nearest-gene locus definition (LocDef) baseline 87 ENCODE ChIP-seq datasets for 34 TFs not stated
TF-paired Wilcoxon rank-sum test (labeled 'sum-rank' in Fig. 1 caption) Sequential pairwise comparison of the top-ranked EnTDef against each lower-ranked EnTDef to identify the set not significantly worse (p > 0.01 threshold) 34 TFs not stated
Paired Wilcoxon test Assessing the relative contribution of each individual enhancer-definition, extension, or linking method by comparing EnTDefs containing vs. excluding that method (Fig. 2B) 741 top-ranked EnTDefs not stated
Wilcoxon signed-rank test Comparison of EnTDef_plus5kb (top 10) vs. nearest TSS method; top 10 EnTDefs and 5 kb LocDef vs. >5 kb LocDef; top 10 EnTDefs vs. nearest TSS (non-significant for top half) 87 ChIP-seq datasets / 34 TFs not stated
Wilcoxon rank-sum test Comparison of top 10 EnTDefs vs. GREAT/FET/5 kb LocDef methods (average F1 = 0.47 vs 0.45, p < 0.007) 34 TFs not stated
Friedman test Omnibus comparison of three GSE testing methods (Poly-Enrich, GREAT, Fisher's exact test with 5 kb LocDef) for overall equivalence (p = 0.91) 34 TFs not stated
Approaches that could also have been used
  • Performance was summarized as average F1 score at a single significance threshold across TFs, reported as a point estimate with no dispersion.
    Could also: Area under the precision-recall curve (AUPRC) or AUROC, plus a standard deviation or 95% CI around mean F1 across TFs, could also be reported. — F1 depends on the significance threshold chosen; AUPRC/AUROC integrate across thresholds providing a threshold-independent summary, and dispersion measures would convey how consistently an EnTDef performs across diverse TFs rather than just on average.
  • Sequential pairwise Wilcoxon tests were performed between the top-ranked EnTDef and each lower-ranked one to identify the set not significantly worse (p > 0.01), without explicit multiplicity correction for the 1,859 comparisons.
    Could also: A single multiplicity-corrected ANOVA or mixed-effects model with TF as a random effect, followed by Dunnett-style contrasts against the reference EnTDef, could also control the family-wise error rate across all pairwise comparisons. — Sequential pairwise tests against a reference accumulate type-I error across comparisons; a model-based approach with formal FWER control would provide tighter guarantees on which EnTDefs are genuinely equivalent to the top performer.
  • FDR correction was applied to the comparison of 1,860 EnTDefs vs. the nearest-gene baseline, but the specific procedure (e.g., Benjamini-Hochberg, Storey q-value) and the exact family of tests were not stated.
    Could also: Explicitly naming the FDR procedure and the complete set of tests it covered could also be included. — Different FDR procedures (BH vs. Storey q-value vs. Bonferroni) have different assumptions and power; stating the method allows readers to reproduce the threshold and evaluate its appropriateness for the correlation structure among the 1,860 EnTDefs.
  • The three GSE methods (Poly-Enrich, GREAT, FET) were declared equivalent based on a non-significant Friedman test (p = 0.91).
    Could also: Equivalence testing (e.g., two one-sided tests, TOST) or reporting confidence intervals around pairwise F1 differences could also be used to formally establish equivalence. — A non-significant omnibus test does not demonstrate equivalence; TOST or a CI-based approach would quantify the magnitude of any difference and whether it falls within a pre-specified equivalence margin, supporting a positive claim of similarity.
  • The evaluation used GO biological process annotations from the GO database as the ground truth for TF function, both to rank and to describe EnTDef performance.
    Could also: Cross-validation with held-out TFs, or ranking on one independent benchmark (e.g., curated enhancer-gene pairs from ENCODE functional validation) while reporting performance on GO annotations separately, could also be used. — Using the same annotation source for both ranking EnTDefs and evaluating them can favor definitions that align with GO database coverage rather than with biological ground truth; an orthogonal benchmark provides a less circular performance estimate.
  • Method contribution (Fig. 2B) was assessed by comparing EnTDefs containing vs. excluding each method using paired Wilcoxon tests, reporting the percent of EnTDefs showing significant improvement.
    Could also: A permutation-based or bootstrap importance analysis, or a factorial design decomposing variance in F1 attributable to each method factor and their interactions, could also quantify method contributions. — Reporting the percent of pairwise comparisons that are significant conflates effect size with statistical power; a variance-decomposition or effect-size approach would more directly quantify how much each method contributes to overall performance.
Software: Not explicitly stated in available text

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
18
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GM12878 IGSR/1000 Genomes in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
also used by 5 papers:
GSE140341 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE180260 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

No individual results have been recorded for this entry yet.

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 44/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

This is a controlled DROP (no_code): the only resolvable code artifact, TrimGalore, is a generic adapter/quality-trimmer and not the enhancer→target-gene assignment / gene-set pipeline behind the paper's headline claims (Genome Biology 2022). No claims were pinned, GSE180260 was never processed, and no «our HPC» compute was run, so derivability and the core conclusion are undetermined, not refuted. The blocker is on our side — a likely text-mining false positive in code-link extraction — rather than an authors' defect or fabrication, so q4 is our-methodology and q5/q7/q8 stay yellow (cautionary) rather than red. Recommendation: re-screen the code link; if the real analysis repo is found this RU can be upgraded to a genuine reproduction attempt.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

19.7 k
tokens (I/O) · 461 k incl. cache
10 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.