Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genome-wide prediction of DNase I hypersensitivity using gene expression.

Nat Commun · 2017
L1 79/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
79/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 55% of all assessed papers rank 514 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1 on the deterministic core. BIRD is a third-party-style deterministic C++ tool whose released exon-array model (BIRD-model v1.1, SHA f5d388...) + paper's own training data (BIRD-data v1.0 = GSE19090 exon arrays + ENCODE DNase) reproduce the paper's headline STRUCTURAL numbers exactly: 1,108,603 DHS loci x 57 cell types, genome-wide prediction lands on exactly those 1,108,603 loci and is bit-deterministic, and the loci cluster into 1000/2000/5000 clusters as described. The held-out CV ACCURACY (r_L=0.82, r_C=0.50) is the deliberately-skipped ~20%: the shipped .bin is the full 57-cell model, so only IN-SAMPLE correlations could be computed (r_L=0.908, r_C=0.756) -- upper bounds that exceed the held-out figures in the correct direction, corroborating but not 1:1-reproducing them. NOT attempted: 40/17 CV retraining (BigKmeans + R scripts + BIRD_build_library) and the 912,886-loci CV model. No fabrication signal: every reproduced value is regenerable from the public code+data at the pinned commit/SHAs.

💻 Code ↗ 🗄 Data: GSE19090

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 79
    assessed: 2026-06-14 ⛓ 05e98c3b24eb
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a biological sample's transcriptome (gene expression) be used to predict its genome-wide regulatory element activities measured by DNase I hypersensitivity (DH), and to what extent are regulatory elements' activities predictable from the whole transcriptome?

Core claims
  • Gene expression substantially predicts genome-wide DNase I hypersensitivity (DH), demonstrating transcriptome-based prediction as a feasible approach for regulome mapping finding
  • BIRD (Big Data Regression for predicting DH) handles the ultra-high-dimensional prediction problem by clustering co-expressed genes and aggregating locus-level and pathway-level models method
  • Information useful for predicting DH is contained in the whole transcriptome rather than limited to a regulatory element's neighboring genes finding
  • BIRD-predicted DH can be used to predict transcription factor-binding sites (TFBSs), build a regulome database from GEO expression samples, predict differential regulatory element activities, and serve as pseudo-replicates resource
  • DH correlates in trans with expression of TFs binding the locus and co-expressed genes, explaining why whole-transcriptome prediction outperforms neighboring-gene prediction mechanism
  • BIRD produces the best prediction performance among compared methods while remaining computationally efficient for big data regression finding
Experimental setups
Assay System Perturbation Readout Platform
DNase-seq (chromatin accessibility / DNase I hypersensitivity) 57 distinct human cell types with normal karyotype (40 training, 17 test) from ENCODE none DH bin read count / DH signal level at genomic loci (912,886 DHSs) ENCODE DNase-seq
Exon array (gene expression profiling) 57 distinct human cell types from ENCODE (40 training, 17 test) none gene expression levels of 18,000+ genes used as predictors Exon array (ENCODE)
Computational regression prediction (BIRD) ENCODE human cell types; applied to GEO gene expression samples none predicted genome-wide DH levels per locus BIRD algorithm
Key results
  • BIRD achieved high cross-locus prediction accuracy (mean P–T correlation r_L) in the 17 test cell types r_L = 0.82 (mean)
  • BIRD predicted cross-cell-type DH variation, more challenging than cross-locus variation r_C = 0.50 (mean) vs r_L = 0.82
  • Random/permuted models also yielded high cross-locus correlation due to locus-specific DH propensity, but BIRD exceeded them BIRD r_L = 0.82 vs random r_L = 0.65
  • Random prediction models had cross-cell-type correlation centered around zero, while BIRD substantially increased r_C random r_C = -0.03 (mean)
  • Whole-transcriptome BIRD prediction substantially outperformed best neighboring-gene approach across r_L, r_C, and τ
  • Fused lasso achieved similar accuracy to BIRD's locus-level model but was vastly slower >10^5 times slower
  • Cross-cell-type prediction accuracy varied greatly among loci, with a subset of loci showing high r_C 6% of loci (partial statement)
Key statistics
  • correlation r_L = 0.82 (mean cross-locus P–T correlation for BIRD across 17 test cell types)
  • correlation r_C = 0.50 (mean cross-cell-type P–T correlation for BIRD across genomic loci)
  • correlation r_L = 0.65 (mean cross-locus correlation for random (permuted) prediction models)
  • correlation r_C = -0.03 (mean cross-cell-type correlation for random prediction models)
  • count 912,886 genomic loci (DHSs) (DHSs retained with unambiguous DNase-seq signal in at least one training cell type)
  • count 57 human cell types (40 training, 17 test) (ENCODE cell types used for training and testing)
  • pvalue p < 10^-15 (Two-sided Wilcoxon signed-rank test comparing methods on r_C)
  • pvalue p < 10^-4 (Two-sided Wilcoxon signed-rank test comparing methods on r_L)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational methods paper that develops and benchmarks a prediction algorithm (BIRD) using a train/test design on ENCODE data: 57 human cell types were randomly partitioned into 40 training and 17 test cell types, and prediction accuracy was assessed with Pearson correlations between predicted and true DH (cross-locus r_L and cross-cell-type r_C) and a normalized squared prediction error (τ). Method-vs-method comparisons of these accuracy statistics were assessed with two-sided Wilcoxon signed-rank tests, and distributions were summarized with boxplots (median, quartiles, 1.5×IQR whiskers).

Replicationunclear Sample size57 cell types randomly partitioned into 40 training and 17 test; 912,886 genomic loci retained after filtering; no formal power/sample-size calculation described GroupsBIRD vs. alternative/random/neighboring-gene prediction methods Pairingpaired Randomization/blindingstated DispersionIQR Exact p-valuesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Two-sided Wilcoxon signed-rank test Comparing cross-locus P–T correlation (r_L) between methods (Fig. 2b) 17 (test cell types) not stated
Two-sided Wilcoxon signed-rank test Comparing cross-cell-type P–T correlation (r_C) between methods (Fig. 2c) 912,886 (genomic loci) not stated
Pearson's correlation (predicted vs. true DH) Evaluation metric for r_L (across loci) and r_C (across cell types) (Fig. 2) not stated
Approaches that could also have been used
  • Method comparisons used the two-sided Wilcoxon signed-rank test on paired accuracy statistics.
    Could also: A paired t-test (when approximate normality holds) or a permutation/bootstrap test on the paired differences could also be used. — A paired t-test can offer more power when its assumptions hold, while a bootstrap can directly yield confidence intervals for the mean difference in accuracy.
  • Significance was reported as p-value thresholds (e.g., p < 10^-4, p < 10^-15).
    Could also: Reporting exact p-values alongside an effect-size measure (e.g., the median paired difference or a 95% CI for it) could also be presented. — Exact values and effect sizes convey the magnitude of improvement, not just that a difference exists, which is informative when n is very large and tiny differences become significant.
  • Distributions of r_L and r_C were summarized with boxplots showing median and IQR.
    Could also: Adding mean ± SD or a 95% CI, or overlaying the raw points, could also describe the spread. — These complementary summaries make central tendency and uncertainty explicit, which can aid comparison across methods.
  • Prediction accuracy was evaluated on a single random 40/17 train/test split.
    Could also: Cross-validation or repeated random splits could also be used to estimate accuracy. — Resampling provides a distribution of performance estimates and a sense of variability due to the particular partition chosen.
  • Prediction agreement was quantified with Pearson's correlation between predicted and true DH.
    Could also: Spearman's rank correlation or concordance/CCC, alongside the reported squared error, could also be used. — Rank-based or concordance measures are robust to nonlinearity and outliers and capture agreement (not just linear association), complementing the squared-error metric.
  • For r_C, the test was based on n = 912,886 loci treated as the unit of comparison.
    Could also: A multilevel/clustered analysis that accounts for correlation among co-activated loci (e.g., DHS 'pathways') could also be considered. — Modeling the dependence structure among loci can give a more conservative assessment of uncertainty when many loci co-vary.

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
53
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE15805 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE19090 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE24976 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE32219 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE46837 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE51004 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE93012 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE9703 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-29051481 (BIRD)

Paper: Zhou et al. 2017, "Genome-wide prediction of DNase I hypersensitivity using gene expression", Nat Commun. BIRD = a deterministic C++ regression tool + a pre-built model. The released model and training data ARE the paper's pipeline artifacts (third-party-tool-style reproduction, P16-valid).

In scope (pipeline-derived) — attempted

  • Model/data structure (C1, C2, C4): number of DHS loci (1,108,603), number of cell types (57), and the DHS clustering (1000/2000/5000). Pipeline = the released BIRD-data v1.0 matrices + cluster files. DETERMINISTIC, exactly reproduced.
  • Prediction (C3): running the released exon-array model (BIRD_predict) on the shipped K562 example. Pipeline = the BIRD C++ predictor. DETERMINISTIC, reproduced.
  • Accuracy (C5, C6) — partial: cross-locus r_L and cross-cell r_C. Pipeline = BIRD prediction + correlation against measured DNase. Reproduced only IN-SAMPLE (full model on its own training cells); the paper's HELD-OUT CV values were not re-derived (see below).

In scope but NOT attempted (the hard ~20%)

  • Held-out cross-validation (true r_L=0.82, r_C=0.50): requires retraining a BIRD model on 40 cells and testing on 17 — BigKmeans + R_script/get_param.r, get_DHS_cluster.r, get_model_data.r + BIRD_build_library. Feasible but the deliberately-skipped 20%. The 912,886-loci 40-cell CV model is not shipped as a .bin.

Out of scope (wet-lab / external / manual)

  • Generation of the ENCODE DNase-seq and exon-array assays themselves.
  • The PDDB web prediction database (2,000 GEO samples) and its server.
  • Roadmap (70-cell) and ENCODE RNA-seq (167-cell) model variants — different releases, not the original paper's exon-array model.

Pipelines named per result

  • C1/C2/C4: data inspection of BIRD-data v1.0 (.rds + cluster .txt).
  • C3: BIRD BIRD_predict (C++), exon-array full model.
  • C5/C6: BIRD BIRD_predict + R Pearson correlation vs measured DNase.
Figures / tables: Fig 2
C1
Reported
1,108,603 DHS loci (full model)
Reproduced
1108603 rows in DNase_data_57_cells.rds
exact
C2
Reported
57 distinct human cell types
Reproduced
57 cell columns in DNase & Exon training matrices
exact
C3
Reported
genome-wide DNase prediction at the model loci (~1.1M)
Reproduced
BIRD_predict output = 1,108,603 loci, bit-deterministic on rerun
exact
C4
Reported
DHS clustered into 1000/2000/5000 clusters
Reproduced
DH_cluster_{1000,2000,5000}.txt each assign all 1,108,603 loci
exact
C5
Reported
cross-locus correlation r_L = 0.82 (held-out)
Reproduced
in-sample mean r_L = 0.908 (min 0.838, max 0.948, n=57)
partial
C6
Reported
cross-cell-type correlation r_C = 0.50 (held-out)
Reproduced
in-sample mean r_C = 0.756 (50k random loci)
partial
C7
Reported
912,886 loci (40-cell CV model)
Reproduced
not attempted (CV model not shipped as .bin)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 79/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

The deterministic core reproduces exactly from the pinned public BIRD model+data: 1,108,603 DHS loci × 57 cell types, genome-wide prediction landing on exactly those loci and bit-deterministic across reruns, plus the 1000/2000/5000 clustering. The paper's held-out accuracy (r_L=0.82, r_C=0.50) was not 1:1 reproduced because only the full 57-cell .bin is shipped — so in-sample correlations (0.908, 0.756) were computed instead, which upper-bound the held-out values in exactly the expected direction, corroborating but not re-validating them. The 40/17 CV retraining is the deliberately-skipped hard 20%. No fabrication signal — every value is regenerable at the pinned SHAs; a solid partial limited by the accuracy not being held-out-reproduced.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

136.4 k
tokens (I/O) · 7.3 M incl. cache
15 min
runtime · 0.03 CPU-h
2.6 GB
peak RAM
3
HPC jobs
hummel
machine