Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A computational pipeline to visualize DNA-protein binding states using dSMF data.

STAR Protoc · 2022
L1 89/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
89/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 77% of all assessed papers rank 246 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED. STAR Protocols dSMF visualization pipeline (Rao & Ramachandran 2022) run faithfully on «our HPC» from the authors' pinned repo @028c14ca + shipped demo data. 6 of 7 claims reproduced: C1 (Fig1 peak_229 heatmap) and C3 (Fig2 peak_110 cobinding heatmap) reproduced 1:1 in structure; C2 (25 discarded molecules) and C5 (8 of 9 cobinding states, missing D-N) are EXACT integer matches; C4 (78bp peak separation -> 77bp) and C7 (median molecule length 269bp -> 267bp) within-tol. The two exact count matches plus the two within-1-2bp matches on the shipped demo are strong evidence the published figures are faithfully derivable from the deposited data+code (no fabrication detected). C6 (open-enhancer population state distribution, an 'Expected outcomes' soft range) was not attempted: it is an aggregate over all open enhancers requiring the full ~33GB GSE77369 S2 dataset and an externally-defined open-enhancer set, beyond the shipped demo that the protocol provides to reproduce the figures. Out of scope: wet-lab dSMF library prep / bisulfite conversion / sequencing (GSE77369 generation). Environment hurdles solved: /home quota (redirected HOME+conda cache+TMPDIR to «infra»), gnuplot libGL.so.1 (added conda-forge libgl/libglx/libopengl), snakemake5 --configfile arg-order, /tmp full on nodes.

💻 Code ↗ 🗄 Data: GSE77369

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 89
    assessed: 2026-06-21 ⛓ 01ff9ab240f2
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper presents (rather than tests) the premise that dual-enzyme single-molecule footprinting (dSMF) data can be computationally processed to map and quantify in vivo protein-DNA binding states, including cooperative (cobinding) events, at individual DNA molecules.

Core claims
  • The pipeline maps states of protein-binding DNA in vivo using dSMF data and identifies binding states at an enhancer in Drosophila S2 cells method
  • The pipeline infers and quantifies cooperative (cobinding) events between transcription factors on chromatinized DNA finding
  • Unlike MNase-seq and DNase-seq, dSMF maps the unbound (naked) state of genomic DNA finding
  • In open Drosophila S2 enhancers, one should expect ~40%-50% naked-DNA, ~15%-20% TF-bound, and ~30%-40% nucleosomal DNA molecules finding
  • A footprint is defined as one or more unmethylated cytosines flanked by methylated cytosines (>=10 bp); footprints of 10-50 bp are classified as TF-bound and larger footprints as nucleosomal method
  • Eight of nine possible cobinding states were found at the enhancer Peak 110 locus finding
  • The pipeline is built on Snakemake to enable scalable, reproducible workflow execution on high-performance clusters method
  • The pipeline requires input data from cells predominantly lacking endogenous DNA methylation (e.g., Drosophila S2 cells) as a prerequisite finding
Experimental setups
Assay System Perturbation Readout Platform
dual-enzyme single-molecule footprinting (dSMF) via paired-end bisulfite sequencing Drosophila S2 cells none CpG/GpC methylation status per DNA molecule used to call footprints and assign protein-DNA binding states (naked, TF-bound, nucleosome-bound) Illumina paired-end sequencing; Bismark/Bowtie2 alignment
STARR-seq (source data for enhancer region definition) Drosophila S2 cells none cis-regulatory enhancer peak summits used to define regions of interest
Key results
  • Three binding states (naked-DNA, TF-bound, nucleosome-bound) observed at enhancer Peak 229
  • Eight of nine possible cobinding states found at enhancer Peak 110, with MNase peaks separated by 78 bp
  • Expected proportions of chromatin states in open Drosophila S2 enhancers: naked-DNA, TF-bound, nucleosomal DNA 40%-50% / 15%-20% / 30%-40%
  • 25 DNA molecules were discarded from the heatmap in Figure 1 due to lack of valid footprints 25 molecules
  • Standard Illumina paired-end 150 bp reads yield sufficiently long DNA molecules for mapping TF and nucleosomal binding median length 269 bp
Key statistics
  • count 40%-50% (proportion of naked-DNA molecules in open Drosophila S2 enhancers)
  • count 15%-20% (proportion of TF-bound molecules in open Drosophila S2 enhancers)
  • count 30%-40% (proportion of nucleosomal DNA molecules in open Drosophila S2 enhancers)
  • other 269 bp (median length of DNA molecules formed from concatenated paired-end reads)
  • count 25 (DNA molecules discarded from Figure 1 heatmap due to no valid footprint)
  • other 78 bp (separation between the two MNase peaks analyzed in Figure 2 cobinding example)
  • count 8/9 (number of possible cobinding states observed at enhancer Peak 110)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper is a computational protocol paper (methods/workflow description) rather than a hypothesis-testing study. It presents a bioinformatics pipeline for processing dual-enzyme single-molecule footprinting (dSMF) bisulfite sequencing data to classify individual DNA molecules into protein-binding states (naked-DNA, TF-bound, nucleosome-bound) and co-binding states at pairs of sites. No inferential statistical tests are reported; results are presented as counts of DNA molecules per state visualized in heatmaps.

Replicationunclear Sample sizeNot described; the protocol uses a publicly available demo dataset (subset of Krebs et al. 2017) for illustration; sample size or power considerations are not discussed GroupsNo group comparisons; single-sample workflow demonstration on Drosophila S2 cells Pairingna Randomization/blindingnot stated Dispersionnone
Approaches that could also have been used
  • Binding states are reported as raw counts of DNA molecules per state with no uncertainty quantification
    Could also: Reporting proportions with 95% confidence intervals (e.g., Wilson or Clopper-Pearson intervals for proportions) per state would also convey the precision of the state estimates — When molecule counts are the basis for downstream biological interpretation (e.g., assessing cooperative vs. independent binding), confidence intervals would let readers judge how reliably each state proportion is estimated from the sequencing depth available at a given locus
  • Co-binding states across nine possible two-site combinations are enumerated by raw counts without a formal test of independence between the two sites
    Could also: A chi-squared test of independence or a log-linear model on the 3×3 contingency table of state combinations could also be applied to assess whether binding at one site is statistically associated with binding at the other — Formal testing of independence would allow quantification of cooperativity or anti-cooperativity beyond visual inspection of the count distribution, which is the biological question motivating the co-binding analysis
  • The footprint-calling thresholds (e.g., ≥10 bp minimum footprint length, ±15 bp TF window, 10–50 bp for TF vs. nucleosome) are fixed rule-based cutoffs with no sensitivity analysis reported
    Could also: A systematic parameter sweep or bootstrap resampling over threshold choices would also characterize how sensitive state assignments are to the chosen cutoffs — Reporting classification stability across a range of threshold values would help users understand how robustly their biological conclusions depend on these specific algorithmic choices
  • Expected state proportions (~40–50% naked-DNA, ~15–20% TF-bound, ~30–40% nucleosomal) are cited from a prior paper as qualitative benchmarks for validating a new run
    Could also: A formal goodness-of-fit test (e.g., chi-squared or multinomial likelihood ratio) comparing observed proportions to those expected benchmarks could also serve as a quantitative quality-control check — A formal comparison would make the quality-control criterion reproducible and threshold-based rather than relying on qualitative agreement, which is particularly useful when the pipeline is applied to new cell types or organisms
Software: Snakemake · Trim Galore · Bowtie2 · Bismark · Bamtools · Samtools · deepTools2 · Gnuplot · Anaconda/Python Python 3.6

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35463472

Paper: Rao S, Ramachandran S. A computational pipeline to visualize DNA-protein binding states using dSMF data. STAR Protoc 2022. PMID 35463472 / PMC9026571 / doi:10.1016/j.xpro.2022.101299.

Type: This is a STAR Protocols paper — it is a computational pipeline description. Reproducing it = running the shipped Snakemake pipeline on the shipped demo data and obtaining the protocol's example outputs (Figures 1 & 2) and the stated "expected outcomes". This is therefore an almost-ideal pipeline reproduction target (own code + shipped demo input + concrete expected figures).

Code: https://github.com/satyanarayan-rao/star_protocol_enhancer_cooperativity

  • Pinned commit: 028c14ca9a67c938f590ad2063d5fd918125aa39 (release tag v2_for_star_protocol / "STAR_PROTOCOL", 2022-01-28). Zenodo: doi:10.5281/zenodo.5914775.
  • Engine: Snakemake. Steps: Trim Galore (adapter trim) → Bismark (bisulfite-aware align to dm3) → suppress HCH/DGD methylation contexts → process BAM (merge, flag-select <99,147>/<83,163>, concatenate mate pairs into "DNA molecules") → call footprints (unmethylated stretches; min 10 bp, merge across 1 bp gaps) → build footprint/methylation matrices over a region of interest (ROI ± flanks) → assign binding states (0 naked / 1 TF-bound 10–50 bp / 2 nucleosome >50 bp / 3 discarded) → gnuplot heatmaps.
  • Languages: Python 3.6, Shell, R. Env via conda env dsmf_viz (install_required_packages.sh: snakemake, trim-galore=0.6.2, bedtools, bowtie2, bismark, samtools=1.9, pyBigWig=0.3.14, bamtools=2.4.1, pandas=0.25.3, numpy=1.16.3, tbb=2020.2, gnuplot, ghostscript, perl=5.26.0).

Data: GEO GSE77369 (Krebs et al. 2017), Drosophila S2 paired-end bisulfite (dSMF) sequencing. Repo ships a demo subset of 4 paired-end runs (SRR3133326–SRR3133329, ~20 KB each, 8 files) as data_from_geo/demo_data/, combined into sample merged_demo_S2. Reference genome dm3 (UCSC) downloaded via download_reference_genome.sh. MNase peak track shipped as bigwigs/mnase/sm.sub.comb_50_c.bigwig (28 MB). STARR-seq enhancer summits referenced externally.

In scope (pipeline-derived; will attempt)

id reported result pipeline paper location
C1 Single binding-site footprint+methylation heatmap at peak_229 (chr2L:480305), 3 states sorted by molecule length (Figure 1) full smk single-binding target Fig 1 / "Step 5", demo cmd
C2 "25 DNA molecules are discarded" from the peak_229 input matrix due to invalid edge footprints edge_cases / select_valid_states Fig 1 text
C3 Cobinding footprint+methylation heatmap at peak_110_4 & peak_110_6 (chr2L:19155173 / 19155250), spanning two sites (Figure 2) cobinding smk target Fig 2 / demo cmd
C4 The two MNase peaks are ~78 bp apart input bedpe coords Fig 2 text
C5 "8 of 9 cobinding states observed" at peak_110 assign_cobinding_states Fig 2 text
C6 "Expected" open-enhancer state distribution: ~40–50% naked-DNA, ~15–20% TF-bound, ~30–40% nucleosomal assign_binding_states over all open enhancers "Expected outcomes"
C7 Concatenated DNA molecules have median length 269 bp process_bam / mate concatenation Methods

Out of scope (not attempted — not pipeline-derived)

  • Wet-lab dSMF library prep, bisulfite conversion, Illumina sequencing (GSE77369 generation) — experimental, external.
  • The original Krebs et al. 2017 biological conclusions — different paper.

Reproduction strategy / honesty notes

  • Floor (quick minimum): run the two shipped demo Snakemake targets (merged_demo_S2, peak_229 single-site + peak_110 cobinding) and obtain the .fp.pdf + .methylation.pdf and the underlying state matrices. This reproduces the pipeline mechanics and figure format (C1, C3, C4 directly from inputs).
  • Caveat for C2/C5/C6/C7: the precise counts ("25 discarded", "8/9 states",
Figures / tables: Figure 1Figure 2
C1
Reported
Figure 1 single-site footprint+methylation heatmap at peak_229 (chr2L:480305), 3 length-sorted states + MNase track
Reproduced
Reproduced 1:1: Naked(D)=105, TF(T)=114, Nuc(N)=233 ordered blocks + MNase track, x=-150..+150
exact
C2
Reported
25 DNA molecules discarded from peak_229 input matrix
Reproduced
25 state-3 (discard) fragments (105+114+233 valid + 25 = 477 spanning)
exact
C3
Reported
Figure 2 cobinding heatmap peak_110_4 & peak_110_6 (chr2L:19155173/19155250)
Reproduced
Reproduced 1:1: 2-peak MNase track + 8 ordered pair-state blocks
exact
C4
Reported
MNase peaks ~78 bp apart
Reproduced
77 bp summit-to-summit (19155250-19155173); 78 bp end-to-start
within tolerance
C5
Reported
8 of 9 cobinding states observed at peak_110
Reproduced
8 of 9: D-D17,D-T11,T-D13,T-T24,T-N4,N-D7,N-T5,N-N16 (missing D-N)
exact
C6
Reported
~40-50% naked / ~15-20% TF-bound / ~30-40% nucleosomal (open enhancers)
Reproduced
NOT ATTEMPTED on demo: population aggregate over all open enhancers; requires full ~33GB GSE77369 S2 run + externally-defined open-enhancer set
partial
C7
Reported
median concatenated DNA-molecule length 269 bp
Reproduced
267 bp median over 1046 demo molecules (min111/max299/mean258)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 89/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

315.9 k
tokens (I/O) · 28.8 M incl. cache
54 min
runtime · 0.1 CPU-h
2.7 GB
peak RAM
1
HPC jobs
hummel
machine