Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

MoDLE: high-performance stochastic modeling of DNA loop extrusion interactions.

Genome Biol · 2022
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

IN PROGRESS — building MoDLE (bioconda) on «our HPC» and running the documented genome-wide hg38 example simulation. Will report .cool structure + runtime vs paper's performance claims. OpenMM comparison (rho=0.93) and full Fig5 ENCODE pipeline are out of scope (heavy). Profiling GSE90994 in same pass.

💻 Code ↗ 🗄 Data: GSE90994

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 80731763fc44
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can DNA loop extrusion interactions be modeled stochastically with high performance to simulate realistic genome-wide contact patterns orders of magnitude faster than existing molecular dynamics approaches?

Core claims
  • MoDLE is a high-performance stochastic model that simulates DNA-DNA contacts from loop extrusion genome-wide in minutes using less than 1 GB of RAM resource
  • MoDLE accurately simulates contact maps in concordance with OpenMM molecular dynamics simulations and with Micro-C data finding
  • MoDLE is orders of magnitude (4000–5000 times) faster than OpenMM at simulating loop extrusion contacts finding
  • MoDLE models LEF binding/release/extrusion as an iterative stochastic process and extrusion barriers (e.g. CTCF sites) as a two-state Markov process method
  • Altering MoDLE parameters to mimic in silico CTCF and WAPL depletion reproduces expected loss of TAD insulation and more pronounced stripe/dot patterns finding
  • MoDLE runs efficiently on hardware ranging from laptops to high-performance computing clusters, scaling close to theoretical optimum with CPU cores finding
  • Bayesian optimization with Gaussian processes can optimize CTCF binding kinetics parameters to maximize similarity with Micro-C data method
Experimental setups
Assay System Perturbation Readout Platform
Stochastic loop extrusion simulation (MoDLE) Human genome (GRCh38) / H1-hESC-derived barriers none (wildtype) and in silico CTCF/WAPL depletion, LEF density/processivity/collision alterations Simulated DNA-DNA contact frequencies (cooler format) MoDLE CLI
Molecular dynamics simulation (OpenMM, 1D+3D polymer) Synthetic genome regions (1–500 Mbp) and 10 Mbp regions on five chromosomes, H1-hESC input none Simulated contact maps including random polymer contacts OpenMM (CPU on server C; GPU on server D, NVIDIA Tesla P100)
CTCF ChIP-seq (input data) H1-hESC cells none CTCF barrier positions and occupancy/binding probabilities
RAD21 (cohesin) ChIP-seq (input data) H1-hESC cells none Cohesin/LEF binding signal
Micro-C / Hi-C (reference contact data) H1-hESC cells none Experimental contact frequency matrices for comparison and pixel-pattern classification
Execution time and memory benchmarking Synthetic genomic regions 1–500 Mbp; full human genome (3088 Mbp) varying CPU cores (1–52), multithreading on/off Median wall clock time, CPU time, peak memory usage Servers A–D, Laptop A (Table 1)
Key results
  • MoDLE output and OpenMM output correlate strongly Pearson ρ = 0.93
  • MoDLE achieves similar median pixel accuracy to OpenMM in reproducing stripe/dot patterns vs Micro-C 0.69 (MoDLE) vs 0.68 (OpenMM)
  • MoDLE is dramatically faster than OpenMM across compared genome regions 4000–5000x faster
  • Genome-wide human simulation with default settings completes quickly generating >370 million contacts ~40 s (server A) / ~5 min (laptop A); >370 million contacts
  • MoDLE strong-scaling: human genome simulation time drops as CPU cores increase from 1 to 52 1 h 21 min (1 core) to 1 min 48 s (52 cores)
  • OpenMM execution times are very long, requiring GPUs in practice 2 h 35 min (smallest region) to >41 h (250 Mbp); up to 35 h 20 min CPU below 5 Mbp
  • MoDLE simulates 500 Mbp quickly with multithreading vs single thread ~1 min (52 cores) vs ~12 min (1 core)
  • In silico CTCF depletion loses TAD insulation while WAPL-depletion mimic strengthens stripe/dot patterns
Key statistics
  • correlation Pearson ρ = 0.93 (MoDLE vs OpenMM contact output correlation)
  • other median pixel accuracy 0.69 (MoDLE), 0.68 (OpenMM) (correct classification of dot/stripe pixels vs Micro-C/Hi-C)
  • other 4000–5000 times faster (MoDLE speedup over OpenMM)
  • count 38,815 CTCF barriers and 61,766 LEFs (genome-wide human run barriers from H1-hESC)
  • count >370 million contacts (contacts generated in genome-wide default human run)
  • other ~40 s (server A); ~5 min (laptop A) (genome-wide human simulation runtime)
  • other 1 h 21 min to 1 min 48 s (human genome simulation time from 1 to 52 CPU cores)
  • count over 500 simulation instances per chromosome; <10 MB memory per instance (default parallel simulation instances and per-instance memory)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational methods/software paper (MoDLE, a stochastic simulator of DNA loop extrusion) rather than a hypothesis-testing biology study. Validation was performed by comparing MoDLE's simulated output to Micro-C data and to an existing molecular-dynamics tool (OpenMM) using descriptive and correlational metrics (e.g., Pearson correlation, median pixel classification accuracy) and by benchmarking runtime/memory with repeated runs summarized as medians. Parameter optimization was performed using Bayesian optimization with Gaussian processes rather than classical inferential statistics.

Replicationtechnical Sample sizeBenchmarking runs were repeated a stated number of times (10 for MoDLE, 5 for OpenMM); simulated comparison regions were five 10-Mbp regions on five chromosomes; no formal power analysis described GroupsMoDLE simulated output vs. OpenMM (molecular dynamics) simulated output vs. Micro-C/Hi-C experimental data Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Pearson correlation comparison of MoDLE vs OpenMM simulated contact matrices (overall and diagonal-by-diagonal) five 10-Mbp regions across five chromosomes not stated
Median pixel classification accuracy (descriptive proportion, no significance test reported) comparison of MoDLE and OpenMM output against Micro-C dot/stripe patterns (Fig. 2E) modeled regions in H1-hESC cells na
Median elapsed wall-clock time and peak memory usage (descriptive summary statistics, no formal hypothesis test reported) runtime/memory benchmarking of MoDLE vs OpenMM across genome region sizes (Fig. 3) 10 repeated measurements for MoDLE, 5 for OpenMM per condition na
Bayesian optimization using Gaussian processes optimization of CTCF binding kinetics parameters (P_UU, P_BB) to maximize similarity to Micro-C data not stated
Approaches that could also have been used
  • Similarity between MoDLE and OpenMM outputs was summarized with a single Pearson correlation coefficient (ρ = 0.93).
    Could also: Reporting a confidence interval around the correlation coefficient, or a rank-based correlation (e.g., Spearman) alongside Pearson's, could also be used — A CI would convey the precision of the estimated correlation, and a rank-based measure is often used as a complement when the linearity of the relationship between contact-matrix values is uncertain.
  • Benchmarking measurements were repeated a fixed number of times (10 for MoDLE, 5 for OpenMM) and summarized with medians.
    Could also: Reporting the spread of these repeated measurements (e.g., interquartile range, min-max, or SD) alongside the median could also be used — Showing dispersion in addition to a central-tendency estimate helps convey the variability/reproducibility of runtime and memory measurements across repeated runs.
  • Pixel classification accuracy for MoDLE versus OpenMM was compared descriptively (0.69 vs. 0.68) without a formal statistical test.
    Could also: A paired comparison test (e.g., paired t-test or Wilcoxon signed-rank test across the compared regions) or a bootstrap confidence interval on the accuracy difference could also be used — Such an approach would let readers assess whether the small observed difference in accuracy is distinguishable from measurement variability across the compared genomic regions.
  • Optimization of CTCF binding parameters used Bayesian optimization with Gaussian processes to minimize an objective function.
    Could also: A grid search or classical gradient-based optimization with reported convergence diagnostics could also be used — These alternatives can offer more transparent, reproducible convergence behavior and are sometimes preferred when the objective function is cheap enough to evaluate exhaustively.
  • Comparisons between MoDLE and OpenMM were framed largely in terms of overall similarity of simulated contact patterns.
    Could also: A quantitative distance/similarity metric with an associated null distribution (e.g., permutation test comparing observed similarity to similarity expected by chance) could also be used — This would allow a formal statistical statement about whether the observed agreement between methods exceeds what would be expected under a null model.
Software: OpenMM · MoDLE · cooler (output file format)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36451166 (MoDLE)

Paper: Rossini, Kumar, Mathelier, Rognes, Paulsen. MoDLE: high-performance stochastic modeling of DNA loop extrusion interactions. Genome Biol 2022. DOI 10.1186/s13059-022-02815-7.

Tool under test: MoDLE — authors' own C++ tool (P16: own code). Repo: https://github.com/paulsengroup/modle (bioconda: modle, latest v1.1.0). Full paper pipeline: https://github.com/paulsengroup/2021-modle-paper-001-data-analysis (Nextflow + Singularity; fetch_data → ENCODE chip-seq-pipeline2 → preprocess → benchmarks/optimization/heatmap).

What MoDLE does

Stochastic simulation of DNA loop-extrusion contacts genome-wide. Inputs: a chrom.sizes file + a BED6+ extrusion-barrier file (occupancy 0–1, strand). Output: a .cool contact matrix (5 kbp bins by default), a .log, a _config.toml, and a _lef_1d_occupancy.bw LEF-occupancy track. Ships example data examples/data/hg38.chrom.sizes + hg38_extrusion_barriers.bed.xz.

IN SCOPE (pipeline-derived, attempt)

id result how difficulty
C1 MoDLE installs/builds and runs the documented example simulation, producing a valid genome-wide .cool (5 kbp) + the 4 documented output files bioconda install on «our HPC»; modle simulate on shipped hg38 example data FLOOR
C2 Output cooler structure: 5 kbp resolution, correct chrom set/bin counts vs hg38.chrom.sizes, non-trivial contact count inspect .cool with cooler/h5 FLOOR
C3 Runtime for genome-wide human simulation (paper: ~40 s server A many-core, ~5 min laptop A 1-core). Honest 1:1 on «our HPC» EPYC, NOT same hardware time the run on N cores + 1 core FLOOR (hardware-different)
C4 Determinism/reproducibility of MoDLE: same seed+threads → identical/near-identical contacts; default stochastic run summary stats stable run twice, compare pixels STRETCH
C5 Advanced "deep" example (--ncells 16 --target-contact-density 10) runs and yields denser matrix second simulate run STRETCH
C6 Single-chromosome memory footprint (<120 MB chr1 per paper) — observe peak RSS «path» -v STRETCH

DATASET to profile

  • GSE90994 — ChIP-seq CTCF + RAD21 in JM8.N4 mouse ES cells. Paper uses 6 SRA runs SRR5085152–SRR5085157 (EBI/SRA mirror), processed via ENCODE chip-seq- pipeline2 → fold-change-over-control → extrusion-barrier annotation for the HoxD Fig 5 simulation. Profile: N reported vs observed, data type, completeness, QC.

OUT OF SCOPE (not attempted, or only noted)

  • OpenMM molecular-dynamics comparison (Fig 2; Pearson ρ=0.93 MoDLE vs OpenMM; "4000–5000× faster"): requires the full OpenMM polymer MD benchmark harness — a large separate simulation campaign. Not attempted (heavy; out of FLOOR/STRETCH).
  • Fig 5 HoxD end-to-end (ENCODE chip-seq-pipeline2 on GSE90994 → barriers → sim → comparison): the ENCODE pipeline is a multi-day GPU/CPU campaign. We profile GSE90994 but do not re-run the full ENCODE pipeline.
  • All wet-lab / experimental Micro-C/Hi-C generation (4DN datasets) — external.

Honesty notes

  • MoDLE is stochastic; exact pixel values depend on RNG seed AND thread count (per docs). Runtime claims are hardware-specific → our times are a DIFFERENT machine, reported as such, never graded "exact".
  • Headline biological-accuracy claim (ρ=0.93 vs OpenMM) is NOT reproduced here.
C1
Reported
MoDLE simulate runs on shipped hg38 example, emits .cool + 3 sidecar files
Reproduced
partial
C2
Reported
5 kbp cooler, genome-wide hg38 bins, non-trivial contacts
Reproduced
partial
C3
Reported
human genome sim ~40s (server, many cores) / ~5min (laptop, 1 core)
Reproduced
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

This is a software/performance paper (MoDLE loop-extrusion simulator) with fully open code (bioconda/github) and a shipped hg38 example, so there is no fabrication concern and the claims are derivable in principle. However the reproduction is incomplete: status is partial/IN PROGRESS, agreement.json is not-run-yet, and all reproduced values are blank, while the quantitative validations (OpenMM rho=0.93, Fig5 ENCODE pipeline) were declared out of scope on our side. The deviation, such as it is, sits on our methodology (scope reduction + unfinished run), not the authors'. Net: a solid, low-risk setup graded yellow across the board — nothing contradicts the central claim, but nothing is yet confirmed end-to-end.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

78.2 k
tokens (I/O) · 4.8 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.