iceDP: identifying inter-chromatin engagement via density peaks clustering algorithm.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH TO RUN, and the pipeline reproduces 1:1 as software. iceDP is a small self-contained Python (pandas/numpy/scipy) density-peaks NHCC caller. We cloned it (commit 985cba3) onto «infra» and ran the exact documented procedure on the repo's own shipped demo data (play_data/chr4_chr11.txt, 1,750,743 chr4xchr11 mm10 contacts) with shipped default params on «our HPC» SLURM 2176247 (COMPLETED, ~5.4 min, 4 cpus, py3.7.12/pandas1.3.5). It runs end-to-end, is deterministic, and regenerates the documented 15-column output table -> C2 EXACT. WHAT IS NOT 1:1: the repo's OWN documented worked example -- the reference spot at 72275000/101650000 shipped as image/bin1_72275000_bin2_101650000.png -- is NOT reproducible from the shipped data (C1 MISMATCH). That genomic locus is near-empty in the shipped file (3 dots) whereas the shipped figure shows a dense field with a ~15-dot red density-peak there; a peak cannot arise from 3 dots. The example figure was evidently made from a different, denser input than the file shipped, and the README even names play_data/chr4_chr11_mm10.txt which does NOT exist in the repo (it ships chr4_chr11.txt). Flagged as a data/figure inconsistency / possible fabrication for human review (see AUDIT.md + artifacts/2026-06-14_icedp_refspot_sidebyside.png). NOT ATTEMPTED (hard 20%, out of scope): the journal paper's biology/benchmark numbers (NHCC counts in OSNs/mESCs, diffHiC & FitHiC comparison, Pol II/H3K27ac enrichment) which require downloading and fully reprocessing GSE230332 + Table S3 Hi-C/Micro-C/HiChIP datasets with per-dataset parameters not specified in the methods.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 53assessed: 2026-06-14 ⛓ f6cc13a605ec
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusExisting Hi-C analysis tools cannot reliably identify non-homologous inter-chromatin contacts (NHCCs); can a density-based clustering approach (Density Peaks) specifically and accurately detect these inter-chromosomal interactions while filtering out false positives?
- ★ iceDP is a tool that uses the Density Peaks clustering algorithm plus a distribution test and fold-change filter to identify non-homologous inter-chromatin contacts (NHCCs) from Hi-C data method
- ★ iceDP accurately identifies known NHCCs, including olfactory receptor gene hubs in mature olfactory sensory neurons and Polycomb-regulated developmental genes in mESCs finding
- ★ iceDP uncovers previously unreported transcriptionally active NHCCs finding
- ★ iceDP outperforms diffHiC and FitHiC, exhibiting the highest positive rate finding
- ★ iceDP is compatible with multiple chromatin conformation capture techniques including in-situ Hi-C, Micro-C, HiChIP, and BL-HiC resource
- NHCCs appear as local high-density regions in Hi-C contact matrices, making density-based clustering suitable for their detection mechanism
- Default cutoffs of delta = 1x10^6 and rho = 30 are appropriate for Hi-C data based on systematic parameter analysis method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| in-situ Hi-C | mouse olfactory sensory neurons (mOSN), mESC/ES, NPC, PSC, human islet, TC-797 cells | none | inter-chromatin contact frequency / NHCC identification | — |
| H3K27ac HiChIP | mOSN, extraembryonic cells, Jurkat cells | none | H3K27ac-associated chromatin contacts | — |
| Micro-C | H1 hESC and JM8-N4 cells | none | chromatin contact frequency | — |
| BL-HiC | DM WT and GM WT | none | chromatin contact frequency | — |
| PolII ChIP-seq | (cell type per GSE52071) | none | RNA Pol II occupancy | — |
| H3K27ac ChIP-seq and ATAC-seq | (cells per GSE90895) | none | H3K27ac occupancy / chromatin accessibility | — |
| H3K4me3 ChIP-seq | multiple datasets (GSE32218, GSE39513, GSE31039, GSE23943, GSE60749, GSE62380) | none | H3K4me3 occupancy | — |
| Aggregate peak analysis (APA) | Hi-C contact data | none | peak score (central pixel intensity vs lower-left corner average) | Juicer tools v1.11.09 |
- – iceDP accurately identified known NHCCs (OR genes in mOSN; Polycomb-regulated developmental genes in mESCs)
- – iceDP discovered novel previously unreported transcriptionally active NHCCs
- ▲ iceDP achieved the highest positive rate compared to diffHiC and FitHiC
- – >99% of density (delta) points fall below 1x10^6 across tested delta values (1x10^4 to 1x10^8) >99%
- – Cluster center counts stabilize when rho > 30 (tested rho 5 to 50 in steps of 5)
- other delta = 1x10^6 (default cutoff) (default delta cutoff for Hi-C data)
- other rho = 30 (default cutoff) (default rho cutoff; cluster center counts stable above this)
- count >99% (density points below delta = 1x10^6)
- pvalue P > 10^-7 (chi-square test threshold for true positive definition)
- fold_change FC1 > 2 (fold change criterion for true positive sites)
- count >15 (interaction count within central region for true positive)
- other |FC2 - FC3| < 1.5 (true positive criterion for difference between fold change values)
- other default fold change threshold = 2 (fold change filter to distinguish local hot spots from cluster centers)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
iceDP is a computational tool that applies the Density Peaks clustering algorithm to Hi-C inter-chromosomal contact matrices to nominate candidate non-homologous chromosomal contact (NHCC) sites. Candidate cluster centers are filtered through a chi-square goodness-of-fit test (to remove strip-pattern artifacts) and fold-change ratio filters (FC1–FC3, to remove non-focal high-density regions), and interaction strength is quantified via a negative binomial test incorporated into a composite interaction score. Tool performance was benchmarked against diffHiC and FitHiC2 using a positive rate metric derived from dataset-specific true-positive region definitions (mESC H3K27me3 domains, mOSN olfactory-receptor gene clusters, TC-797 H3K27ac mega-domains), with aggregate peak analysis (APA) performed using Juicer tools v1.11.09.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Chi-square goodness-of-fit test | Distribution test step: each candidate cluster center subdivided into a 3×3 grid; tested against expected frequency patterns {1,1,1,1,2,1,1,1,1} and {1,1,1,1,1,1,1,1,1}; larger P-value of the two reported; P < 0.05 used to filter strip-pattern false positives; P > 10^-7 additionally required as one of four true-positive classification criteria | — | not stated |
| Negative binomial test | Interaction score calculation: tests whether contact counts within a 5× extended interaction rectangle exceed levels expected under a negative binomial distribution; -log10(p-value) used as numerator of the composite interaction score | — | not stated |
| Fold-change threshold filters (FC1, FC2, FC3) | Fold change filter step: density ratio of central rectangle to vertical (FC1), horizontal (FC2/FC3), and larger surrounding rectangles; default threshold of 2 for FC1; |FC2 - FC3| < 1.5 used as true-positive criterion | — | na |
| Aggregate peak analysis (APA) peak score | Validation of identified NHCCs; ratio of central pixel intensity to mean intensity of lower-left corner pixels, computed via Juicer tools v1.11.09 | — | not stated |
-
Chi-square tests are applied independently to each candidate cluster center with no correction for the large number of comparisons made genome-wide↳ Could also: Benjamini-Hochberg FDR correction applied across all candidate cluster centers tested simultaneously could also control the expected false discovery rate — When testing many candidate regions across an entire inter-chromosomal contact matrix, per-family FDR control provides a principled framework for calibrating the proportion of false positives among retained candidates; the choice between per-test thresholds and FDR reflects a design trade-off between sensitivity and specificity that can be made explicit
-
The chi-square goodness-of-fit test on a 3×3 grid of counts is used to detect strip-pattern artifacts around each cluster center↳ Could also: A multinomial likelihood ratio test (G-test) or exact multinomial test could also assess spatial uniformity of contact point distributions in the 3×3 grid — G-tests and exact multinomial tests can perform better than chi-square approximations when expected cell counts are low, a common situation in sparse inter-chromosomal contact matrices; exact tests avoid the large-sample approximation requirement entirely
-
Contact counts within interaction rectangles are modeled with a negative binomial distribution for the interaction score↳ Could also: A zero-inflated negative binomial model or Poisson regression could also model sparse Hi-C count data, as implemented in tools such as HiC-DC+ or FitHiC2 — Zero-inflated models explicitly account for the excess of zero-count bins that is characteristic of inter-chromosomal contact matrices; Poisson models are a simpler alternative when overdispersion is limited and can be tested against the NB to select the better fit
-
Candidate NHCCs are identified using the Density Peaks algorithm with fixed thresholds on ρ (≥30) and δ (≥10^6) chosen by systematic parameter sweeps on one dataset↳ Could also: DBSCAN or kernel density estimation (KDE)-based peak calling could also identify local high-density regions in the 2D contact map — DBSCAN does not require pre-specification of cluster count and adapts to varying local density levels; KDE provides a continuous density surface whose threshold can be calibrated consistently across datasets with different sequencing depths
-
Tool performance is compared to diffHiC and FitHiC2 at a single operating point using a positive rate metric↳ Could also: Precision-recall curves or ROC curves computed across a range of score thresholds, summarized as AUPRC or AUROC, could also characterize and compare tool performance — A single positive rate at a fixed threshold does not capture performance across the full score spectrum; AUPRC is particularly informative for imbalanced settings where true NHCCs are rare relative to background contacts, and threshold-free metrics allow fairer comparison between tools with different default parameters
-
True positive regions are defined by dataset-specific epigenomic features (H3K27me3 domains, OR gene clusters, H3K27ac mega-domains) as surrogate ground truth↳ Could also: Cross-validation against orthogonal experimental measurements such as DNA-FISH co-localization, SPRITE, or split-pool barcoding could also serve as ground truth for benchmarking — Epigenomic-feature-based ground truth ties the definition of true positives to the biological signal being studied, which can circularly favor methods sensitive to that signal; orthogonal, technology-independent measurements provide a benchmark that is agnostic to the Hi-C data being evaluated
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41499218 (iceDP)
Paper: iceDP: identifying inter-chromatin engagement via density peaks clustering
algorithm. Chen R, Chen J, Shi L, He J. Brief Bioinform 2026. DOI 10.1093/bib/bbaf704.
Code: https://github.com/JiekaiLab/iceDP (AGPL-3.0, public, not archived).
Commit pinned: 985cba3903171701cdbe94d4a5e48a00d00cd318 (2025-05-19).
Data (paper): GSE230332 + others (Table S3) — full Hi-C/Micro-C/HiChIP datasets.
What the tool is
A small self-contained Python package (pandas/numpy/scipy/matplotlib) implementing
a Density-Peaks-clustering caller for non-homologous inter-chromatin contacts
(NHCCs) from a 3-column genome-interaction table (start1, start2, counts). The
repo ships its own demo input play_data/chr4_chr11.txt (mouse mm10, chr4×chr11,
~40 MB) AND a reference output image image/bin1_72275000_bin2_101650000.png
whose filename encodes a detected spot at start1=72275000, start2=101650000 (the
README worked example plot_one_spot(x.data_filted2.values[1], x)).
In scope (pipeline-derived, low-hanging 80%)
- R1 — Demo-data spot detection (PRIMARY). Run the documented iceDP procedure
(
readData → get_rho → get_delta → do_chi_square_test → define_border → horizontal_and_vertical_fold_change → save_reult) on the SHIPPEDplay_data, with the SHIPPED default parameters (dc=150000, window_size=600000, n_cpu=4, min_rho=30). Expected, derivable purely from shipped data+code: the detected interaction-spot set (data_filted2) contains a spot at (72275000, 101650000) — the spot shown in the repo's reference image. Grade = exact if present. - R2 — Output table reproduction (auditable artifact). The 15-column DPresult table is regenerated; we report N spots and the full small table for hand audit.
- R3 — Reference-figure reproduction (optional/visual). Re-render the
plot_one_spotfigure for the target spot and compare to the shipped PNG.
Out of scope (hard ~20%, NOT attempted — and why)
- The paper's headline biology/benchmark numbers: NHCC counts in mature OSNs / mESCs, the diffHiC & FitHiC comparison ("highest positive rate"), Pol II/H3K27ac enrichment, applicability across Micro-C/HiChIP/BL-HiC. These require downloading and fully reprocessing the GSE230332 (+Table S3) Hi-C datasets to per-chromosome contact tables, with parameters/thresholds not fully specified per dataset, and external ChIP/ATAC overlays. Large compute + under-specified — deliberately skipped per the 80/20 rule. We reproduce that the published, shipped pipeline runs and regenerates its own documented example output 1:1.
Compute plan
All compute on «our HPC» SLURM (account kubisch_std, partition std, cpus-per-task=4, no --mem). Repo + 40 MB demo data fetched onto «infra» inside the compute job; only small result tables/figures pulled back to «host». conda prefix-env on «infra» (python 3.7 to match the shipped .pyc, + pandas/numpy/scipy/matplotlib).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The iceDP tool itself reproduces 1:1 as software — deterministic end-to-end on the shipped demo data with shipped defaults, regenerating the documented 15-column DPresult table (C2 exact). However, the repo's own documented worked example — the reference spot at 72275000/101650000 shipped as image/bin1_72275000_bin2_101650000.png — is not derivable from the shipped data: that locus holds only 3 dots (rho ~3 << min_rho=30) versus the figure's dense ~15-dot red peak, and the README even names an absent chr4_chr11_mm10.txt. This sits on the authors' side as a data/figure inconsistency flagged as possible fabrication; the paper's central NHCC-biology claims were out of scope (full GSE230332 reprocessing) so could not be confirmed, leaving overall confirmation limited and the case critical.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.