Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Single-Cell Hi-C Technologies and Computational Data Analysis.

Adv Sci (Weinh) · 2025
L1 80/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 56% of all assessed papers rank 484 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the paper's central quantitative result 1:1. This is a scHi-C data-quality benchmark (Dautle & Chen 2025); its headline numeric output is Table 3 (per-dataset, per-genome average & median Total Contacts and cis/trans ratio across all public scHi-C datasets). The authors' repo (chenyongrowan/scHIC_Evaluation @3de55e3) ships the underlying per-cell master table (AllStatistics_ForBoxplot_09_01_2024.csv, 67,957 cells x 23 dataset-labels) plus the scripts that aggregate it (getStatistics.py / makeFigures.py). C-AGG (DONE): I re-aggregated the shipped per-cell CSV with the repo's OWN formula (cis/trans = Cis/(Total-Cis); drop cis/trans>=10000 outliers; group by Authors x Reference_Genome; mean/median) and compared every field of Table 3 -> 86 EXACT, 10 within-tol, 5 partial, 7 mismatch of 108 numeric fields (~89% reproduce to the printed digits; e.g. Tan2018 937763->937763.2, Mulqueen 1199765->1199764.7, Stevens 12.136->12.136, Wen 0.912->0.912, Liu 279855->279855.4). The reported table IS the rounded aggregation of the shipped data -> no fabrication evident. Honest flags (in AUDIT.md/agreement.json): (1) Tan2019 cis/trans reproduces EXACTLY but absolute totals are ~2x reported (ratio is scale-invariant -> a contact-counting/scaling definition difference, not a random error); (2) Ramani-hg19 CSV has 1896 cells vs reported 2972 and the Ramani 'Mixed' barnyard row is missing -> CSV is a partial/QC-filtered subset; (3) Luo labeled '2022' in CSV vs '2019' in paper, values differ ~20% -> ambiguous, not claimed. NOT attempted: (a) C-RAW, the one-cell raw->statistic recompute from the only shipped .hic (GSM7678878) to verify the per-cell numbers themselves derive from real raw data -- staged (run.sbatch + analyze.py with hicstraw) but BLOCKED because «our HPC» requires a browser SAML + 2FA VPN that the human operator did not action across 3 connect attempts (client also returned 'keine SSO-URL'); (b) the full from-raw recomputation of all 67,957 cells (per-cell counting code not shipped; ~TBs FASTQ across ~20 GEO series) -- the hard >>20%; (c) Figure 5 contact maps (visualization, non-deterministic random.sample). Compute note: C-AGG is a sub-second deterministic groupby over the 3MB shipped CSV (host-independent, byte-identical on «our HPC»); no data was stored on «host» (CSV streamed through memory and deleted). Verdict provisional; human audit sheet in AUDIT.md.

💻 Code ↗ 🗄 Data: GSE80006

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 80
    assessed: 2026-06-15 ⛓ 424cbfc0208c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

How can the thirteen available single-cell Hi-C (scHi-C) protocols be quantitatively evaluated and which computational analysis strategies best address the extreme data sparsity and analytical challenges of scHi-C data? This review aims to assess protocol efficiency and provide practical guidance on computational topics.

Core claims
  • Thirteen scHi-C protocols currently exist—eight capturing chromatin interactions exclusively and five combining scHi-C with other assays for multi-omics data. resource
  • Among protocols detecting chromatin interactions only, the Nagano et al. 2017 protocol and scNanoHi-C have the highest total captured contacts and cis/trans ratios. finding
  • Data quality of scHi-C protocols can be evaluated by total number of contacts recovered per cell and the cis/trans (intra- vs inter-chromosomal) ratio. method
  • Extreme data sparsity is the central computational challenge of scHi-C data, complicating clustering, compartment calling, TAD calling, loop calling, 3D reconstruction, simulation, and differential interaction analysis. finding
  • Combining scHi-C with assays such as RNA-seq, DNA methylation, or DNA-seq provides a more comprehensive understanding of chromatin architecture and its functional outcomes and improves rare-cell-type detection. finding
  • Standardization, automation, high-throughput combinatorial barcoding, reagent minimization, and alternative amplification (e.g., META) are proposed solutions to improve scHi-C protocol scalability and reduce bias. method
  • The snHi-C protocol performs about average on mouse and fruit fly cells but performs poorly on human cells (median cis/trans ratio 0.428). finding
  • Dip-C recovers an average number of contacts but yields some of the lowest cis/trans ratios among all datasets. finding
Experimental setups
Assay System Perturbation Readout Platform
scHi-C (chromatin conformation capture) mouse cells (mm9/mm10) none total chromatin contacts per cell and cis/trans ratio Nagano et al. 2013/2017 protocol (MboI digestion, PCR)
scHi-C (Stevens et al.) mouse cells (mm10) none total chromatin contacts per cell and cis/trans ratio Stevens 2017 protocol (AluI digestion, PCR)
snHi-C (single nucleus Hi-C) mouse (mm9), human (hg19), fruit fly (dm3) cells/oocytes/zygotes none total chromatin contacts and cis/trans ratio Flyamer et al. protocol (MDA whole-genome amplification)
sci-Hi-C (combinatorial indexing Hi-C) mouse (mm10) and human (hg19) cells none total chromatin contacts and cis/trans ratio Ramani et al. protocol (96-well combinatorial barcoding, PCR)
Dip-C human (hg19) and mouse (mm10) cells none haplotype-separated Hi-C maps, contacts and cis/trans ratio META (multiplex end-tagging amplification)
scSPRITE mixed mouse (mm9) and human (hg19) cells none total chromatin contacts and cis/trans ratio Arrastia et al. protocol (multiple spatial/nucleus barcoding, PCR)
scNanoHi-C mouse (mm10) and human (hg38) cells none total chromatin contacts and cis/trans ratio Li et al. protocol (tagmentation/barcoding, nanopore-based)
sn-m3C-seq (Hi-C + DNA methylation) mouse (mm10) and human (hg19/hg38) cells/brain none chromatin interactions plus DNA methylation state Lee et al. protocol (bisulfite conversion)
scMethyl Hi-C (Hi-C + methylation) mouse cells (mm9) none chromatin interactions plus methylation state Li et al. 2019 protocol (biotinylated end fill, bisulfite)
HiRES (Hi-C + RNA-seq) mouse cells (mm10) none chromatin interactions plus scRNA-seq Liu et al. protocol (reverse transcription before digestion)
MUSIC (Hi-C + RNA-seq + RNA-chromatin) human (hg38, AD and H1) and mouse (mm10) cells none chromatin interactions, RNA-seq, RNA-chromatin associations Wen et al. protocol (split-pool barcoding, 3 rounds)
s3-GCC (Hi-C + scDNA-seq) human cells (hg38) none chromatin interactions plus scDNA-seq Mulqueen et al. protocol (tagmentation, PCR)
Key results
  • Nagano et al. 2017 protocol and scNanoHi-C have the highest total captured contacts and cis/trans ratios among interaction-only protocols.
  • Stevens et al. 2017 protocol has high cis/trans ratio but recovers only about a quarter of the total contacts compared to the Nagano protocols. ~1/4 of Nagano contacts; mean cis/trans 12.136, median 13.268
  • snHi-C performs poorly on human cells with a low median cis/trans ratio. median cis/trans = 0.428
  • Dip-C datasets show average contact recovery but some of the lowest cis/trans ratios. Dip-C Tan 2018 mean cis/trans 2.552; Tan 2019 median 0.992
  • sci-Hi-C (Ramani 2017, Kim 2020) and scSPRITE exhibit high cis/trans ratios but low numbers of total contacts. sci-Hi-C mixed mean contacts 19248; scSPRITE mean 176982
  • scNanoHi-C recovers a higher-than-average number of total contacts on both mouse and human cells. mm10 mean 800199 contacts, cis/trans 10.129
Key statistics
  • count 13 scHi-C protocols (8 interaction-only, 5 multi-omics) (Total protocols reviewed and evaluated)
  • other median cis/trans ratio 0.428 (snHi-C (Flyamer 2017) on human hg19 cells, 34 cells)
  • other mean cis/trans 12.136, median 13.268 (Stevens et al. 2017 scHi-C, mm10, 8 cells)
  • count mean interactions 800199, cis/trans 10.129 (scNanoHi-C mm10, 672 cells)
  • other mean cis/trans 17.696 (Rappoport 2023 scHi-C (Nagano 2017), mm9, highest cis/trans mean)
  • other mean cis/trans 2.552, median 2.415 (Dip-C Tan et al. 2018, hg19, 35 cells, low cis/trans)
  • count mean interactions 1247004 (sn-m3C-seq Luo 2019, hg19, 4238 cells, highest contact count)
  • other mean cis/trans 0.912, median 0.195 (MUSIC hg38 (10366 AD + 2267 H1), lowest cis/trans)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a narrative review with an embedded quantitative benchmarking component. Protocol quality was assessed by downloading 13 existing public scHi-C datasets and computing per-cell descriptive statistics—mean and median total chromatin contacts and mean and median cis/trans ratios—summarized in Table 3 and visualized in Figures 3–5. No formal inferential statistical tests were applied; protocol comparisons are made descriptively by inspecting these summary values and scatter plots. No p-values, effect sizes, or confidence intervals are reported.

Replicationunclear Sample sizeNumber of cells per dataset is listed in Table 3 (ranging from 8 to 19,388 cells); no power analysis or formal sample-size justification is stated GroupsscHi-C protocols compared by mean and median per-cell total contacts and cis/trans ratio across re-analyzed public datasets Pairingna Randomization/blindingnot stated Dispersionnone Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • Protocol performance was compared by inspecting means and medians of per-cell contact counts and cis/trans ratios narratively, without formal statistical tests
    Could also: A non-parametric test such as Kruskal-Wallis with Dunn's post-hoc correction could also be applied to formally compare the distributions of per-cell values across protocols — Formal tests would quantify the probability that observed differences arise by chance and produce a structured family-wise error control, which is particularly relevant given the large variation in cell counts across datasets (8 to 19,388 cells)
  • Central tendency is summarized with mean and median; no measures of within-dataset spread are reported in Table 3
    Could also: Interquartile range (IQR), standard deviation, or 95% bootstrap confidence intervals could also be reported alongside the mean and median — Per-cell contact distributions in scHi-C data are typically right-skewed; reporting spread alongside central tendency helps readers judge whether differences between protocols are consistent across cells or driven by outliers
  • Datasets from different protocols are compared directly on raw contact counts without adjusting for sequencing depth or species/genome differences
    Could also: Rarefaction (downsampling to a common sequencing depth) or depth-normalized contact counts could also be computed before cross-protocol comparison — Differences in sequencing effort between datasets may partly explain differences in total contacts; depth normalization isolates protocol-specific capture efficiency from the amount of sequencing applied
  • Figure 5 displays per-cell cis/trans ratio as a function of total contacts in a scatter plot, but no correlation statistic is reported
    Could also: Spearman's rank correlation (or Pearson's r after log-transformation of counts) could also be computed and reported to quantify the association between sequencing depth and cis/trans ratio — A reported correlation coefficient makes the relationship quantitative and allows readers to compare the strength of this association across protocols or species
  • Protocol comparisons are made across datasets that differ in species (mouse, human, Drosophila) and reference genome simultaneously
    Could also: Stratified analysis restricted to datasets sharing the same species and reference genome, or a mixed-effects model treating species as a covariate, could also be used — Biological differences across species may confound protocol-level comparisons; stratification or covariate adjustment would allow more direct attribution of differences to the protocol itself

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
7
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE100569 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
GSE131811 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
GSE48262 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
GSE80006 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
GSE80280 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
GSE94489 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39887949

Paper: Dautle MA, Chen Y. Single-Cell Hi-C Technologies and Computational Data Analysis. Adv Sci (Weinh) 2025. PMID 39887949 · PMCID PMC11884588 · DOI 10.1002/advs.202412232. Repo: https://github.com/chenyongrowan/scHIC_Evaluation @ 3de55e3 (main, pushed 2024-11-24), GPL-3.0. Primary accession (registry): GEO GSE80006 (one of ~20 source series; the actual per-cell quality table aggregates many series — see below).

What kind of paper this is

A review + data-quality benchmark of single-cell Hi-C (scHi-C) protocols. The benchmark portion is a real bioinformatic pipeline: for every publicly available scHi-C cell across 13 protocols / 23 dataset-labels, the authors compute per-cell Total_Contacts, Cis_Contacts, Cis/Total, and Cis/Trans, then aggregate per dataset (means/medians) and visualize (Figures 2–4). One contact map figure (Fig 5) is derived from cooler files of one GEO series.

Repo contents (what is shipped)

  • CalculateStatistics/AllStatistics_ForBoxplot_09_01_2024.csvmaster per-cell table, 67,957 cells × 23 dataset-labels (the central derived dataset).
  • CalculateStatistics/getStatistics.py — aggregates the master CSV into per-dataset mean/median Total_Contacts, Cis/Total, Cis/Trans (drops Cis/Trans ≥10000 outliers).
  • FigureGenerationScripts/Figures2-4/makeFigures.py (+ same CSV) — regenerates Fig 2 (total contacts boxplots), Fig 3 (cis/trans boxplots), Fig 4 (scatterplots).
  • FigureGenerationScripts/GSM7678878_p003-bdf1_001.hicone raw .hic file (24 MB, mm10, Dip-C/LiMCA cell from GSE239969) — the only shipped raw input.
  • FigureGenerationScripts/Figure5/ — CreateFig5.py + MakeCoolerFiles.sh build cooler files from GSE129029 (Collombet et al 2020) contact lists → contact maps.

In scope (pipeline-derived, will attempt)

  • C-AGG: per-dataset summary statistics (paper Table 3 / abstract numbers). Re-run the repo's aggregation logic on the shipped master CSV → per-dataset mean & median Total_Contacts and Cis/Trans, and Cis/Total. Deterministic. Compare to the exact numbers reported in the paper (e.g. Stevens 2017 cis/trans mean 12.136 / median 13.268; Nagano 2013 10.462/10.608; Tan 2018/Dip-C total 937763, cis/trans 2.552/2.415; Mulqueen 2021 total 1199765; Ramani 2017 total 5534). This is the central quantitative result of the benchmark.
  • C-RAW: raw→statistic fabrication check (one cell). Recompute Total_Contacts and Cis_Contacts directly from the shipped raw .hic (GSM7678878, mm10) by summing genome-wide and intra-chromosomal contacts at base resolution, and check the pair appears as a row under "Liu et al 2023" (mm10) in the master CSV. Tests whether the upstream per-cell numbers (whose generation code is NOT shipped) are honestly derived from raw data vs fabricated.
  • C-FIG (optional): regenerate Fig 2/3/4 PDFs from the CSV and confirm they reproduce the shipped figure PDFs (visual + summary-stat consistency).

Out of scope (and why)

  • Full from-raw recomputation of all 67,957 cells. The per-cell statistic computation code (HiC-Pro/Juicer valid-pairs counting) is NOT shipped; reproducing it for all 23 datasets means downloading ~20 large GEO/SRA series (TBs of FASTQ/ pairs) and rerunning each protocol's mapping pipeline. This is the hard >>20% and is not attempted; we instead verify ONE cell from the one shipped raw file (C-RAW).
  • Figure 5 contact maps (Collombet GSE129029 download + cooler build) — a visualization, not a quantitative claim; the random.sample() makes it non-deterministic. Skipped (low value for fabrication detection).
  • Wet-lab / protocol-description content of the review — not computational.

Compute plan

All data on «infra», compute on «our HPC» («infra») per hard rules: clone repo on «infra» inside a compute job (front1 has no internet), build a small conda env (pandas/numpy/cooler/hic-straw), run C-AGG + C-RAW, write small result JSONs,

Figures / tables: Table
Tan2018_hg19_tot_mean
Reported
937763
Reproduced
937763.2
exact
Mulqueen2021_hg38_tot_mean
Reported
1199765
Reproduced
1199764.7
exact
Stevens2017_mm10_ct_mean
Reported
12.136
Reproduced
12.136
exact
Kim2020_hg19_tot_mean
Reported
5509
Reproduced
5508.8
exact
Wen2024_hg38_ct_mean
Reported
0.912
Reproduced
0.912
exact
Liu2023_mm10_tot_mean
Reported
279855
Reproduced
279855.4
exact
Li2023_mm10_ct_mean
Reported
10.129
Reproduced
10.129
exact
Nagano2017_mm9_tot_mean
Reported
225687
Reproduced
225686.6
exact
TABLE3_AGG_overall
Reported
Table 3: 108 numeric fields across 25 dataset-genome rows
Reproduced
86 exact / 10 within-tol / 5 partial / 7 mismatch
partial
Tan2019_mm10_tot_mean_FLAG
Reported
252292
Reproduced
496972.5 (cis/trans ratio reproduces EXACTLY at 2.594/0.992; absolute totals ~2x -> counting/scaling definition diff)
did not match
Ramani2017_hg19_ncells_FLAG
Reported
2972 cells
Reproduced
1896 cells in shipped CSV (means agree ~2%); paper's Ramani 'Mixed' barnyard row absent from CSV
partial
C-RAW_hic_recompute
Reported
per-cell Total/Cis from raw .hic (GSM7678878)
Reproduced
NOT RUN — blocked on «our HPC» VPN (Cisco SSO 2FA human-gated; client returned 'keine SSO-URL'); job staged
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

204.7 k
tokens (I/O) · 10.5 M incl. cache
16 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.