Widespread allele-specific topological domains in the human genome are not confined to imprinted gene clusters.
The main results reproduced, with only marginal, non-material deviations.
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL (honest). The authors' own HiCFlow Snakemake pipeline was reproduced end-to-end on a «our HPC» compute node, on its bundled Drosophila S2 example data (the demo the authors ship to validate the pipeline), via --use-conda (16 per-rule envs). It regenerated the complete documented output set (HiC matrices, OnTAD TADs, HiCcompare differential, viewpoint, HiCRep, insert-size, MultiQC) with concrete harvested numbers (HiCRep replicate SCC 0.963-0.978 — within the paper's RC1 range 0.97-0.98; OnTAD 11-17 TADs/sample; HiCExplorer Hi-C contacts 85-94%). Only ditagLength.svg + the optional fastQScreen contamination QC were skipped (no external genome DBs bundled; not a default target). NOT reproduced (documented blockers, not faked): the paper's headline human genome-scale numbers (99.8% genotyping / 99.9% phasing concordance, 39 conserved ASTADs, ASE/imprinted enrichment Z-scores) require regenerating EVERY HiCFlow intermediate from multi-TB GSE63525 across 3 cell lines + the AS-HiC-Analysis notebooks (which reference intermediates by external local paths and ship none) = node-weeks + multi-TB, out of bounded compute. RC1/RC2 on the paper's own PRJNA926951 RC-Hi-C is tractable in size but blocked by the missing capture-region BED.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-30no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates whether allele-specific (parent-of-origin) differences in 3D chromatin conformation are limited to classical imprinted gene clusters, or whether allele-specific topologically associating domains (TADs) exist more widely across the human genome and relate to allele-specific gene expression.
- ★ HiCFlow, a new bioinformatic pipeline, performs de novo haplotype assembly, phasing, and visualization of allele-specific (parental) chromatin conformation directly from Hi-C data without requiring pre-phased haplotypes. method
- ★ At the IGF2-H19 locus, known stable allele-specific CTCF-mediated looping interactions (ICR-dependent) are robustly detected in both RC-Hi-C and Hi-C datasets, consistent with loop-extrusion models. finding
- ★ SNRPN and DLK1 loci show more variable, less canonical allele-specific 3D structures than IGF2-H19, with allele-specific A/B compartmentalization detected instead of a single conserved imprinted structure. finding
- ★ Genome-wide unbiased ranking of TADs by allele-specific contact frequency identifies a defined set of allele-specific TADs (ASTADs), which occur in regions of high sequence variation. finding
- ★ ASTADs are enriched not only for imprinted loci but also for allele-specific expressed (ASE) genes genome-wide, including previously unreported loci such as the TAS2R bitter taste receptor gene cluster. finding
- ★ HiCFlow's Hi-C-based genotyping and haplotype phasing show high concordance (~99.8-99.9%) with an experimentally validated high-confidence haplotype in GM12878 cells. finding
- ★ Most allele-specific chromatin interactions occur within subTADs rather than defining entire TADs, and imprinted gene clusters share TADs with non-imprinted neighboring genes. finding
- ★ 8-32% of genes exhibiting allele-specific expression are located within ASTADs. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Region Capture Hi-C (RC-Hi-C) | 1-7HB2 human breast epithelial cell line | none | allele-specific chromatin contact frequency at imprinted loci (IGF2-KCNQ1, SNRPN, DLK1-DIO3) | MboI 4-base cutter restriction enzyme, tiled capture probes |
| Hi-C (haplotype phasing benchmark) | GM12878 human lymphoblastoid cell line | none | genotyping and phasing accuracy compared to high-confidence experimentally validated haplotype | prototype haplotype-phased Hi-C data |
| Region Capture Hi-C / Hi-C | IMR-90 and H1-hESC human cell lines | none | allele-specific interactions at imprinted loci | — |
| Genome-wide Hi-C / TAD analysis | human cell lines (unspecified panel) | none | TAD insulation score, identification of allele-specific TADs (ASTADs) genome-wide | — |
| Allele-specific expression (ASE) analysis | human cell lines | none | overlap of ASE genes with ASTAD locations | — |
- – Known allele-specific enhancer interactions with IGF2 and H19 promoters robustly detected as opposing A1/A2 signals in the subtraction matrix at the IGF2-H19 locus.
- – HiCFlow genotyping agreed with the high-confidence GM12878 dataset for the vast majority of common variant calls. 99.8% agreement (n=3,799,226 common loci)
- – HiCFlow haplotype phasing agreed closely with the high-confidence dataset. 99.9% agreement (n=1,942,361 informative loci)
- – SNRPN locus shows weak/poorly defined TAD boundaries with directional allelic bias in long-range associations (A2 biased leftward, A1 biased rightward).
- – DLK1 locus subTAD structure differs by allele: A1 forms a larger subTAD anchored upstream of DLK1, while A2 subTAD is anchored at the imprinting control region (ICR).
- – A defined fraction of genes with allele-specific expression are located within genome-wide-identified ASTADs. 8-32%
- – RC-Hi-C library generated high sequencing depth and coverage comparable to published high-resolution Hi-C datasets. ~40 million valid read pairs, ~1700 read pairs/kb
- other 99.8% agreement in variant identity (GM12878 HiCFlow genotyping vs. high-confidence dataset)
- other 99.9% agreement in phasing (GM12878 HiCFlow haplotype phasing vs. high-confidence dataset)
- count 4,267,624 (HiCFlow) vs 4,049,512 (high-confidence) variants identified (GM12878 genotyping comparison)
- count 2,147,688 (HiCFlow) vs 2,063,320 (high-confidence) phased variants identified (GM12878 phasing comparison)
- count 34,399 probes covering ~4.1Mb (~16.1%) of capture regions (RC-Hi-C probe design across 5 imprinted loci (25Mb total))
- count ~40 million valid read pairs, mean coverage ~1700 read pairs per kilobase (RC-Hi-C sequencing depth in 1-7HB2 cells)
- other 8-32% (proportion of allele-specific expressed genes located within ASTADs)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper describes a bioinformatic pipeline (HiCFlow) for haplotype phasing and allele-specific analysis of Hi-C/Region Capture Hi-C data, applied to imprinted gene loci and genome-wide TAD calling. Results are reported largely as descriptive comparisons (visual subtraction matrices, percent concordance between genotyping/phasing pipelines, and qualitative/enrichment observations about allele-specific TAD distributions) rather than through a dedicated inferential statistics section. Note: the supplied text is a partial excerpt (cuts off before a Methods/Statistics section), so this assessment is based only on what is shown.
-
Allele-specific TADs (ASTADs) were defined by unbiasedly ranking TADs according to their allele-specific contact-frequency differences.↳ Could also: A formal statistical/count-based test for differential chromatin interactions (e.g., paired permutation testing, or Hi-C-adapted count models such as diffHic or multiHiCcompare) could also be applied to each TAD. — This would pair the ranking with a significance threshold and false-discovery-rate control, letting readers distinguish TADs with strong statistical support from those near the top of the ranking by chance.
-
Genotyping and phasing concordance between HiCFlow and the high-confidence reference dataset was reported as single percentage values (99.8% and 99.9% agreement).↳ Could also: Reporting these proportions alongside a confidence interval (e.g., Wilson or Clopper-Pearson) or a chance-corrected agreement statistic such as Cohen's kappa would also be a standard way to summarize concordance. — A CI or kappa statistic conveys the precision of the concordance estimate given the number of loci compared, complementing the raw percentage.
-
Enrichment of allele-specific TADs in regions of high sequence variation and for allele-specific expressed genes was described narratively.↳ Could also: A formal enrichment test (e.g., Fisher's exact test, hypergeometric test, or permutation-based enrichment tools like regioneR/GREAT) with multiple-testing correction could also be used. — This would provide a quantitative estimate (odds ratio or fold enrichment) and an FDR-adjusted p-value indicating how unlikely the observed overlap is under a null model of random genomic overlap.
-
Differences between alleles at individual loci (e.g., IGF2-H19, SNRPN, DLK1) were assessed via visual inspection of subtraction matrices denoised with a median filter.↳ Could also: A quantitative, bin-pair-level statistical comparison of matched contact counts between alleles (e.g., using Hi-C-specific differential interaction tools such as diffHic, HiCcompare, or a paired non-parametric test on normalized counts) could also complement the visualization. — This would let each highlighted allelic difference be accompanied by an effect size and significance value rather than relying solely on visual/color-based interpretation of the subtraction matrix.
-
Variation in ASTAD distribution across cell lines was described qualitatively ("The ASTAD distribution varied between cell lines").↳ Could also: A formal statistical comparison of TAD-count proportions across cell lines (e.g., a chi-square or Fisher's exact test on contingency counts) could also be used. — This would give a quantitative basis (test statistic and p-value) for the claim of between-cell-line variability, in addition to the descriptive comparison.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36869353
Paper: Richer S, Tian Y, Schoenfelder S, Hurst L, Murrell A, Pisignano G. "Widespread allele-specific topological domains in the human genome are not confined to imprinted gene clusters." Genome Biol 2023. PMCID PMC9983196, DOI 10.1186/s13059-023-02876-2.
Code (MIT):
- Pipeline: https://github.com/StephenRicher/HiCFlow (Snakemake; default branch
master, last push 2024-01-26) - Downstream analysis: https://github.com/StephenRicher/AS-HiC-Analysis (branch
main, last push 2022-05-04) - Zenodo mirror: https://zenodo.org/record/7563515
Data:
- Public Hi-C: GM12878, IMR-90, H1-hESC in-situ Hi-C — GEO GSE63525 (Rao et al. 2014; 200 samples, 4.9 billion contacts in GM12878) / also distributed via 4D Nucleome.
- Generated RC-Hi-C (1-7HB2 / HB2 cells): NCBI PRJNA926951 (2 runs SRR23199640 HB2_WT2, SRR23199639 HB2_WT1; Hi-C; ~175M + ~161M read pairs).
- Auxiliary (not pipeline-derived here): CTCF ChIP (ERX115548 + ENCODE), Roadmap chromHMM 15-state, ASMdb allele-specific methylation, geneimprint.com imprinted genes, Peakachu loops (3D Genome Browser), GTEx eQTL, GWAS catalog.
Pipeline-derived results (candidate reproduction targets)
| id | result | reported value | paper loc | pipeline | feasibility |
|---|---|---|---|---|---|
| EX1 | HiCFlow bundled example (Wang 2018 Drosophila S2 cells) reproduces the documented HiC/HiCcompare/viewpoint/HiCRep/MultiQC outputs | "produces the exact figures as shown" (README) | HiCFlow README | HiCFlow (whole) | HIGH — authors' own code + bundled ~550 MB data; bounded compute |
| KT1 | Karyotype / ploidy & CNV per chromosome (GM12878/IMR90/H1hESC) | aneuploidy/CNV calls (Add. file, Fig CNVstatus) | AS-HiC 0.checkKaryotype, 0.processCNV |
nQuire + binned matrices (.bin shipped) | MED — only step whose inputs are bundled in the repo |
| PH1 | Genotyping concordance GM12878 Hi-C vs GIAB high-confidence | 4,267,624 variants called vs 4,049,512 truth; 3,799,226 common; 99.8% agreement | Results / Fig (phasing) | HiCFlow CallVariant (GATK) | LOW — needs full GM12878 raw Hi-C (multi-TB) |
| PH2 | Phasing concordance GM12878 vs GIAB | 2,147,688 phased vs 2,063,320; 99.9% over 1,942,361 loci | Results | HiCFlow Phase (HapCUT2/SNPsplit) | LOW — same blocker |
| AT1 | # conserved ASTADs across 3 cell lines (90% reciprocal overlap) | 39 conserved ASTADs; 5 imprinted genes within | Results / Fig | AS-HiC 2.summariseASTAD,8.overlapAnalysis |
LOW — needs all 3 cell lines fully processed; repo ships no intermediates |
| AT2 | ASE enrichment in ASTADs | GM12878 153/480 (32%) Z=4.64 p=1.7e-6; IMR90 73/409 (18%) Z=3.43 p=3e-4; H1 182/2398 (8%) | Results | AS-HiC 1.ASE_and_Imprinted randomisation |
LOW — same blocker |
| AT3 | Imprinted-gene enrichment in ASTADs | H1 45/115 (39%); IMR90 38 (33%); GM12878 42 (37%); p≤0.001 | Results | AS-HiC 1.ASE_and_Imprinted |
LOW — same blocker |
| RC1 | RC-Hi-C reproducibility between WT replicates | 0.97–0.98 at 5 kb (HiCRep) | Results | HiCFlow on PRJNA926951 | MED — paper's own ~20 GB data; needs capture-region BED |
In scope (attempted)
- EX1 (primary, clean 1:1 of the authors' pipeline on bundled data) — run on «our HPC».
- KT1 (stretch) — repo-bundled
.binmatrices → nQuire ploidy. - RC1 (stretch) — PRJNA926951 RC-Hi-C HiCRep, if capture BED resolvable.
Out of scope / not attempted (honest blockers)
- PH1/PH2/AT1/AT2/AT3 — the headline genome-scale numbers. Reproducing them
requires regenerating every HiCFlow intermediate (aligned BAMs, GATK VCFs,
HapCUT2 phases, SNPsplit allele matrices, OnTAD ASTAD calls) from the multi-TB
GSE63525 raw Hi-C for all three cell lines, then running the AS-HiC-Analysis
notebooks. The AS-HiC-Analysis repo references those intermediates by
local/external paths (
../../GM12878/...,/media/stephen/Elements/...) and ships none of them except the karyotype.binmatrices. Thi
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.