Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Determination of complete chromosomal haplotypes by bulk DNA sequencing.

Genome Biol · 2021
L1 97/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
97/100
Reproducibility score
1.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 92% of all assessed papers rank 80 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> 1:1 reproduction. The paper's method is mLinker (C++ Ising-spin haplotype solver integrating 10x linked-reads + Hi-C), NOT dipcall (registry code_url was a harvesting artifact; dipcall only built the NA12878 benchmark). The repo ships whole-genome final-haplotype solutions for NA12878 and RPE-1; from these I recomputed every Table-2 and Results completeness/accuracy number. Total sites (2652381), phased SNVs (2319027), final accuracies (0.997/0.996) and the scaffold completeness/accuracy values reproduce EXACTLY, including the named worst chromosomes (chr19 for NA12878 at 98.5%, chr9 for RPE-1 at 96.1%). A few completeness ratios differ by ~0.3% (within-tol) due to slightly different reference-denominator subsetting. NOT attempted: upstream extract from raw BAMs (needs ~96GB linked-reads + proprietary LongRanger), dipcall reference construction, and aneuploid cancer-genome / EGA-restricted Fig5 results (qualitative, controlled-access).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ 109ee20b1a85
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether combining bulk long-range (linked-read) sequencing with Hi-C sequencing can computationally reconstruct complete, accurate whole-chromosome haplotypes without single-chromosome isolation.

Core claims
  • A hierarchical computational strategy that first builds high-confidence local haplotype blocks from long-range/linked-read linkage and then concatenates them into whole-chromosome haplotypes using Hi-C contacts method
  • The method resolves parental chromosome haplotypes in diploid human genomes (NA12878, RPE-1) with >99% precision and >98% completeness finding
  • The approach can assemble the syntenic structure of rearranged chromosomes in aneuploid cancer genomes at base-pair resolution finding
  • Local and long-range haplotype inference are both formulated as a minimization problem solvable by steepest descent method
  • Aggregating Hi-C links between megabase-scale segments amplifies sparse long-range linkage signal enough to bridge gaps and low-variant-density regions mechanism
  • A digital karyotype of the K-562 aneuploid cancer genome was constructed from published bulk long-range and Hi-C sequencing data resource
  • Linked-reads linkage range is capped by input DNA molecule size, limiting it to local phasing finding
  • Hi-C linkage density decays as a power law with genomic distance and is predominantly intramolecular (cis) finding
Experimental setups
Assay System Perturbation Readout Platform
Linked-read (10x Genomics) sequencing RPE-1 cell line none variant calling and local haplotype phasing 10x Genomics linked-reads
Linked-read (10x Genomics) sequencing NA12878 lymphoblastoid cell line (v1 and v2 libraries) none local haplotype phasing 10x Genomics linked-reads
Hi-C sequencing RPE-1 cell line none long-range (whole-chromosome) haplotype phasing via chromatin contact linkage
Hi-C sequencing NA12878 cell line none long-range haplotype phasing
Bulk whole-genome sequencing RPE-1 cell line none variant calling
PacBio circular-consensus sequencing (CCS) long reads, low-pass (11x) RPE-1 cell line none local haplotype phasing PacBio CCS
Single-cell sequencing RPE-1 monosomic single cells none generation of reference/ground-truth haplotypes
Bulk WGS and Hi-C / cytogenetic data (aneuploid genome analysis) Aneuploid RPE-1 cells and K-562 cancer cell line aneuploidy/rearrangement (endogenous) haplotype-specific sequence coverage and Hi-C contact used to resolve chromosomal alterations and construct digital karyotype
Key results
  • Haplotype inference reproduces parental chromosome haplotypes with high accuracy and completeness in diploid genomes >99% precision, >98% completeness
  • Maximum range of molecular haplotype linkage from linked-reads data ~100kb (RPE-1), ~300kb (NA12878)
  • Fraction of Hi-C links consistent with cis linkage >90%
  • Residual distance-independent linkage signal in linked-reads data beyond molecular size, from random barcode tagging of unrelated fragments ~50% cis/trans
  • Probability that two variant sites 100kb apart are linked by Hi-C reads <10^-3
  • Haplotype solution converged within 10 rounds of iteration for all chromosomes; Chr.2 took the longest 1,000 seconds for Chr.2
  • Chosen block-switching penalty cutoffs producing high-confidence local haplotype blocks with no apparent intra-block switching error ΔE=1000 (RPE-1), ΔE=5000 (NA12878)
  • SMA region (5q13.2) segmental duplications prevent short-read resolution of variants, causing apparent low-accuracy blocks that reflect false variants rather than true switching errors ~200kb duplications, >98% sequence similarity
Key statistics
  • other >99% precision, >98% completeness (haplotype benchmarking against reference data for NA12878 and RPE-1)
  • other >90% (fraction of Hi-C links consistent with cis linkage)
  • other ~50% (residual cis/trans linkage accuracy in linked-reads data beyond molecular size, reflecting random barcode collisions)
  • other <10^-3 (probability of Hi-C linkage between variants 100kb apart)
  • count 941,518,426 reads, 60x mean depth (RPE-1 linked-reads sequencing data)
  • count 486,848,169 reads, 91,428,507 contacts (>1Mb) (NA12878 Hi-C sequencing data)
  • other ΔE=1000 (RPE-1), ΔE=5000 (NA12878) (chosen block-switching penalty cutoffs for high-confidence haplotype blocks)
  • other 1,000 seconds (convergence time for Chr.2 local haplotype inference, longest among all chromosomes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational methods paper introducing a two-stage haplotype phasing algorithm that combines linked-read and Hi-C sequencing data. Validation was performed by benchmarking the inferred haplotypes of two diploid human genomes (NA12878 and RPE-1) against independent reference datasets using accuracy metrics (fraction of genotypes consistent with the majority haplotype per block, N50 block length, and overall precision/completeness). The paper relies exclusively on descriptive performance metrics rather than formal inferential hypothesis tests, consistent with the algorithmic benchmarking context.

Replicationtechnical Sample sizeTwo diploid genomes (NA12878 and RPE-1) used as validation samples; sample size is not justified by power analysis — the samples were chosen based on availability of reference haplotype data GroupsComputationally inferred haplotypes vs. independent reference haplotypes (GIAB consortium VCF; diploid PacBio de novo assembly; single-cell monosomy-derived RPE-1 haplotypes) Pairingna Randomization/blindingna Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Intra-block phasing accuracy (max(f, 1-f) where f = fraction of phased genotypes agreeing with reference) Benchmarking of local haplotype blocks vs. GIAB and diploid-assembly reference haplotypes (Fig. 4, Chr.5 and genome-wide) not stated
Fraction of Hi-C links consistent with cis linkage (descriptive proportion) Assessment of Hi-C and linked-read molecular haplotype linkage quality across genomic distances (Fig. 2e, f) not stated
N50 haplotype block length (descriptive summary statistic) Comparison of haplotype block continuity across switching-penalty cutoffs and samples (Fig. 4c) na
Overall precision and completeness (>99% and >98%) against reference haplotypes Top-level benchmark of whole-chromosome haplotype inference for NA12878 and RPE-1 not stated
Average number of Hi-C links between genomic segments at varying distances (descriptive mean) Quantifying Hi-C linkage signal between 0.5, 1, and 2 Mb segments (Fig. 3a) not stated
Approaches that could also have been used
  • Phasing accuracy is reported as a single point estimate (e.g., >99% precision) without any measure of uncertainty
    Could also: Bootstrap resampling over chromosomal segments or binomial confidence intervals around the accuracy proportion could also be reported — Uncertainty estimates around the benchmark accuracy would allow readers to assess how stable the performance metric is and to formally compare it against competing methods at a specified confidence level
  • Validation was performed on two genomes (NA12878 and RPE-1) selected by data availability rather than a systematic sample
    Could also: Leave-one-chromosome-out or k-fold cross-validation across chromosomes within each genome could also be used to estimate generalization error — Cross-validation would provide an internal estimate of variance in accuracy across genomic regions and reduce dependence on the specific choice of benchmark genome
  • The N50 block length and accuracy are compared across switching-penalty cutoffs visually (Fig. 4c) without a formal quantitative comparison
    Could also: A receiver operating characteristic (ROC) or precision-recall curve plotting accuracy against completeness across the full range of cutoff values could also be constructed — A ROC/PR curve would make the accuracy-completeness tradeoff across all cutoffs explicit and would enable selection of an optimal operating point using a principled criterion such as the F1 score
  • The decay of Hi-C linkage density with genomic distance is described as following a power-law (Fig. 2a, b) based on visual inspection
    Could also: A formal log-log linear regression with goodness-of-fit statistics (R², residual plot) could also be used to quantify how well the power-law model fits the observed decay — Quantifying the fit would allow the degree of power-law behavior to be compared across datasets or against competing decay models (e.g., exponential), and would surface any deviations at specific distance ranges
  • Performance of the method is characterized by a single threshold-based binary classification (≥98% accuracy = high-confidence block, colored gray; <98% = low-confidence, colored red)
    Could also: A continuous scoring approach using the block-switching penalty score as a calibrated confidence measure, reported alongside the accuracy estimate, could also be used — A calibrated confidence score would give downstream users a probabilistic handle on block reliability rather than requiring a fixed accuracy threshold, enabling more flexible downstream use
  • Benchmarking uses two reference datasets for NA12878 (GIAB and diploid assembly) and one for RPE-1 (single-cell monosomy), but concordance between reference datasets is not formally quantified
    Could also: Cohen's kappa or percent raw agreement between the two NA12878 reference haplotypes could also be reported before using them as ground truth — Quantifying inter-reference agreement would bound the ceiling on evaluable accuracy: if the two reference datasets disagree at 0.5% of sites, a measured accuracy of 99.5% cannot be distinguished from the reference noise floor
Software: Custom Python/algorithmic implementation (steepest descent minimizer for haplotype inference) · 10X Genomics linked-reads pipeline (implied by data source references) · hifiasm (used for diploid de novo assembly of NA12878 reference haplotype, cited as [40]) r253 · dipcall (used to call phased variants from diploid assembly reference)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Table
C1
Reported
2652381 total SNV sites
Reproduced
2652381
exact
C2
Reported
2319027 SNVs phased (mLinker)
Reproduced
2319027
exact
C3
Reported
final accuracy vs GIAB 0.997
Reproduced
0.99651 (comparable 1824401 exact)
exact
C4
Reported
final completion vs GIAB 0.980
Reproduced
0.97984
exact
C5
Reported
final accuracy vs diploid-asm 0.996
Reproduced
0.99637 (comparable 2122256 exact)
exact
C10
Reported
scaffold completeness vs dipAsm 2037593/2312059=88.1%
Reproduced
2037593/2312059=88.13%
exact
C11
Reported
scaffold accuracy 99.6%, worst chr19 98.5%
Reproduced
99.67-99.75%, worst chr19 98.3-98.5%
exact
C12
Reported
RPE-1 scaffold completeness 2071147/2320153=89.3%
Reproduced
2071147/2320153=89.27%
exact
C13
Reported
RPE-1 scaffold accuracy 98.3%, worst chr9 96.1%
Reproduced
98.34%, worst chr9 96.10%
exact
C6
Reported
final completion vs dipAsm 0.969
Reproduced
0.97212
within tolerance
C9
Reported
scaffold completeness vs GIAB 93.5%
Reproduced
93.8% (1752014)
within tolerance
C16
Reported
chr21 end-to-end shipped Final-Haplotype-Solution
Reproduced
PENDING (binary rerun in progress)
m.public.grade.pending

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 97/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

224.1 k
tokens (I/O) · 9.9 M incl. cache
43 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.