Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Landscape of allele-specific transcription factor binding in the human genome.

Nat Commun · 2021
82/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
82/100
Reproducibility score
0.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 60% of all assessed papers rank 459 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

ADASTRA (Abramov et al., Nat Commun 2021) Part-A SNP-calling reproduced on the assigned accession SRR8652105 (CCLE MCF7 deep WGS), targeting the one self-contained Methods sub-analysis that uses this exact accession: 'Construction of an independent BAD map for MCF7 cells', which reports 5 explicit per-stage counts (C1-C5). The full faithful chain executed to completion on «our HPC» (SLURM/«infra», «job»-2208547): bowtie2 2.4.5 default settings -> UCSC hg38 analysisSet (no-alt) -> samtools sort -> PICARD MarkDuplicates(REMOVE_DUPLICATES=true) -> MAPQ>=10 accounting -> GATK 4.1.9.0 BQSR -> ApplyBQSR -> HaplotypeCaller(--dbsnp, scatter by chr) -> MergeVcfs -> bcftools counts. RESULTS vs paper: C1 total reads 1,136,666,560 EXACT (also matches ENA metadata 568,333,280 pairs x2); C2 duplicates 30,877,024 vs 28,278,026 (+9.2%, within-tol; PERCENT_DUPLICATION 0.0282); C3a reads-left 973,194,243 vs 996,064,609 (-2.3%, within-tol); C3b MAPQ<10 filtered 132,595,293 vs 112,323,925 (+18%, partial - the most reference/aligner-version-sensitive count); C4 HaplotypeCaller variants 4,064,130 all-records vs 3,969,250 (+2.4%, within-tol; SNV-only 3,361,639 = -15.3%, paper's 'SNPs' definition is ambiguous and both bracket the reported value); C5 het+>=5-reads SNVs 1,584,694 vs 1,427,492 (+11.0%, within-tol). FABRICATION CHECK PASSES: the paper's read accounting is digit-exact internally (1,136,666,560 - 28,278,026 - 112,323,925 = 996,064,609) and a fully independent re-run lands within 2-18% of every reported count with identical internal structure (dup ~2.7%, het/SNP ~39%) - the reported values are reproducible and genuine, no fabrication signal. Deviations are fully consistent with the paper not specifying hg38 flavor, tool versions, or dbSNP build. NOT ATTEMPTED (out of scope, infeasible corpus scale): genome-wide ASB over 7,669 GTRD ChIP-seq experiments, BABACHI BAD-map fit, BaalChIP/GWAS/eQTL downstream analyses. All heavy compute targeted «our HPC» compute nodes («infra» workdir); «host» held only orchestration + small results. Overall: a strong within-tolerance concordant reproduction of the in-scope sub-analysis; graded 'partial' because only this single self-contained sub-analysis of a corpus-scale paper was in scope and one of its six counts (C3b) lands outside tight tolerance.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 82
    assessed: 2026-06-20 ⛓ 110fab237c3b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Differential transcription factor binding at heterozygous single-nucleotide variants in ChIP-Seq data can be used to systematically identify allele-specific binding (ASB) events genome-wide, provided the confounding effects of aneuploidy and local copy-number variation (background allelic dosage) and reference mapping bias are properly corrected for.

Core claims
  • A novel statistical framework (ADASTRA) calls allele-specific TF binding from existing ChIP-Seq alignments by jointly correcting for background allelic dosage (BAD, from aneuploidy/CNVs) and reference mapping bias. method
  • The ADASTRA database assembles genome-wide ASB events across 674 human TFs and 337 cell types, listing more than half a million ASB entries at nearly 270,000 SNPs. resource
  • SNPs with ASB events are enriched for associations with medically relevant phenotypes and frequently overlap eQTLs, nominating candidate causal regulatory variants. finding
  • A subset of 'switching sites' exists where different TFs preferentially bind alternative alleles, revealing allele-specific rewiring of regulatory circuitry. finding
  • A Bayesian changepoint identification algorithm reconstructs genome-wide BAD maps directly from SNV read counts in ChIP-Seq data without requiring matched input/control CNV data. method
  • Predicted BAD maps show strong agreement with COSMIC copy-number ground truth (AUC-ROC > 0.83 for BAD values 1-3). finding
  • Segregating variant read counts by BAD removes much of the overdispersion previously requiring separate correction in ASB statistical models. mechanism
  • MCF7 BAD maps agree poorly with COSMIC copy-number data regardless of SNV count, likely reflecting genuine genomic instability of this cell line rather than a method failure. finding
Experimental setups
Assay System Perturbation Readout Platform
ChIP-Seq (reprocessed alignments, meta-analysis) 566 human cell types, 1025 human TFs (GTRD database) none (endogenous heterozygous SNVs) allele-specific read counts at SNVs / ASB calls GTRD database, hg38 alignments
Variant calling from ChIP-Seq read alignments GTRD ChIP-Seq data sets (7669 alignments) none filtered SNV genotype calls (dbSNP common subset, >=5 reads/allele)
Bayesian changepoint BAD segmentation K562 cells (ENCODE biosample ENCBS725WFV), chr2 and chr6 none predicted background allelic dosage (BAD) segments vs COSMIC ground truth ENCODE
Deep genomic (non-ChIP) sequencing MCF7 cell line none (ChIP-independent control) BAD map from genomic DNA, compared to ChIP-derived MCF7 BAD maps
Microarray-based CNV estimation 13 cell types matched across COSMIC and study data none CNV estimates compared with COSMIC and BAD maps microarray
COSMIC CNV database comparison 76 matched cell types (predominantly K562, MCF7) none ground-truth copy number vs predicted BAD (Kendall tau_b, ROC/PRC) COSMIC database
Key results
  • ADASTRA identified 233,290 TF-level ASBs at 147,909 SNPs and 351,967 cell-type-level ASBs at 252,469 SNPs, passing FDR<0.05.
  • 7669 ChIP-Seq alignments covering 1025 TFs and 566 cell types were processed, using a filtered set of more than 54 million variant call entries. 54 million entries
  • BAD maps classified individual SNVs with AUC-ROC > 0.83 and AUPRC 0.66-0.85 for the most common BAD values (1-3) against COSMIC ground truth. AUC>0.83
  • Kendall tau_b rank correlation between predicted and COSMIC BAD improved as the number of SNV calls in a data set group increased.
  • The ChIP-independent genomic BAD map for MCF7 still correlated poorly with COSMIC copy numbers, despite validating well against ChIP-based MCF7 BAD maps. Kendall tau_b ~0.2
  • The top 8 TFs and top 5 cell types accounted for only about half of ASB calls, indicating broad diversity across samples.
  • ASB SNPs were enriched for associations with medically relevant phenotypes and frequently overlapped eQTLs.
Key statistics
  • count 233,290 ASBs at 147,909 SNPs (ADASTRA TF-level ASB dataset)
  • count 351,967 ASBs at 252,469 SNPs (ADASTRA cell-type-level ASB dataset)
  • pvalue P < 0.05 (Benjamini-Hochberg FDR adjusted) (significance threshold for calling ASBs)
  • count 7669 ChIP-Seq alignments; 1025 TFs; 566 cell types (input data processed by the ADASTRA pipeline)
  • count >54 million (filtered variant call entries used for the statistical framework)
  • other AUC-ROC >0.83; AUPRC 0.66-0.85 (BAD map classifier performance for BAD values 1-3 vs COSMIC)
  • correlation Kendall tau_b ~0.2 (MCF7 ChIP-independent BAD map vs COSMIC copy-number data)
  • count 2556 groups of variant calls; 76 matched cell types (BAD calling groups and COSMIC ground-truth comparison set)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper presents a computational meta-analysis of 7,669 human ChIP-Seq alignments to identify allele-specific transcription factor binding (ASB) events genome-wide. The core statistical workflow involves: (1) a novel Bayesian changepoint identification algorithm using dynamic programming to segment the genome into regions of stable background allelic dosage (BAD) from SNV read counts; (2) a mixture of two negative binomial distributions as the null background model for allelic read imbalance, conditioned on BAD and read coverage; and (3) per-experiment P values aggregated across experiments sharing a TF or cell type using the George–Mudholkar method, followed by Benjamini–Hochberg FDR correction at 0.05. BAD map validity was assessed against COSMIC copy-number ground truth using Kendall τb rank correlation and ROC/precision-recall curve analysis.

Replicationunclear Sample size7669 ChIP-Seq read alignments from GTRD covering 1025 TFs and 566 cell types; more than 54 million filtered variant call entries; 2556 BAD-calling groups defined by cell type and GEO series or ENCODE biosample ID GroupsReference allele vs. alternative allele read counts at heterozygous SNVs; ASBs aggregated separately at TF level (TF-ASB, 233,290 events at 147,909 SNPs) and cell type level (CT-ASB, 351,967 events at 252,469 SNPs) Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini–Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
Bayesian changepoint identification (dynamic programming maximizing marginal likelihood, followed by maximum-posterior BAD assignment per segment) Genome-wide segmentation into regions of stable background allelic dosage (BAD) from SNV read counts, performed for 2556 groups of ChIP-Seq data sets More than 54 million filtered variant call entries across 7669 alignments not stated
Mixture of two negative binomial distributions (with BAD-determined p parameters) for null background of allelic read imbalance Background distribution fitting for ASB calling, applied per BAD segment and read-coverage stratum, separately for reference-allele and alternative-allele read counts not stated
George–Mudholkar P-value aggregation (meta-analytic combination of independent P values across experiments) Combining per-experiment allele-read-bias P values across all experiments sharing a TF (TF-ASB) or a cell type (CT-ASB) 7669 ChIP-Seq alignments; 1025 TFs; 566 cell types not stated
Kendall τb rank correlation Validation of predicted BAD maps against COSMIC copy-number ground truth across 76 matched cell types (516 data-set groups) 516 groups of related data sets; 76 matched cell types not stated
ROC curve / AUC and precision-recall curve (PRC) analysis Evaluating predicted BAD maps as binary classifiers of individual SNVs for BAD values 1, 3/2, 2, and 3 against COSMIC ground truth BAD values 1–3 cover more than 97% of candidate ASB variants na
Approaches that could also have been used
  • Per-experiment P values for allelic read imbalance were combined across experiments using the George–Mudholkar method
    Could also: Fisher's combined probability test or Stouffer's Z-score method are also standard approaches for meta-analytic P-value combination — Stouffer's method allows weighting individual P values by read depth or experiment quality, which may be advantageous when experiments vary substantially in coverage; Fisher's method is the most widely known and reported, facilitating comparability across studies; the choice of aggregation method can affect sensitivity to small but consistent allelic signals
  • Allelic read imbalance at heterozygous SNVs was modeled using a mixture of two negative binomial distributions conditioned on BAD
    Could also: A beta-binomial distribution is another established model for overdispersed allelic read counts and is used in tools such as WASP, RASQUAL, and MBASED — The beta-binomial directly models extra-binomial variation as a compound distribution and has a single overdispersion parameter, which can simplify fitting; comparing the two frameworks would clarify how the choice of background model affects ASB call rates and false-discovery control
  • Genome-wide BAD segmentation was performed with a Bayesian changepoint identification algorithm using dynamic programming
    Could also: Hidden Markov Model (HMM)-based segmentation, as implemented in tools such as HMMcopy, CNVnator, or GATK ModelSegments, is a widely used alternative for segmenting allelic dosage along chromosomes — HMM-based approaches provide explicit transition probabilities between copy-number states and are well-characterized in the CNV literature; they offer a natural comparison point for evaluating whether the Bayesian changepoint approach yields different segment boundaries or BAD assignments, particularly in regions of complex structural variation
  • Benjamini–Hochberg FDR correction was applied separately within each TF family and each cell type family, and separately for Ref-ASB and Alt-ASB
    Could also: Storey's q-value approach, which estimates the proportion of true null hypotheses (π0) from the P-value distribution, could also be applied; alternatively, a single joint FDR correction across the full set of tests could be used — Storey's q-value can increase power when π0 is substantially less than 1, as expected in a large-scale ASB search; a joint correction across all TFs and cell types simultaneously would control the global database-wide FDR rather than local per-TF/per-cell-type FDR, offering a different sensitivity–specificity trade-off
  • BAD map concordance with COSMIC copy-number profiles was quantified using Kendall τb rank correlation
    Could also: Spearman's rank correlation coefficient (ρ) or the concordance correlation coefficient (CCC) are also commonly used to assess agreement between ordered or continuous measurements from two methods — Spearman's ρ is more widely familiar and directly comparable across publications; CCC additionally penalizes systematic location and scale differences between the two measurement systems (beyond rank agreement), which could reveal whether BAD estimates are biased relative to COSMIC copy numbers, not merely mis-ranked
  • Reference read mapping bias was accounted for statistically within the negative binomial model rather than at the alignment stage
    Could also: Alignment-stage bias correction approaches such as N-masking of known SNP positions in the reference, alignment to a personalized diploid genome, or post-alignment filtering with WASP are established alternatives — Alignment-stage correction removes the source of bias before statistical modeling and can improve accuracy at SNVs with strong reference preference without requiring parametric assumptions about bias magnitude; the statistical correction used here was chosen to enable processing of pre-made alignments at scale, and noting this trade-off helps readers assess applicability to other study designs
Software: ADASTRA pipeline (custom, described in paper) · GTRD (Gene Transcription Regulation Database, source of ChIP-Seq alignments) · dbSNP (common SNP subset, used for variant filtering and annotation) · COSMIC (copy-number variant data, used as ground-truth for BAD validation)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33980847 (ADASTRA: Landscape of allele-specific TF binding)

  • Paper: Abramov et al., Nat Commun 2021. PMID 33980847 / PMC8115691 / DOI 10.1038/s41467-021-23007-0.
  • Pipeline code: https://github.com/autosome-ru/ADASTRA-pipeline (master, commit 833cbc9d9fd5d6ec1561611bd4a5bb7e1740bae4, last push 2021-10-20). No license file.
  • Assigned data accession: SRR8652105 = CCLE deep whole-genome sequencing of the MCF7 breast-cancer cell line (PRJNA523380 / SRX5449795 / SAMN10988179). Paired-end, 568,333,280 read pairs (= 1,136,666,560 reads), ~67 GB gz.

What the paper reports (and from what)

The headline ADASTRA results are a full-database meta-analysis over 7,669 ChIP-Seq alignments from GTRD, yielding >500k ASB events at ~270k SNPs for hundreds of TFs and cell types. Re-running that whole corpus (terabytes of input, weeks of compute) is out of scope / infeasible here and is not attempted.

However, the paper contains one self-contained sub-analysis that uses exactly our assigned accession: the section "Construction of an independent BAD map for MCF7 cells" (Methods). It runs ADASTRA Part A (SNP calling) on SRR8652105 alone and reports explicit per-stage counts. This is the reproduction target.

Reported method (verbatim paraphrase)

The paired-end reads of MCF7 deep genome sequencing (SRR8652105) were aligned to hg38 using bowtie2 with default settings. 28,278,026 (2.5%) of 1,136,666,560 paired reads were marked as duplicates; 112,323,925 (9.9%) were filtered by GATK mapping-quality ≥10, leaving 996,064,609 reads for SNP calling. 3,969,250 SNPs were reported by GATK HaplotypeCaller, among which 1,427,492 were heterozygous and passed the basic ADASTRA filter (≥5 reads on each allele), and were used to build the MCF7 BAD map with BABACHI.

General ADASTRA variant-calling method (Methods "Variant calling from GTRD alignments"): PICARD dedup → GATK BQSR → GATK HaplotypeCaller with dbSNP151-common annotation → filter to biallelic heterozygous (GT=0/1), ≥5 reads each allele, and (for candidate ASB) present in dbSNP151 common. For the MCF7 BAD-map count, only het + ≥5-reads is required (dbSNP-common membership is the additional filter for genome-wide candidate-ASB sites, not for this count).

In scope (this room)

id claim reported how reproduced
C1 Total paired reads in SRR8652105 1,136,666,560 direct from ENA/SRA metadata (568,333,280 pairs × 2) — exact, no compute
C2 Reads marked as duplicates 28,278,026 (2.5%) bowtie2→sort→MarkDuplicates (PICARD/GATK)
C3 Reads filtered by MAPQ≥10 / reads left 112,323,925 (9.9%) / 996,064,609 samtools flag/MAPQ counts
C4 SNPs reported by GATK HaplotypeCaller 3,969,250 HaplotypeCaller VCF, count SNV records
C5 Heterozygous SNPs passing ≥5 reads/allele 1,427,492 VCF filter GT=0/1 & AD_ref≥5 & AD_alt≥5

Internal consistency cross-check (cheap fabrication test) passes: 1,136,666,560 − 28,278,026 − 112,323,925 = 996,064,609 (exact); dup=2.488%≈2.5%; mapq=9.882%≈9.9%; het/SNP=35.96%.

Out of scope (not attempted; stated honestly)

  • Genome-wide ASB calling over 7,669 GTRD ChIP-Seq experiments (>500k ASB / ~270k SNPs / 1232 TFs etc.) — infeasible corpus scale.
  • BAD-map construction with BABACHI, NB-mixture fit, FDR aggregation, effect sizes — depend on the full corpus / the MCF7 BAD map is a downstream product; we stop at the SNP-calling counts that define its input.
  • BaalChIP comparison (Fig. 2d, 4,976,303 sites), GWAS/eQTL enrichment, TF-switching — corpus-level, non-single-dataset.

Reproduction plan («our HPC» SLURM, «infra»; 12 h/job cap)

  1. ref-prep: hg38 fasta + bowtie2 index + .fai/.dict + dbSNP151-common VCF (cache on «infra»).
  2. align: download SRR8652105 → bowtie2 (default) → sort → MarkDuplicates (→C2) → MAPQ≥10 filter (→C3).
  3. call: GATK HaplotypeCaller (scatter-gather)
C1_total_reads
Reported
1136666560
Reproduced
1136666560
exact
C2_duplicates
Reported
28278026 (2.5%)
Reproduced
30877024 (2.72%)
within tolerance
C3a_reads_left
Reported
996064609
Reproduced
973194243
within tolerance
C3b_mapq_filtered
Reported
112323925 (9.9%)
Reproduced
132595293 (11.7%)
partial
C4_gatk_snps
Reported
3969250
Reproduced
4064130 all-records / 3361639 SNVs
within tolerance
C5_het_5reads
Reported
1427492
Reproduced
1584694
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

1.1 M
tokens (I/O) · 64.1 M incl. cache
446 min
runtime · 120.95 CPU-h
92.4 GB
peak RAM
5
HPC jobs
hummel
machine