Landscape of allele-specific transcription factor binding in the human genome.
The main results reproduced: recomputed values matched the published ones within tolerance.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
ADASTRA (Abramov et al., Nat Commun 2021) Part-A SNP-calling reproduced on the assigned accession SRR8652105 (CCLE MCF7 deep WGS), targeting the one self-contained Methods sub-analysis that uses this exact accession: 'Construction of an independent BAD map for MCF7 cells', which reports 5 explicit per-stage counts (C1-C5). The full faithful chain executed to completion on «our HPC» (SLURM/«infra», «job»-2208547): bowtie2 2.4.5 default settings -> UCSC hg38 analysisSet (no-alt) -> samtools sort -> PICARD MarkDuplicates(REMOVE_DUPLICATES=true) -> MAPQ>=10 accounting -> GATK 4.1.9.0 BQSR -> ApplyBQSR -> HaplotypeCaller(--dbsnp, scatter by chr) -> MergeVcfs -> bcftools counts. RESULTS vs paper: C1 total reads 1,136,666,560 EXACT (also matches ENA metadata 568,333,280 pairs x2); C2 duplicates 30,877,024 vs 28,278,026 (+9.2%, within-tol; PERCENT_DUPLICATION 0.0282); C3a reads-left 973,194,243 vs 996,064,609 (-2.3%, within-tol); C3b MAPQ<10 filtered 132,595,293 vs 112,323,925 (+18%, partial - the most reference/aligner-version-sensitive count); C4 HaplotypeCaller variants 4,064,130 all-records vs 3,969,250 (+2.4%, within-tol; SNV-only 3,361,639 = -15.3%, paper's 'SNPs' definition is ambiguous and both bracket the reported value); C5 het+>=5-reads SNVs 1,584,694 vs 1,427,492 (+11.0%, within-tol). FABRICATION CHECK PASSES: the paper's read accounting is digit-exact internally (1,136,666,560 - 28,278,026 - 112,323,925 = 996,064,609) and a fully independent re-run lands within 2-18% of every reported count with identical internal structure (dup ~2.7%, het/SNP ~39%) - the reported values are reproducible and genuine, no fabrication signal. Deviations are fully consistent with the paper not specifying hg38 flavor, tool versions, or dbSNP build. NOT ATTEMPTED (out of scope, infeasible corpus scale): genome-wide ASB over 7,669 GTRD ChIP-seq experiments, BABACHI BAD-map fit, BaalChIP/GWAS/eQTL downstream analyses. All heavy compute targeted «our HPC» compute nodes («infra» workdir); «host» held only orchestration + small results. Overall: a strong within-tolerance concordant reproduction of the in-scope sub-analysis; graded 'partial' because only this single self-contained sub-analysis of a corpus-scale paper was in scope and one of its six counts (C3b) lands outside tight tolerance.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 82assessed: 2026-06-20 ⛓ 110fab237c3b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetDifferential transcription factor binding at heterozygous single-nucleotide variants in ChIP-Seq data can be used to systematically identify allele-specific binding (ASB) events genome-wide, provided the confounding effects of aneuploidy and local copy-number variation (background allelic dosage) and reference mapping bias are properly corrected for.
- ★ A novel statistical framework (ADASTRA) calls allele-specific TF binding from existing ChIP-Seq alignments by jointly correcting for background allelic dosage (BAD, from aneuploidy/CNVs) and reference mapping bias. method
- ★ The ADASTRA database assembles genome-wide ASB events across 674 human TFs and 337 cell types, listing more than half a million ASB entries at nearly 270,000 SNPs. resource
- ★ SNPs with ASB events are enriched for associations with medically relevant phenotypes and frequently overlap eQTLs, nominating candidate causal regulatory variants. finding
- ★ A subset of 'switching sites' exists where different TFs preferentially bind alternative alleles, revealing allele-specific rewiring of regulatory circuitry. finding
- ★ A Bayesian changepoint identification algorithm reconstructs genome-wide BAD maps directly from SNV read counts in ChIP-Seq data without requiring matched input/control CNV data. method
- ★ Predicted BAD maps show strong agreement with COSMIC copy-number ground truth (AUC-ROC > 0.83 for BAD values 1-3). finding
- ★ Segregating variant read counts by BAD removes much of the overdispersion previously requiring separate correction in ASB statistical models. mechanism
- MCF7 BAD maps agree poorly with COSMIC copy-number data regardless of SNV count, likely reflecting genuine genomic instability of this cell line rather than a method failure. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ChIP-Seq (reprocessed alignments, meta-analysis) | 566 human cell types, 1025 human TFs (GTRD database) | none (endogenous heterozygous SNVs) | allele-specific read counts at SNVs / ASB calls | GTRD database, hg38 alignments |
| Variant calling from ChIP-Seq read alignments | GTRD ChIP-Seq data sets (7669 alignments) | none | filtered SNV genotype calls (dbSNP common subset, >=5 reads/allele) | — |
| Bayesian changepoint BAD segmentation | K562 cells (ENCODE biosample ENCBS725WFV), chr2 and chr6 | none | predicted background allelic dosage (BAD) segments vs COSMIC ground truth | ENCODE |
| Deep genomic (non-ChIP) sequencing | MCF7 cell line | none (ChIP-independent control) | BAD map from genomic DNA, compared to ChIP-derived MCF7 BAD maps | — |
| Microarray-based CNV estimation | 13 cell types matched across COSMIC and study data | none | CNV estimates compared with COSMIC and BAD maps | microarray |
| COSMIC CNV database comparison | 76 matched cell types (predominantly K562, MCF7) | none | ground-truth copy number vs predicted BAD (Kendall tau_b, ROC/PRC) | COSMIC database |
- – ADASTRA identified 233,290 TF-level ASBs at 147,909 SNPs and 351,967 cell-type-level ASBs at 252,469 SNPs, passing FDR<0.05.
- – 7669 ChIP-Seq alignments covering 1025 TFs and 566 cell types were processed, using a filtered set of more than 54 million variant call entries. 54 million entries
- ▲ BAD maps classified individual SNVs with AUC-ROC > 0.83 and AUPRC 0.66-0.85 for the most common BAD values (1-3) against COSMIC ground truth. AUC>0.83
- ▲ Kendall tau_b rank correlation between predicted and COSMIC BAD improved as the number of SNV calls in a data set group increased.
- ▼ The ChIP-independent genomic BAD map for MCF7 still correlated poorly with COSMIC copy numbers, despite validating well against ChIP-based MCF7 BAD maps. Kendall tau_b ~0.2
- – The top 8 TFs and top 5 cell types accounted for only about half of ASB calls, indicating broad diversity across samples.
- – ASB SNPs were enriched for associations with medically relevant phenotypes and frequently overlapped eQTLs.
- count 233,290 ASBs at 147,909 SNPs (ADASTRA TF-level ASB dataset)
- count 351,967 ASBs at 252,469 SNPs (ADASTRA cell-type-level ASB dataset)
- pvalue P < 0.05 (Benjamini-Hochberg FDR adjusted) (significance threshold for calling ASBs)
- count 7669 ChIP-Seq alignments; 1025 TFs; 566 cell types (input data processed by the ADASTRA pipeline)
- count >54 million (filtered variant call entries used for the statistical framework)
- other AUC-ROC >0.83; AUPRC 0.66-0.85 (BAD map classifier performance for BAD values 1-3 vs COSMIC)
- correlation Kendall tau_b ~0.2 (MCF7 ChIP-independent BAD map vs COSMIC copy-number data)
- count 2556 groups of variant calls; 76 matched cell types (BAD calling groups and COSMIC ground-truth comparison set)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper presents a computational meta-analysis of 7,669 human ChIP-Seq alignments to identify allele-specific transcription factor binding (ASB) events genome-wide. The core statistical workflow involves: (1) a novel Bayesian changepoint identification algorithm using dynamic programming to segment the genome into regions of stable background allelic dosage (BAD) from SNV read counts; (2) a mixture of two negative binomial distributions as the null background model for allelic read imbalance, conditioned on BAD and read coverage; and (3) per-experiment P values aggregated across experiments sharing a TF or cell type using the George–Mudholkar method, followed by Benjamini–Hochberg FDR correction at 0.05. BAD map validity was assessed against COSMIC copy-number ground truth using Kendall τb rank correlation and ROC/precision-recall curve analysis.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Bayesian changepoint identification (dynamic programming maximizing marginal likelihood, followed by maximum-posterior BAD assignment per segment) | Genome-wide segmentation into regions of stable background allelic dosage (BAD) from SNV read counts, performed for 2556 groups of ChIP-Seq data sets | More than 54 million filtered variant call entries across 7669 alignments | not stated |
| Mixture of two negative binomial distributions (with BAD-determined p parameters) for null background of allelic read imbalance | Background distribution fitting for ASB calling, applied per BAD segment and read-coverage stratum, separately for reference-allele and alternative-allele read counts | — | not stated |
| George–Mudholkar P-value aggregation (meta-analytic combination of independent P values across experiments) | Combining per-experiment allele-read-bias P values across all experiments sharing a TF (TF-ASB) or a cell type (CT-ASB) | 7669 ChIP-Seq alignments; 1025 TFs; 566 cell types | not stated |
| Kendall τb rank correlation | Validation of predicted BAD maps against COSMIC copy-number ground truth across 76 matched cell types (516 data-set groups) | 516 groups of related data sets; 76 matched cell types | not stated |
| ROC curve / AUC and precision-recall curve (PRC) analysis | Evaluating predicted BAD maps as binary classifiers of individual SNVs for BAD values 1, 3/2, 2, and 3 against COSMIC ground truth | BAD values 1–3 cover more than 97% of candidate ASB variants | na |
-
Per-experiment P values for allelic read imbalance were combined across experiments using the George–Mudholkar method↳ Could also: Fisher's combined probability test or Stouffer's Z-score method are also standard approaches for meta-analytic P-value combination — Stouffer's method allows weighting individual P values by read depth or experiment quality, which may be advantageous when experiments vary substantially in coverage; Fisher's method is the most widely known and reported, facilitating comparability across studies; the choice of aggregation method can affect sensitivity to small but consistent allelic signals
-
Allelic read imbalance at heterozygous SNVs was modeled using a mixture of two negative binomial distributions conditioned on BAD↳ Could also: A beta-binomial distribution is another established model for overdispersed allelic read counts and is used in tools such as WASP, RASQUAL, and MBASED — The beta-binomial directly models extra-binomial variation as a compound distribution and has a single overdispersion parameter, which can simplify fitting; comparing the two frameworks would clarify how the choice of background model affects ASB call rates and false-discovery control
-
Genome-wide BAD segmentation was performed with a Bayesian changepoint identification algorithm using dynamic programming↳ Could also: Hidden Markov Model (HMM)-based segmentation, as implemented in tools such as HMMcopy, CNVnator, or GATK ModelSegments, is a widely used alternative for segmenting allelic dosage along chromosomes — HMM-based approaches provide explicit transition probabilities between copy-number states and are well-characterized in the CNV literature; they offer a natural comparison point for evaluating whether the Bayesian changepoint approach yields different segment boundaries or BAD assignments, particularly in regions of complex structural variation
-
Benjamini–Hochberg FDR correction was applied separately within each TF family and each cell type family, and separately for Ref-ASB and Alt-ASB↳ Could also: Storey's q-value approach, which estimates the proportion of true null hypotheses (π0) from the P-value distribution, could also be applied; alternatively, a single joint FDR correction across the full set of tests could be used — Storey's q-value can increase power when π0 is substantially less than 1, as expected in a large-scale ASB search; a joint correction across all TFs and cell types simultaneously would control the global database-wide FDR rather than local per-TF/per-cell-type FDR, offering a different sensitivity–specificity trade-off
-
BAD map concordance with COSMIC copy-number profiles was quantified using Kendall τb rank correlation↳ Could also: Spearman's rank correlation coefficient (ρ) or the concordance correlation coefficient (CCC) are also commonly used to assess agreement between ordered or continuous measurements from two methods — Spearman's ρ is more widely familiar and directly comparable across publications; CCC additionally penalizes systematic location and scale differences between the two measurement systems (beyond rank agreement), which could reveal whether BAD estimates are biased relative to COSMIC copy numbers, not merely mis-ranked
-
Reference read mapping bias was accounted for statistically within the negative binomial model rather than at the alignment stage↳ Could also: Alignment-stage bias correction approaches such as N-masking of known SNP positions in the reference, alignment to a personalized diploid genome, or post-alignment filtering with WASP are established alternatives — Alignment-stage correction removes the source of bias before statistical modeling and can improve accuracy at SNVs with strong reference preference without requiring parametric assumptions about bias magnitude; the statistical correction used here was chosen to enable processing of pre-made alignments at scale, and noting this trade-off helps readers assess applicability to other study designs
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33980847 (ADASTRA: Landscape of allele-specific TF binding)
- Paper: Abramov et al., Nat Commun 2021. PMID 33980847 / PMC8115691 / DOI 10.1038/s41467-021-23007-0.
- Pipeline code: https://github.com/autosome-ru/ADASTRA-pipeline (master, commit
833cbc9d9fd5d6ec1561611bd4a5bb7e1740bae4, last push 2021-10-20). No license file. - Assigned data accession:
SRR8652105= CCLE deep whole-genome sequencing of the MCF7 breast-cancer cell line (PRJNA523380 / SRX5449795 / SAMN10988179). Paired-end, 568,333,280 read pairs (= 1,136,666,560 reads), ~67 GB gz.
What the paper reports (and from what)
The headline ADASTRA results are a full-database meta-analysis over 7,669 ChIP-Seq alignments from GTRD, yielding >500k ASB events at ~270k SNPs for hundreds of TFs and cell types. Re-running that whole corpus (terabytes of input, weeks of compute) is out of scope / infeasible here and is not attempted.
However, the paper contains one self-contained sub-analysis that uses exactly our assigned accession: the section "Construction of an independent BAD map for MCF7 cells" (Methods). It runs ADASTRA Part A (SNP calling) on SRR8652105 alone and reports explicit per-stage counts. This is the reproduction target.
Reported method (verbatim paraphrase)
The paired-end reads of MCF7 deep genome sequencing (SRR8652105) were aligned to hg38 using bowtie2 with default settings. 28,278,026 (2.5%) of 1,136,666,560 paired reads were marked as duplicates; 112,323,925 (9.9%) were filtered by GATK mapping-quality ≥10, leaving 996,064,609 reads for SNP calling. 3,969,250 SNPs were reported by GATK HaplotypeCaller, among which 1,427,492 were heterozygous and passed the basic ADASTRA filter (≥5 reads on each allele), and were used to build the MCF7 BAD map with BABACHI.
General ADASTRA variant-calling method (Methods "Variant calling from GTRD alignments"): PICARD dedup → GATK BQSR → GATK HaplotypeCaller with dbSNP151-common annotation → filter to biallelic heterozygous (GT=0/1), ≥5 reads each allele, and (for candidate ASB) present in dbSNP151 common. For the MCF7 BAD-map count, only het + ≥5-reads is required (dbSNP-common membership is the additional filter for genome-wide candidate-ASB sites, not for this count).
In scope (this room)
| id | claim | reported | how reproduced |
|---|---|---|---|
| C1 | Total paired reads in SRR8652105 | 1,136,666,560 | direct from ENA/SRA metadata (568,333,280 pairs × 2) — exact, no compute |
| C2 | Reads marked as duplicates | 28,278,026 (2.5%) | bowtie2→sort→MarkDuplicates (PICARD/GATK) |
| C3 | Reads filtered by MAPQ≥10 / reads left | 112,323,925 (9.9%) / 996,064,609 | samtools flag/MAPQ counts |
| C4 | SNPs reported by GATK HaplotypeCaller | 3,969,250 | HaplotypeCaller VCF, count SNV records |
| C5 | Heterozygous SNPs passing ≥5 reads/allele | 1,427,492 | VCF filter GT=0/1 & AD_ref≥5 & AD_alt≥5 |
Internal consistency cross-check (cheap fabrication test) passes: 1,136,666,560 − 28,278,026 − 112,323,925 = 996,064,609 (exact); dup=2.488%≈2.5%; mapq=9.882%≈9.9%; het/SNP=35.96%.
Out of scope (not attempted; stated honestly)
- Genome-wide ASB calling over 7,669 GTRD ChIP-Seq experiments (>500k ASB / ~270k SNPs / 1232 TFs etc.) — infeasible corpus scale.
- BAD-map construction with BABACHI, NB-mixture fit, FDR aggregation, effect sizes — depend on the full corpus / the MCF7 BAD map is a downstream product; we stop at the SNP-calling counts that define its input.
- BaalChIP comparison (Fig. 2d, 4,976,303 sites), GWAS/eQTL enrichment, TF-switching — corpus-level, non-single-dataset.
Reproduction plan («our HPC» SLURM, «infra»; 12 h/job cap)
- ref-prep: hg38 fasta + bowtie2 index + .fai/.dict + dbSNP151-common VCF (cache on «infra»).
- align: download SRR8652105 → bowtie2 (default) → sort → MarkDuplicates (→C2) → MAPQ≥10 filter (→C3).
- call: GATK HaplotypeCaller (scatter-gather)
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.