Transcriptional regulation and chromatin architecture maintenance are decoupled functions at the Sox2 locus.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (pipeline-derived results, from deposited GSE195906 .cis on «our HPC»). The deposited 4C .cis profiles are confirmed to be genuine mm10 DpnII fragment-resolved profiles: 391374/391374 (100%) fragment midpoints match an independently rebuilt mm10 chr3 DpnII (GATC) fragment map (4See makefrags/coord2frag), and the bait self-ligation peak sits 238-3929bp from the declared viewpoint (chr3:34,750,100, within the SCR) -> strong provenance, no fabrication signal (R1-structural EXACT). The headline deltaSCR claim reproduces within tolerance: paper reports a 28% decrease (P=0.02) in SCR-proximal-bait<->Sox2 contact frequency in delSCR/delSCR vs WT; our limma quantile-normalized + two-tailed t-test pipeline on the Sox2-spanning region (chr3:34,644,922-34,664,967) gives 30.4% decrease, P=0.006 (n=4 vs 4) (R3a within-tol). peakC (window=21) calls the SCR-proximal<->Sox2 interaction (22 significant fragments in the Sox2-spanning region; 3/4 WT replicates individually) (R2 within-tol). NOT attempted/clean: R3b (P=0.74 het-vs-homozygous; allele-track assignment ambiguous from GEO labels), full .cis regeneration from raw fastq (needs bait primers, Suppl Table S6), and all wet-lab/imaging/qPCR results (out of scope). Code is a valid third-party + author-utility stack (najoshi/sabre + TomSexton00/4See + deWitLab/peakC) applied to the paper's own data (P16-valid).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 56assessed: 2026-06-19 ⛓ 2e670348900c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetHow distal regulatory elements (enhancers) control gene transcription versus chromatin topology is unclear; the paper tests, at the Sox2 locus in mouse ESCs, whether transcriptional activation and chromatin–chromatin interaction/TAD maintenance are mediated by the same or different cis-regulatory sequences.
- ★ Sox2 transcriptional activation is traced almost entirely to two key transcription factor-bound regions (SRR107 and SRR111) within the SCR finding
- ★ Deletion of SRR107/SRR111 (the transcription-driving regions) has no effect on promoter–enhancer (SCR–Sox2) interaction frequency or TAD organization finding
- ★ Local chromatin architecture maintenance is distributed over multiple transcription factor-bound regions across the locus and is CTCF-independent finding
- ★ Ectopic chromatin loop formation that partially disrupts promoter–enhancer interactions has no effect on Sox2 transcription finding
- ★ Significant disruption of chromatin interaction frequency and TAD boundary insulation requires deletion of the entire SCR, not just the transcription-driving subregions finding
- ★ Deletion of the sole CTCF-bound site within the SCR (SRR109) does not affect chromatin topology or Sox2 transcription finding
- ★ SRR107 (via an OCT4:SOX2 composite motif) and SRR111 (via two KLF4 motifs) act in a partially redundant manner to drive SCR enhancer activity mechanism
- Allele-specific CRISPR/Cas9 genome editing combined with allele-specific 4C-seq and RT-qPCR was used to dissect cis-regulatory function at the Sox2 locus method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| allele-specific 4C-seq | F1 mouse ESCs (129 x castaneus), including ΔSCR/ΔSCR and heterozygous ΔSCR clones | CRISPR/Cas9 deletion of the SCR (homozygous and heterozygous) | chromatin interaction frequency across the Sox2 locus from an SCR-proximal or Sox2 promoter bait | — |
| ChIP-seq | mouse ESCs | none | binding of CTCF, RAD21, SMC1A, MED1, EP300, and H3K27ac enrichment across the Sox2 locus | — |
| allele-specific RT-qPCR | F1 mouse ESC clones (129 x CAST) | CRISPR/Cas9 heterozygous deletion of SRR106, SRR107, SRR109, SRR111 individually or in combination | allele-specific Sox2 transcript levels relative to total transcript | — |
| allele-specific RT-qPCR | F1 mouse ESC clones | microdeletion of OCT4:SOX2 motif in SRR107 or KLF4 motifs in SRR111 (on background lacking the other SRR) | allele-specific Sox2 transcript levels | — |
| allele-specific ChIP | mouse ESCs with OCT4:SOX2 motif-deleted SRR107 | OCT4:SOX2 motif deletion | OCT4 and RNA polymerase II association at the altered SRR | — |
| motif discovery (JASPAR GeneReg database) | in silico sequence analysis of SRR107/SRR111 | none | high-scoring transcription factor binding motifs (OCT4:SOX2, KLF4) | JASPAR GeneReg database tool |
| ChIP-seq data overlap (CODEX database) | compiled ESC ChIP-seq data sets | none | confirmation of transcription factor occupancy over identified motifs | CODEX database |
- ▼ ΔSCR/ΔSCR ESCs show decreased SCR-proximal bait to Sox2 contact frequency versus wild type 28%, P=0.02
- ▼ In heterozygous ΔSCR cells, the deleted allele shows reduced Sox2–SCR contact frequency while the WT allele is unchanged 24% reduction (P=0.04) on ΔSCR allele; P=0.46 on WT allele
- ▼ Deletion of SRR107 alone reduces allele-specific Sox2 transcript levels 27% reduction, significant
- ▼ Deletion of SRR111 alone causes a weaker, nonsignificant reduction in Sox2 transcript levels 14% reduction, not significant
- ▼ Compound deletion of SRR107 and SRR111 on the same allele causes a large decrease in Sox2 expression, similar to full SCR deletion 70% decrease
- ▼ Deletion of the OCT4:SOX2 motif in SRR107 reduces Sox2 transcript levels close to the level seen with full SRR107 loss; off-target microdeletions retaining the motif do not
- ▼ Deletion of both KLF4 motifs in SRR111 (SRR107-deleted background) significantly reduces Sox2 transcript levels
- ▼ Prior work found only a slight decrease in SCR–Sox2 interaction frequency after removing the core SRR109 CTCF motif, and acute CTCF depletion did not affect Sox2 transcription slight decrease (cited, de Wit et al. 2015)
- pvalue P = 0.02 (ΔSCR/ΔSCR vs WT relative contact frequency between SCR-proximal bait and Sox2)
- fold_change 28% decrease (ΔSCR/ΔSCR vs WT relative SCR–Sox2 contact frequency)
- pvalue P = 0.04 (ΔSCR allele in heterozygous ΔSCR cells vs WT)
- fold_change 24% reduction (ΔSCR allele in heterozygous ΔSCR cells vs WT levels)
- pvalue P = 0.46 (WT allele in heterozygous ΔSCR cells vs WT (no significant difference))
- pvalue P = 0.74 (comparison of homozygous vs heterozygous ΔSCR reduction in Sox2–SCR interaction)
- fold_change 27% reduction (Sox2 transcript levels from 129 allele in ΔSRR107/+ clones)
- fold_change 70% decrease (allele-specific Sox2 expression in ΔSRR107+111/+ clones)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper compares chromatin interaction frequencies (via allele-specific 4C-seq) and allele-specific Sox2 transcript levels (via RT-qPCR) between wild-type ESCs and a series of CRISPR/Cas9-generated heterozygous and homozygous deletion lines (SCR, SRR107, SRR111, SRR109, and specific transcription-factor motifs). Results are reported as relative interaction/expression values with exact P-values for the 4C comparisons and significance-threshold asterisks for the qPCR comparisons, based on biological replicate 4C experiments and independent ESC clones (n = 3–4). 4C interaction calling used a previously published fitted background statistical model (Geeven et al. 2018). The text excerpt available does not name the specific statistical test(s) used for the P-value/asterisk comparisons, nor does it describe a multiplicity-correction method.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| not explicitly named in the available text; reported as exact P-values for pairwise comparisons | 4C-seq relative contact frequency between the SCR-proximal bait and the Sox2 gene, comparing WT, homozygous ΔSCR/ΔSCR, and the WT/ΔSCR alleles of heterozygous ΔSCR cells (Fig. 1C) | n = 4 biological replicates (WT), n = 4 (ΔSCR/ΔSCR), n = 3 (WT allele, heterozygous), n = 4 (ΔSCR allele, heterozygous) | not stated |
| not explicitly named in the available text; significance denoted by asterisk thresholds (*P<0.05, **P<0.01, ***P<0.001, ****P<0.0001, ns) | Allele-specific Sox2 transcript levels (RT-qPCR) comparing deletion clones (SCR, SRR107, SRR111, SRR107+111, SRR109, OCT4:SOX2 motif, KLF4 motifs) to wild-type values (Fig. 2B–D) | n ≥ 3 independent ESC clones per genotype, as stated in the figure legend | not stated |
-
Many deletion lines/subregions (SCR, SRR107, SRR111, SRR109, and several transcription-factor motifs) are each compared individually against wild-type across multiple figures without a stated multiple-comparison correction.↳ Could also: A one-way ANOVA (or mixed-model equivalent) across all genotypes with a post-hoc correction such as Dunnett's test (each group vs. control) or a Benjamini-Hochberg FDR adjustment across the full set of comparisons — This would control the family-wise or false-discovery error rate across the many pairwise deletion-vs-wild-type comparisons performed throughout the study, which is a standard consideration when numerous related comparisons are drawn from the same experimental system.
-
Variability in the RT-qPCR allele-specific expression data (Fig. 2) is summarized using SD with small per-group sample sizes (n ≥ 3).↳ Could also: Reporting a 95% confidence interval alongside or instead of SD — A CI directly conveys the precision of the estimated group difference and can be more informative than SD alone when group sizes are small.
-
Sample sizes are reported as the number of biological replicates or independent clones per genotype, without a stated a priori power or sample-size justification.↳ Could also: An a priori power analysis or post hoc reporting of achieved effect size with its CI — This would help contextualize whether the chosen n was well suited to detect the effect sizes ultimately observed (e.g., the 14%–70% transcript reductions described).
-
Statistical significance in the qPCR figures is communicated via categorical asterisk thresholds (*, **, ***, ****, ns) rather than exact P-values.↳ Could also: Reporting exact P-values for all comparisons, as was done for the 4C data — Exact P-values preserve more information than threshold categories and facilitate later meta-analysis or reinterpretation by other researchers.
-
4C-seq interaction calling relies on a previously published fitted background statistical model (Geeven et al. 2018) to identify above-background contacts.↳ Could also: Alternative 4C/Hi-C analysis frameworks (e.g., FourCSeq or 4Cker, which use negative-binomial or wavelet-based statistical models) — Different background-modeling approaches make different assumptions about count noise and distance-decay behavior, so comparing results across frameworks can illustrate how sensitive the called interactions are to the chosen statistical model.
-
Allele-specific Sox2 transcription is quantified via RT-qPCR ratios between the 129 and CAST alleles for a targeted set of clones.↳ Could also: Genome-wide allele-specific RNA-seq quantification analyzed with a count-based generalized linear model (e.g., as implemented in DESeq2 or edgeR for allelic counts) — A GLM-based framework can jointly model biological variability and sequencing depth across many loci simultaneously, which is a complementary approach to locus-specific RT-qPCR for allele-specific expression analysis.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35710138
Paper: Blanco-Pose / Sikorska, Sexton et al., "Transcriptional regulation and chromatin architecture maintenance are decoupled functions at the Sox2 locus." Genes Dev 2022. PMID 35710138 · PMC9296009 · DOI 10.1101/gad.349489.122.
Data: GEO GSE195906 — allele-specific 4C-seq of the mouse Sox2 locus
in F1 hybrid (129 × castaneus) mouse ESCs. 4 baits (nearSCR/SCR-proximal, SCR,
SOX9/human-insertion, Sox2 promoter), 14–16 genotypes, 1–4 replicates → 69 GSM
samples, each with raw fastq (SRA SRX...) and a processed .cis file.
Code (P16, third-party + author utilities):
najoshi/sabre(MIT) — fastq demultiplexing by bait primer (-m 2).TomSexton00/4See(GPLv3, paper's last author) —utils/makefrags.pl(restriction-fragment map from fasta) +utils/coord2frag.pl(map aligned reads →.cis), and the 4See R browser for visualization.- Bowtie v1.0.0 (
-a -m 1 --best --strata), mm10. - peakC (Geeven et al. 2018) for interaction calling (window = 21 fragments).
Exact documented pipeline (from Methods + GEO data_processing)
sabredemux fastq by bait primer sequence, 2 mismatches (-m 2).bowtiev1.0.0-a -m 1 --best --strata→ mm10, native.mapoutput.coord2frag.pl <frags> <map> chr=3 coord=4 strand=2 <cis> <chrom> <vp> <name> <rlen>→ assigns intrachromosomal reads to DpnII (GATC) fragments; drops fragments whose count > 2% of kept reads (self-ligation/PCR); emits.cis= header (name chrCHR vp) + rows (frag_midpoint count). Fragment map frommakefrags.pl GATC <mm10_fasta_dir> <frag_len> <out>.- peakC: call interactions per replicate, window = 21 fragments, filter.
- 4See: visualize profiles, quantify interaction frequencies in defined windows.
IN SCOPE (pipeline-derived, attempted)
- R1 —
.cisregeneration (core 1:1): regenerate the.cisprofile from raw fastq for selected WT (and ΔSCR) samples via the exact pipeline above; compare per-fragment counts to the deposited.cis(Spearman/Pearson; bait/self-lig peak position; total kept reads). This is the literal pipeline output deposited in GEO, so it is a direct, auditable data-level reproduction. - R2 — peakC interaction call: call interactions for WT SCR-proximal-bait replicates with peakC (window 21 frags); check the SCR-proximal ↔ Sox2 gene interaction is reproducibly identified (Fig 1C, Suppl Table S1).
- R3 — ΔSCR effect (headline quantitative claim): quantify relative
interaction of the Sox2-spanning region with the SCR-proximal bait in
ΔSCR/ΔSCR vs WT from the
.cisprofiles → paper reports −28%, P = 0.02; and downstream-of-SCR region no change, P = 0.74. Attempt the effect size from deposited + regenerated profiles (exact P depends on their window/stat choices, which are only partly specified → graded provisionally).
OUT OF SCOPE (wet-lab / manual / not pipeline)
- Genome editing / clone generation, RT-qPCR allele-specific transcript levels, RNA-FISH imaging, Western blots, Bioanalyzer QC — experimental, not reproducible from deposited sequencing data.
- 4See GUI visualizations (qualitative browser figures).
Dataset profiling (same pass)
GSE195906 profiled into data/dataset_profile.json: data_type, N reported vs
observed (69 GSM), completeness (raw + processed both present), QC checks on the
files actually handled, delivers-promised, provisional grade.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.