Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Transcriptional regulation and chromatin architecture maintenance are decoupled functions at the Sox2 locus.

Genes Dev · 2022
L1 81/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
81/100
Reproducibility score
0.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 59% of all assessed papers rank 468 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (pipeline-derived results, from deposited GSE195906 .cis on «our HPC»). The deposited 4C .cis profiles are confirmed to be genuine mm10 DpnII fragment-resolved profiles: 391374/391374 (100%) fragment midpoints match an independently rebuilt mm10 chr3 DpnII (GATC) fragment map (4See makefrags/coord2frag), and the bait self-ligation peak sits 238-3929bp from the declared viewpoint (chr3:34,750,100, within the SCR) -> strong provenance, no fabrication signal (R1-structural EXACT). The headline deltaSCR claim reproduces within tolerance: paper reports a 28% decrease (P=0.02) in SCR-proximal-bait<->Sox2 contact frequency in delSCR/delSCR vs WT; our limma quantile-normalized + two-tailed t-test pipeline on the Sox2-spanning region (chr3:34,644,922-34,664,967) gives 30.4% decrease, P=0.006 (n=4 vs 4) (R3a within-tol). peakC (window=21) calls the SCR-proximal<->Sox2 interaction (22 significant fragments in the Sox2-spanning region; 3/4 WT replicates individually) (R2 within-tol). NOT attempted/clean: R3b (P=0.74 het-vs-homozygous; allele-track assignment ambiguous from GEO labels), full .cis regeneration from raw fastq (needs bait primers, Suppl Table S6), and all wet-lab/imaging/qPCR results (out of scope). Code is a valid third-party + author-utility stack (najoshi/sabre + TomSexton00/4See + deWitLab/peakC) applied to the paper's own data (P16-valid).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 56
    assessed: 2026-06-19 ⛓ 2e670348900c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

How distal regulatory elements (enhancers) control gene transcription versus chromatin topology is unclear; the paper tests, at the Sox2 locus in mouse ESCs, whether transcriptional activation and chromatin–chromatin interaction/TAD maintenance are mediated by the same or different cis-regulatory sequences.

Core claims
  • Sox2 transcriptional activation is traced almost entirely to two key transcription factor-bound regions (SRR107 and SRR111) within the SCR finding
  • Deletion of SRR107/SRR111 (the transcription-driving regions) has no effect on promoter–enhancer (SCR–Sox2) interaction frequency or TAD organization finding
  • Local chromatin architecture maintenance is distributed over multiple transcription factor-bound regions across the locus and is CTCF-independent finding
  • Ectopic chromatin loop formation that partially disrupts promoter–enhancer interactions has no effect on Sox2 transcription finding
  • Significant disruption of chromatin interaction frequency and TAD boundary insulation requires deletion of the entire SCR, not just the transcription-driving subregions finding
  • Deletion of the sole CTCF-bound site within the SCR (SRR109) does not affect chromatin topology or Sox2 transcription finding
  • SRR107 (via an OCT4:SOX2 composite motif) and SRR111 (via two KLF4 motifs) act in a partially redundant manner to drive SCR enhancer activity mechanism
  • Allele-specific CRISPR/Cas9 genome editing combined with allele-specific 4C-seq and RT-qPCR was used to dissect cis-regulatory function at the Sox2 locus method
Experimental setups
Assay System Perturbation Readout Platform
allele-specific 4C-seq F1 mouse ESCs (129 x castaneus), including ΔSCR/ΔSCR and heterozygous ΔSCR clones CRISPR/Cas9 deletion of the SCR (homozygous and heterozygous) chromatin interaction frequency across the Sox2 locus from an SCR-proximal or Sox2 promoter bait
ChIP-seq mouse ESCs none binding of CTCF, RAD21, SMC1A, MED1, EP300, and H3K27ac enrichment across the Sox2 locus
allele-specific RT-qPCR F1 mouse ESC clones (129 x CAST) CRISPR/Cas9 heterozygous deletion of SRR106, SRR107, SRR109, SRR111 individually or in combination allele-specific Sox2 transcript levels relative to total transcript
allele-specific RT-qPCR F1 mouse ESC clones microdeletion of OCT4:SOX2 motif in SRR107 or KLF4 motifs in SRR111 (on background lacking the other SRR) allele-specific Sox2 transcript levels
allele-specific ChIP mouse ESCs with OCT4:SOX2 motif-deleted SRR107 OCT4:SOX2 motif deletion OCT4 and RNA polymerase II association at the altered SRR
motif discovery (JASPAR GeneReg database) in silico sequence analysis of SRR107/SRR111 none high-scoring transcription factor binding motifs (OCT4:SOX2, KLF4) JASPAR GeneReg database tool
ChIP-seq data overlap (CODEX database) compiled ESC ChIP-seq data sets none confirmation of transcription factor occupancy over identified motifs CODEX database
Key results
  • ΔSCR/ΔSCR ESCs show decreased SCR-proximal bait to Sox2 contact frequency versus wild type 28%, P=0.02
  • In heterozygous ΔSCR cells, the deleted allele shows reduced Sox2–SCR contact frequency while the WT allele is unchanged 24% reduction (P=0.04) on ΔSCR allele; P=0.46 on WT allele
  • Deletion of SRR107 alone reduces allele-specific Sox2 transcript levels 27% reduction, significant
  • Deletion of SRR111 alone causes a weaker, nonsignificant reduction in Sox2 transcript levels 14% reduction, not significant
  • Compound deletion of SRR107 and SRR111 on the same allele causes a large decrease in Sox2 expression, similar to full SCR deletion 70% decrease
  • Deletion of the OCT4:SOX2 motif in SRR107 reduces Sox2 transcript levels close to the level seen with full SRR107 loss; off-target microdeletions retaining the motif do not
  • Deletion of both KLF4 motifs in SRR111 (SRR107-deleted background) significantly reduces Sox2 transcript levels
  • Prior work found only a slight decrease in SCR–Sox2 interaction frequency after removing the core SRR109 CTCF motif, and acute CTCF depletion did not affect Sox2 transcription slight decrease (cited, de Wit et al. 2015)
Key statistics
  • pvalue P = 0.02 (ΔSCR/ΔSCR vs WT relative contact frequency between SCR-proximal bait and Sox2)
  • fold_change 28% decrease (ΔSCR/ΔSCR vs WT relative SCR–Sox2 contact frequency)
  • pvalue P = 0.04 (ΔSCR allele in heterozygous ΔSCR cells vs WT)
  • fold_change 24% reduction (ΔSCR allele in heterozygous ΔSCR cells vs WT levels)
  • pvalue P = 0.46 (WT allele in heterozygous ΔSCR cells vs WT (no significant difference))
  • pvalue P = 0.74 (comparison of homozygous vs heterozygous ΔSCR reduction in Sox2–SCR interaction)
  • fold_change 27% reduction (Sox2 transcript levels from 129 allele in ΔSRR107/+ clones)
  • fold_change 70% decrease (allele-specific Sox2 expression in ΔSRR107+111/+ clones)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper compares chromatin interaction frequencies (via allele-specific 4C-seq) and allele-specific Sox2 transcript levels (via RT-qPCR) between wild-type ESCs and a series of CRISPR/Cas9-generated heterozygous and homozygous deletion lines (SCR, SRR107, SRR111, SRR109, and specific transcription-factor motifs). Results are reported as relative interaction/expression values with exact P-values for the 4C comparisons and significance-threshold asterisks for the qPCR comparisons, based on biological replicate 4C experiments and independent ESC clones (n = 3–4). 4C interaction calling used a previously published fitted background statistical model (Geeven et al. 2018). The text excerpt available does not name the specific statistical test(s) used for the P-value/asterisk comparisons, nor does it describe a multiplicity-correction method.

Replicationbiological Sample sizeSample sizes given per figure as number of biological replicate 4C experiments (n = 3–4) or independent ESC clones (n ≥ 3) per genotype; no a priori power or sample-size calculation is described in the available text GroupsWild-type ESCs vs. CRISPR/Cas9-generated heterozygous/homozygous deletion lines (SCR, SRR107, SRR111, SRR109, and specific transcription-factor motifs), assessed by chromatin interaction frequency and allele-specific transcript levels Pairingmixed Randomization/blindingnot stated DispersionSD Exact p-valuesyes Effect sizesyes
Statistical tests used
Test Applied to n Assumptions
not explicitly named in the available text; reported as exact P-values for pairwise comparisons 4C-seq relative contact frequency between the SCR-proximal bait and the Sox2 gene, comparing WT, homozygous ΔSCR/ΔSCR, and the WT/ΔSCR alleles of heterozygous ΔSCR cells (Fig. 1C) n = 4 biological replicates (WT), n = 4 (ΔSCR/ΔSCR), n = 3 (WT allele, heterozygous), n = 4 (ΔSCR allele, heterozygous) not stated
not explicitly named in the available text; significance denoted by asterisk thresholds (*P<0.05, **P<0.01, ***P<0.001, ****P<0.0001, ns) Allele-specific Sox2 transcript levels (RT-qPCR) comparing deletion clones (SCR, SRR107, SRR111, SRR107+111, SRR109, OCT4:SOX2 motif, KLF4 motifs) to wild-type values (Fig. 2B–D) n ≥ 3 independent ESC clones per genotype, as stated in the figure legend not stated
Approaches that could also have been used
  • Many deletion lines/subregions (SCR, SRR107, SRR111, SRR109, and several transcription-factor motifs) are each compared individually against wild-type across multiple figures without a stated multiple-comparison correction.
    Could also: A one-way ANOVA (or mixed-model equivalent) across all genotypes with a post-hoc correction such as Dunnett's test (each group vs. control) or a Benjamini-Hochberg FDR adjustment across the full set of comparisons — This would control the family-wise or false-discovery error rate across the many pairwise deletion-vs-wild-type comparisons performed throughout the study, which is a standard consideration when numerous related comparisons are drawn from the same experimental system.
  • Variability in the RT-qPCR allele-specific expression data (Fig. 2) is summarized using SD with small per-group sample sizes (n ≥ 3).
    Could also: Reporting a 95% confidence interval alongside or instead of SD — A CI directly conveys the precision of the estimated group difference and can be more informative than SD alone when group sizes are small.
  • Sample sizes are reported as the number of biological replicates or independent clones per genotype, without a stated a priori power or sample-size justification.
    Could also: An a priori power analysis or post hoc reporting of achieved effect size with its CI — This would help contextualize whether the chosen n was well suited to detect the effect sizes ultimately observed (e.g., the 14%–70% transcript reductions described).
  • Statistical significance in the qPCR figures is communicated via categorical asterisk thresholds (*, **, ***, ****, ns) rather than exact P-values.
    Could also: Reporting exact P-values for all comparisons, as was done for the 4C data — Exact P-values preserve more information than threshold categories and facilitate later meta-analysis or reinterpretation by other researchers.
  • 4C-seq interaction calling relies on a previously published fitted background statistical model (Geeven et al. 2018) to identify above-background contacts.
    Could also: Alternative 4C/Hi-C analysis frameworks (e.g., FourCSeq or 4Cker, which use negative-binomial or wavelet-based statistical models) — Different background-modeling approaches make different assumptions about count noise and distance-decay behavior, so comparing results across frameworks can illustrate how sensitive the called interactions are to the chosen statistical model.
  • Allele-specific Sox2 transcription is quantified via RT-qPCR ratios between the 129 and CAST alleles for a targeted set of clones.
    Could also: Genome-wide allele-specific RNA-seq quantification analyzed with a count-based generalized linear model (e.g., as implemented in DESeq2 or edgeR for allelic counts) — A GLM-based framework can jointly model biological variability and sequencing depth across many loci simultaneously, which is a complementary approach to locus-specific RT-qPCR for allele-specific expression analysis.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35710138

Paper: Blanco-Pose / Sikorska, Sexton et al., "Transcriptional regulation and chromatin architecture maintenance are decoupled functions at the Sox2 locus." Genes Dev 2022. PMID 35710138 · PMC9296009 · DOI 10.1101/gad.349489.122.

Data: GEO GSE195906 — allele-specific 4C-seq of the mouse Sox2 locus in F1 hybrid (129 × castaneus) mouse ESCs. 4 baits (nearSCR/SCR-proximal, SCR, SOX9/human-insertion, Sox2 promoter), 14–16 genotypes, 1–4 replicates → 69 GSM samples, each with raw fastq (SRA SRX...) and a processed .cis file.

Code (P16, third-party + author utilities):

  • najoshi/sabre (MIT) — fastq demultiplexing by bait primer (-m 2).
  • TomSexton00/4See (GPLv3, paper's last author) — utils/makefrags.pl (restriction-fragment map from fasta) + utils/coord2frag.pl (map aligned reads → .cis), and the 4See R browser for visualization.
  • Bowtie v1.0.0 (-a -m 1 --best --strata), mm10.
  • peakC (Geeven et al. 2018) for interaction calling (window = 21 fragments).

Exact documented pipeline (from Methods + GEO data_processing)

  1. sabre demux fastq by bait primer sequence, 2 mismatches (-m 2).
  2. bowtie v1.0.0 -a -m 1 --best --strata → mm10, native .map output.
  3. coord2frag.pl <frags> <map> chr=3 coord=4 strand=2 <cis> <chrom> <vp> <name> <rlen> → assigns intrachromosomal reads to DpnII (GATC) fragments; drops fragments whose count > 2% of kept reads (self-ligation/PCR); emits .cis = header (name chrCHR vp) + rows (frag_midpoint count). Fragment map from makefrags.pl GATC <mm10_fasta_dir> <frag_len> <out>.
  4. peakC: call interactions per replicate, window = 21 fragments, filter.
  5. 4See: visualize profiles, quantify interaction frequencies in defined windows.

IN SCOPE (pipeline-derived, attempted)

  • R1 — .cis regeneration (core 1:1): regenerate the .cis profile from raw fastq for selected WT (and ΔSCR) samples via the exact pipeline above; compare per-fragment counts to the deposited .cis (Spearman/Pearson; bait/self-lig peak position; total kept reads). This is the literal pipeline output deposited in GEO, so it is a direct, auditable data-level reproduction.
  • R2 — peakC interaction call: call interactions for WT SCR-proximal-bait replicates with peakC (window 21 frags); check the SCR-proximal ↔ Sox2 gene interaction is reproducibly identified (Fig 1C, Suppl Table S1).
  • R3 — ΔSCR effect (headline quantitative claim): quantify relative interaction of the Sox2-spanning region with the SCR-proximal bait in ΔSCR/ΔSCR vs WT from the .cis profiles → paper reports −28%, P = 0.02; and downstream-of-SCR region no change, P = 0.74. Attempt the effect size from deposited + regenerated profiles (exact P depends on their window/stat choices, which are only partly specified → graded provisionally).

OUT OF SCOPE (wet-lab / manual / not pipeline)

  • Genome editing / clone generation, RT-qPCR allele-specific transcript levels, RNA-FISH imaging, Western blots, Bioanalyzer QC — experimental, not reproducible from deposited sequencing data.
  • 4See GUI visualizations (qualitative browser figures).

Dataset profiling (same pass)

GSE195906 profiled into data/dataset_profile.json: data_type, N reported vs observed (69 GSM), completeness (raw + processed both present), QC checks on the files actually handled, delivers-promised, provisional grade.

Figures / tables: Fig 1CTableFig S1A
C1
Reported
GSE195906: 4 baits, ~14 genotypes, 1-4 reps
Reproduced
4 baits; 16 genotype labels; 69 GSM (raw fastq + .cis each), confirmed from series matrix
within tolerance
R1_structural
Reported
deposited .cis = mm10 DpnII fragment-resolved 4C profile (sabre->bowtie mm10->4See coord2frag GATC)
Reproduced
391374/391374 (100%) deposited chr3 fragment midpoints match makefrags.pl GATC mm10 chr3 map; bait self-ligation peak 238-3929bp from declared viewpoint chr3:34,750,100
exact
R2
Reported
peakC (win 21) reproducibly calls SCR-proximal<->Sox2 interaction in WT reps
Reproduced
peakC combined.analysis (n=4 WT, wSize=21): 263 sig frags, 22 in Sox2-spanning region (peak range chr3:34,644,922-34,660,291); interaction called; 3/4 reps individually
within tolerance
R3a
Reported
delSCR/delSCR: 28% decrease, P=0.02 (SCR-proximal-bait<->Sox2 contact)
Reproduced
30.4% decrease, P=0.006 (Welch)/0.0056 (equal-var); limma quantile-norm within-1Mb, summed over Sox2 region chr3:34,644,922-34,664,967, two-tailed t-test WT(n=4) vs delSCR/delSCR(n=4)
within tolerance
R3b
Reported
P=0.74 (het deltaSCR-allele loss vs homozygous; or downstream-of-SCR no change)
Reproduced
not cleanly attemptable: het-sample allele->deltaSCR-allele track mapping ambiguous from GEO labels; two conflicting readings of the comparison
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 81/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

94.3 k
tokens (I/O) · 3.7 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.