Genome-wide kinetic properties of transcriptional bursting in mouse embryonic stem cells.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the CORE computational result. The paper has no own analysis-code repo; the burst-kinetics method is an equation-based pipeline in Methods, reproduced here 1:1 from the paper's own deposited allele-resolved scRNA-seq (GEO GSE132589) on «our HPC». The pipeline (TMM -> per-transcript quantile.robust -> geometric-mean allele correction -> Elowitz intrinsic-noise -> b=eta^2*mu-1) runs end-to-end and lands within ~5-8% of the reported transcript-count filters (below-Poisson 27,444 vs 25,481; retained 6,284 vs 5,992) and matches the headline Fig-1G correlation (Spearman 0.875 vs reported 0.869). Two honest negatives: (R1) the deposited ASEcount matrices contain 419 cells, not the 447 the text states (28 cells absent from deposit) -- flagged as a data-vs-text discrepancy, not derivable from the deposit; (R7) the SCALE validation (fig S2) is NOT independently re-runnable from deposited data because SCALE requires ERCC spike-in counts that were not deposited, and a spike-in-free Poisson-Beta surrogate reaches only Spearman 0.50. NOT attempted: R6 (TMM-vs-scran robustness, secondary); the wet-lab validations (smFISH, flow cytometry, CRISPR KI lines, FACS screen, 4sU degradation assays) which are out of scope; molecular-determinant ChIP-seq analyses (Fig 2-3, large external data); exact per-gene degradation rates (Sharova 2009 supplementary is access-gated, so R5 used gamma=1 as the paper does for the SCALE comparison). No evidence of fabricated numeric results.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 57assessed: 2026-06-17 ⛓ c3c602e3f8c9
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-17
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusHow are the kinetic properties of transcriptional bursting (burst size and frequency) and intrinsic gene expression noise regulated at the molecular level in mammalian (mouse embryonic stem) cells?
- ★ Genome-wide transcriptional bursting kinetics (intrinsic noise, burst size, frequency) can be estimated from allele-specific scRNA-seq of hybrid mESCs method
- ★ Transcriptional bursting kinetics is regulated by a combination of promoter- and gene body–binding proteins, including PRC2 and transcription elongation factors finding
- ★ The Akt/MAPK signaling pathway regulates bursting kinetics by modulating transcription elongation efficiency mechanism
- ★ Promoter localization of PRC2 subunits correlates negatively with burst frequency and positively with burst size and intrinsic noise finding
- ★ Gene body localization of transcription elongation factors (H3K36me3, BRD4, AFF4, SPT5, CTR9) is positively correlated with burst frequency, while Pol II pausing index is negatively correlated finding
- ★ Pol II pause release (P-TEFb activity) contributes to regulation of burst size and, context-dependently, burst frequency mechanism
- Promoters with a TATA box show higher burst size and intrinsic noise than those without, with no significant difference in burst frequency finding
- Nuclear retention of mRNA does not buffer intrinsic noise for the genes tested in mESCs finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA-seq (RamDA-seq) | 129/CAST hybrid mESCs (G1 phase, on LN511) | none | allele-specific mRNA levels / intrinsic noise / mean expression | RamDA-seq |
| smFISH (including intron-specific/subcellular) | inbred GFP/iRFP knock-in mESC lines (25 genes) | GFP/iRFP reporter KI | mean transcript number, normalized intrinsic noise, mRNA nuclear retention | — |
| flow cytometry | GFP/iRFP knock-in mESC lines | reporter KI | mean protein expression level and normalized intrinsic noise | — |
| ChIP-seq (informatics analysis of public data) | mESCs | none | promoter/gene body/enhancer enrichment of transcription regulators (RPM) correlated with bursting kinetics | — |
| drug inhibition + flow cytometry | KI mESC lines (2i conditions) | DRB and flavopiridol (P-TEFb/Pol II pause-release inhibitors) | Δnormalized intrinsic noise, Δburst size, Δburst frequency | — |
| CRISPR-Cas9 knockout + Western blot + flow cytometry | Suz12 K/O derived from Dnmt3l, Dnmt3b, Peg3, Ctcf KI mESC lines | Suz12 knockout | H3K27me3 loss; Δnormalized intrinsic noise, burst size, frequency | — |
| large-scale CRISPR-Cas9 library screening | mESCs | genome-wide CRISPR knockout | regulators of transcriptional bursting/elongation (Akt/MAPK pathway) | — |
| CRISPR-Cas9 genome editing (reporter knock-in) | inbred mESC line | GFP and iRFP knock-in into both alleles of target genes | allele-specific reporter expression | — |
- ▲ Normalized intrinsic noise positively correlated with the ratio of burst size to burst frequency Spearman's rho = 0.869
- ▲ smFISH-based mean expression and intrinsic noise correlated significantly with scRNA-seq measurements, validating the method
- – PRC2 promoter localization correlates inversely with burst frequency and positively with burst size and intrinsic noise
- – Suz12 K/O significantly reduced intrinsic noise and burst size of Dnmt3l and Dnmt3b but increased them for Peg3; no change for Ctcf
- – Gene body elongation factors positively correlate with burst frequency; Pol II pausing index negatively correlates with burst frequency
- ▲ DRB and flavopiridol treatment increased normalized intrinsic noise and burst size in most cell lines, with gene-specific effects on burst frequency
- ▲ TATA box–containing promoters had significantly higher burst size and intrinsic noise but no significant burst frequency difference
- – No correlation observed between mRNA nuclear retention rate and intrinsic noise, total noise, or normalized intrinsic noise
- correlation Spearman's rho = 0.869 (normalized intrinsic noise vs burst size/burst frequency ratio)
- count 447 (individual 129/CAST hybrid mESCs analyzed by scRNA-seq)
- count 25 genes (medium-expression genes selected for GFP/iRFP KI validation)
- count top and bottom 5% (transcripts defined as high and low intrinsic noise)
- count mean read count < 20 (low-abundance transcripts excluded from analysis)
- count 10 clusters (genetic/epigenetic feature clusters of high intrinsic noise transcripts; OPLS-DA separated 8 of 10)
- pvalue P < 0.05 (significance threshold for Suz12 K/O and inhibitor effects)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used single-cell RNA-seq (RamDA-seq) of 447 allele-resolved hybrid mESCs to estimate intrinsic noise, mean mRNA levels, and transcriptional bursting kinetics (burst size and frequency), then related these to ChIP-seq binding patterns. Associations were assessed primarily with Spearman's rank correlations, group comparisons (e.g., TATA box present versus absent) were reported as significant differences, and multivariate feature contributions were explored with OPLS-DA modeling and S-plots. Perturbation effects (inhibitor treatment, Suz12 knockout) were reported as residual differences with 95% confidence intervals and significance thresholds of P < 0.05.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Spearman's rank correlation | relationship between normalized intrinsic noise and burst size/frequency ratio (rho = 0.869); correlations between bursting kinetics and ChIP-seq enrichment at promoter/gene body/enhancer (Fig. 2E); validation correlations (Fig. 1, I–L) | — | not stated |
| Group comparison reported as significant difference (test type not specified in available text) | burst size, burst frequency, and normalized intrinsic noise for genes with versus without a TATA box (Fig. 2, A–C) | — | not stated |
| OPLS-DA (orthogonal partial least squares discriminant analysis) with S-plot | separating high versus low intrinsic noise transcripts across 10 clusters based on promoter/gene body features (Fig. 3) | — | not stated |
| Comparison of residual effects with significance threshold P < 0.05 (specific test not stated in available text) | effect of DRB/flavopiridol treatment and Suz12 K/O on Δnormalized intrinsic noise, Δburst size, Δburst frequency (Fig. 2, F and G) | — | not stated |
-
Associations between bursting kinetics and binding factors were quantified with Spearman's rank correlation.↳ Could also: Partial or multiple-regression approaches (or partial correlations) could also be used. — These would additionally describe the joint and independent contributions of co-occurring factors while accounting for shared variance, complementing the pairwise rank correlations.
-
Many genome-wide correlations and multiple per-gene perturbation comparisons were reported with P < 0.05 thresholds.↳ Could also: A multiplicity-control procedure such as Benjamini-Hochberg FDR or Bonferroni could also be applied across the family of tests. — Reporting adjusted P values or q-values would convey the expected proportion of false positives across the large number of simultaneous comparisons.
-
Group differences (e.g., TATA box present versus absent) were reported as significant.↳ Could also: Reporting the specific test name, exact P values, and an accompanying effect size (e.g., difference in medians or rank-biserial correlation) could also be done. — This would make the magnitude and the inferential basis of each comparison transparent and reproducible.
-
Perturbation effects were summarized with 95% confidence intervals on residual differences.↳ Could also: Showing the underlying per-replicate data points alongside the intervals could also be presented. — Overlaying individual measurements conveys the distribution and replicate structure directly, which is especially informative for small numbers of cell lines/replicates.
-
High and low intrinsic noise transcripts were defined as the top and bottom 5% of the normalized intrinsic noise distribution.↳ Could also: Treating intrinsic noise as a continuous variable, or testing sensitivity to alternative thresholds, could also be done. — A continuous or multi-threshold analysis would show how robust the feature associations are to the specific cutoff chosen.
-
Feature contributions distinguishing noise classes were explored with OPLS-DA and S-plots.↳ Could also: Cross-validated classifiers (e.g., penalized logistic regression or random forests with held-out validation and permutation testing) could also be used. — These would provide explicit out-of-sample performance estimates and permutation-based significance for the discriminating features.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Gene body elongation factor enrichment positively correlates with burst frequency while Pol II pausing index negatively correlates with burst frequency in mESCs.ChIP-seq mouse mesc mixed 2020×1papers★ This paper is the founder (earliest)
-
PRC2 promoter occupancy inversely correlates with burst frequency and positively correlates with burst size and intrinsic noise in mESCs.ChIP-seq mouse mesc mixed 2020×1papers★ This paper is the founder (earliest)
-
P-TEFb inhibitors DRB and flavopiridol increase normalized intrinsic noise and burst size in mESCs, with gene-specific effects on burst frequency.flow-cytometry mouse mesc up 2020×1papers★ This paper is the founder (earliest)
-
SUZ12 knockout reduces intrinsic noise and burst size at Dnmt3l and Dnmt3b loci but increases them at Peg3, with no effect at Ctcf, in mESCs.flow-cytometry mouse mesc mixed 2020×1papers★ This paper is the founder (earliest)
-
smFISH-measured mean expression and intrinsic noise correlate significantly with scRNA-seq (RamDA-seq) allele-specific measurements, validating bursting quantification in mESCs.imaging mouse mesc up 2020×1papers★ This paper is the founder (earliest)
-
mRNA nuclear retention rate does not correlate with intrinsic noise, total noise, or normalized intrinsic noise in mESCs.imaging mouse mesc none 2020×1papers★ This paper is the founder (earliest)
-
Normalized intrinsic noise positively correlates with the burst-size-to-burst-frequency ratio (Spearman rho=0.869) in mouse mESCs.scRNA-seq mouse mesc up 2020×1papers★ This paper is the founder (earliest)
-
TATA box-containing promoters exhibit significantly higher burst size and intrinsic noise but no significant difference in burst frequency compared to non-TATA promoters in mESCs.scRNA-seq mouse mesc up 2020×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- Perturbation of placental protein glycosylatio... L1 71/100
- CRISPR/Cas9 Screens Reveal Multiple Layers of... L1 No data access
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32596448
Paper: Ochiai et al. 2020, "Genome-wide kinetic properties of transcriptional bursting in mouse embryonic stem cells", Sci Adv 6:eaaz6699. PMCID PMC7299619.
Correction to auto-enriched metadata
The scaffold enriched code_url=ExpressionAnalysis/ea-utils and data=GSE48895.
Both are wrong as the paper's primary artifacts:
ea-utils(Fastq-mcf) is only an adapter-trimming tool the authors used; not the analysis core. There is no authors' own analysis-code repository — the burst-kinetics method is described as equations in Methods (P16 case: reproduce the described pipeline, not a shipped repo).GSE48895is an externally referenced GRO-seq dataset, not the paper's data.- Paper's OWN data: GEO GSE132593 (SuperSeries) = GSE132589 (allele-specific
scRNA-seq, 129/CAST hybrid mESC), GSE132590 (CRISPR screen), GSE132591/2 (bulk
RNA-seq). Figure source data: Mendeley
10.17632/5rchtsps3z.1.
What the paper computes (pipeline-derived results)
Transcriptional bursting kinetics estimated from allele-resolved single-cell transcript counts via an intrinsic-noise model (Elowitz two-reporter decomposition, here the two alleles are the two reporters):
- total noise: η_tot² = (⟨a1²+a2²⟩ − 2⟨a1⟩⟨a2⟩) / (2⟨a1⟩⟨a2⟩)
- intrinsic noise: η_int² = ⟨(a_gn1 − a_gn2)²⟩ / (2⟨a_gn1⟩⟨a_gn2⟩) (after 3-step normalization: TMM (edgeR) → quantile.robust between alleles (preprocessCore) → geometric-mean allele correction)
- burst size: b = η_int²·μ − 1
- burst frequency: f = μ·γ_m / (η_int²·μ − 1), γ_m = mRNA degradation rate
- normalized intrinsic noise = residual of regression of log(η_int²) on log(mean count) [+ log(transcript length)]
IN SCOPE (computational, reproducible from processed GEO data)
| # | result | inputs | reproducible from |
|---|---|---|---|
| R1 | N cells analysed = 447 (474 − 27 abnormal) | ASEcount matrices | GSE132589 column count |
| R2 | 25,481 transcripts with intrinsic noise below Poisson floor (excluded) | matrices | full pipeline |
| R3 | 5,992 transcripts retained (mean count ≥20, above Poisson) | matrices | full pipeline |
| R4 | burst size b & intrinsic noise per transcript | matrices | b = η²μ−1 (degradation-independent) |
| R5 | Spearman ρ = 0.869 (normalized intrinsic noise vs burst size/frequency ratio), Fig 1G | matrices + γ_m | needs Sharova 2009 half-lives |
| R6 | TMM vs scran: R=0.97 (mean expr), 0.95 (int noise), 0.94 (norm int noise) | matrices | two normalizations |
| R7 | SCALE (Poisson-Beta) vs intrinsic-noise params, R>0.8 (fig S2F-K) | matrices | run SCALE v1.3.0 |
OUT OF SCOPE (wet-lab / not pipeline / not attempted)
- smFISH, flow-cytometry, immunofluorescence validation (wet-lab).
- CRISPR-Cas9 KI cell-line construction; Suz12 K/O; FACS enrichment screen (GSE132590 MAGeCK screen could be partially reproduced — secondary).
- 4sU / BRIC degradation-rate measurement experiments (wet-lab); we instead use the published Sharova 2009 half-life database (ref 23), as the paper did.
- Molecular-determinant ChIP-seq correlations (large external data; Fig 2-3).
Primary reproduction target
R1–R4 (degradation-independent, fully specified) = quick floor; then R5, R6, R7.
Data / code pointers
- data: GEO GSE132589 — GSE132589_ASEcount_G1_{129,CAST}.txt.gz (transcript×cell)
- degradation: Sharova et al. 2009 DNA Res 16:45-58 (mESC mRNA half-lives, ref 23)
- methods source: PMC7299619 full text (equations transcribed above)
- «infra» workdir: «path»
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The core Fig-1 burst-kinetics pipeline reproduces cleanly from the authors' own deposited allele-resolved scRNA-seq: the headline Fig-1G correlation matches (rho 0.869→0.875) and the transcript-count filters land within ~5-8% (25,481→27,444; 5,992→6,284), so the central conclusion holds. The two genuine issues are on the authors'/data-availability side: the paper states 447 cells but only 419 are deposited (28 cells not derivable from the shared data), and the SCALE validation (fig S2, R>0.8) is not independently re-runnable because ERCC spike-ins were never deposited (surrogate gives only 0.50). Severity is moderate — magnitude and direction of the main result are preserved and there is no fabrication suspicion — but the cell-count text/deposit mismatch and the non-reproducible validation keep this short of a clean 1:1.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.