A global change in RNA polymerase II pausing during the Drosophila midblastula transition.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the Chen/Zeitlinger 2013 eLife paper's two headline gene-list numbers with strong fidelity -- 4007/4007 second-wave (MBT) genes (exact match) and 107/117 first-wave (pre-MBT) genes (within-tolerance, corroborated by an exact 23/23 match against the paper's own shipped manual-curation spreadsheet) -- built end-to-end from raw GSE41703 FASTQ through a from-scratch Python port of the paper's bowtie1-alignment-to-gene-selection pipeline. The paper's title-level claim of a global increase in Pol II pausing during the MBT was also qualitatively and directionally confirmed via a simplified stalling-index comparison (~7-11x higher pausing at/after the MBT), though graded only 'partial' because the original figure's custom-transcript-redefinition steps (2/3) were out of scope for this room. A supporting gene-length claim partially reproduced. Roughly 15 of 26 GEO samples (TBP/histone ChIP-seq, MNase-seq) and 4 of 7 pipeline steps were not attempted -- an explicit, honestly-documented scope limitation rather than a silent gap.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusHow the differential transcription of the small set of pre-MBT genes versus the massive zygotic gene set is set up globally in the early embryo, and specifically when RNA polymerase II recruitment and pausing are first established — i.e., whether Pol II can be recruited and paused during the rapid nuclear cycles before the midblastula transition or is recruited 'just in time'.
- ★ Massive de novo recruitment of Pol II (and TBP) with widespread pausing occurs during the Drosophila midblastula transition, at 4007 promoters (~one third of all genes). finding
- ★ Only ~100 genes (117 after curation) are strongly occupied by Pol II before the MBT, and most of them show no apparent Pol II pausing. finding
- ★ The global change in Pol II pausing correlates with distinct core promoter elements, associating a TATA-enriched promoter with rapid early (pre-MBT) transcription, suggesting promoters with distinct dynamic properties are differentially used during zygotic genome activation. mechanism
- ★ Pre-MBT genes are shorter, more often intronless, and more often use the shortest available transcript than MBT-zygotic genes, consistent with a requirement for rapid transcription during the short pre-MBT nuclear cycles making lack of pausing advantageous. finding
- ★ Pre-MBT genes fall into three classes by Pol II occupancy over development: not-paused (n = 77), dual (initially non-paused then paused during/after MBT, n = 30), and paused already at pre-MBT stages (n = 10); thus pausing exists pre-MBT but is rare. finding
- ★ Among MBT-zygotic genes, 251 are actively transcribed during the MBT while 593 are paused and poised for activation at later embryonic stages (in situ first detection mostly stages 9–10). finding
- ★ H3K4me3 is not detectable before the MBT and H3K27me3 is not established at genes prior to transcription, so these histone modifications do not pre-mark genes before zygotic gene activation in Drosophila (unlike reports in zebrafish/ESCs/sperm). finding
- Manual DAPI/DIC-guided hand-sorting of embryo collections to remove out-of-stage embryos yields tightly staged pre-MBT and MBT embryo material suitable for reproducible ChIP-seq. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ChIP-seq (Pol II) | hand-sorted Drosophila melanogaster embryos: pre-MBT (nuclear cycles 8–12), MBT (nc 13–14), post-MBT control | none | Pol II occupancy/enrichment over input at TSS regions (first 200 bp), pause site (+30–50 bp) and gene bodies; pausing index | — |
| ChIP-seq (TBP) | hand-sorted Drosophila embryos, pre-MBT / MBT / post-MBT | none | TBP occupancy relative to TSS (average −20 bp upstream of TSS) | — |
| ChIP-seq (histone modifications H3K4me3, H3K27me3) | hand-sorted Drosophila embryos, pre-MBT / MBT / post-MBT | none | normalized enrichment of H3K4me3 and H3K27me3 at genes/promoters | — |
| Immunostaining (immunofluorescence) | Drosophila embryos at pre-blastoderm (nc 1–7), pre-MBT (nc 8–12) and MBT (nc 13–14) stages | none | nuclear detection of Ser5-phosphorylated Pol II CTD, unphosphorylated Pol II, TBP, H3K4me3, H3K27me3, with Lamin 0 marking nuclei (scale 20 μm) | — |
| DAPI staining with DIC/UV microscopy for embryo staging and hand-sorting | Drosophila embryo collections (1–2 hr for pre-MBT, 2–3 hr for MBT) | none | nuclear number/DAPI intensity and presence of cellularization or gastrulation furrow, used to remove out-of-stage embryos with a pipette | — |
| Public RNA-seq expression data reanalysis (Lott et al., 2011) | Drosophila embryos across nuclear cycles (e.g. nc 10, nc 14D) | none | median transcript levels (RPKM; RPKM > 1 at nc 10 for maternal, RPKM > 5 at nc 14D for active) per gene group | — |
| In situ hybridization pattern analysis (ImaGO database) | Drosophila embryos, all embryonic stages | none | stage of first detected expression for MBT active vs MBT poised genes | — |
| Computational sequence analysis (core promoter motif enrichment; phastCons conservation scoring; gene architecture from FlyBase annotation) | Drosophila melanogaster genome/promoter sequences, insect genome alignments | none | core promoter element/motif enrichment per gene class, mean phastCons conservation score per transcript, transcript length, intron content, TSS usage | — |
- ▲ Pol II and TBP are recruited de novo to 4007 promoters during the MBT, roughly a third of all genes, whereas Pol II occupies only ~100 genes before the MBT 4007 promoters (~1/3 of genes)
- – Curated list of pre-MBT Pol II-occupied genes: 117 genes (after removing 12 false positives from read-through and adding 10 with un-annotated alternative start sites), including 14 non-coding RNA precursors and genes in sex determination, cellularization, A-P and D-V patterning 117 genes
- ▼ Most pre-MBT genes show no notable Pol II enrichment at the pause site and a much lower pausing index than MBT genes, while TBP is positioned upstream of the TSS
- ▼ Pre-MBT genes are substantially shorter than MBT-zygotic genes median 1228 bp vs 6024 bp (Table 2: 6042 bp)
- ▲ Pre-MBT genes are far more often intronless and more often use the shortest transcript isoform than MBT-zygotic genes intronless 54.6% vs 9.2%; shortest-transcript usage 62.9% vs 30.6%
- ▲ MBT-zygotic genes (844 after subtracting 3163 maternally expressed MBT genes) frequently show high Pol II occupancy at the pause site and high pausing index; 251 (30%) are expressed at late nc 14 (MBT active) and 593 are MBT poised 251/844 = 30% active; 593 poised
- – Ser5-P Pol II and TBP are first detectable in nuclei only at nuclear cycles 8–12, while unphosphorylated Pol II is nuclear from the earliest cleavage stages, indicating de novo promoter recruitment at zygotic genome activation
- – H3K4me3 signal (immunostaining and ChIP-seq) only becomes detectable during the MBT; H3K27me3 is seen in nuclei and polar bodies at earliest cleavage stages but not as a pre-transcription gene mark
- count 4007 promoters (Pol II/TBP promoters recruited de novo during MBT (~1/3 of all genes))
- count 117 (curated pre-MBT genes occupied by Pol II before the MBT)
- pvalue p<10−20 (Mann–Whitney) (pre-MBT vs MBT-zygotic transcript size, median 1228 bp vs 6024 bp)
- pvalue p<10−23 (Fisher) (intronless genes, pre-MBT 54.6% (53/97) vs MBT-zygotic 9.2% (68/736))
- pvalue p<0.0003 (Fisher) (use of shortest transcript, pre-MBT 62.9% (22/35) vs MBT-zygotic 30.6% (87/284))
- count 77 not-paused, 30 dual, 10 paused (three pre-MBT gene classes by Pol II occupancy over development)
- count 3163 MBT-maternal, 844 MBT-zygotic (251 active, 593 poised) (classification of genes newly bound by Pol II during MBT)
- other at least twofold enrichment of Pol II over input at the TSS across four Pol II ChIP-seq replicates; nuclear cycle length 8 min (nc 10) to 13 min (nc 12); 5–20% out-of-stage embryos in conventional collections (pre-MBT gene calling threshold, cycle timing motivating rapid transcription, contamination removed by hand-sorting)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study compares gene groups defined by ChIP-seq/RNA-seq occupancy and expression status (pre-MBT vs MBT-zygotic genes, and further subgroups such as not-paused/dual/paused and active/poised) using nonparametric tests for gene-length and categorical-proportion differences, and assesses replicate reproducibility with Pearson correlation. Results for the length and proportion comparisons are reported as p-value thresholds (e.g., p<10^-20) alongside descriptive medians/percentages and a computed 'pausing index' summarized in violin plots.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Mann-Whitney test | Comparison of transcript (gene body) size between pre-MBT and MBT-zygotic genes (Table 2) | 117 pre-MBT genes (97 protein-coding) vs 844 MBT-zygotic genes (736 protein-coding), as listed in Table 2 | not stated |
| Fisher's exact test | Comparison of intronless gene proportion between pre-MBT and MBT-zygotic genes (54.6% vs 9.2%, Table 2) | 97 protein-coding pre-MBT genes vs 736 protein-coding MBT-zygotic genes, as listed in Table 2 | not stated |
| Fisher's exact test | Comparison of shortest-transcript TSS usage between pre-MBT and MBT-zygotic genes (62.9% vs 30.6%, Table 2) | 35 pre-MBT genes with multiple TSS vs 284 MBT-zygotic genes with multiple TSS, as listed in Table 2 | not stated |
| Pearson correlation | Reproducibility between Pol II ChIP-seq replicates across pre-MBT and MBT stages (Figure 1—figure supplement 2B) | — | not stated |
-
Transcript size differences between pre-MBT and MBT-zygotic genes were assessed with a Mann-Whitney test and reported as a p-value threshold (p<10^-20).↳ Could also: A Kolmogorov-Smirnov test comparing the full length distributions, or reporting an exact p-value with an accompanying effect-size measure such as a rank-biserial correlation — This would let readers see both the shape of the distributional difference and the magnitude of the effect alongside the significance threshold.
-
Categorical proportions (intronless genes, shortest-transcript TSS usage) were compared between gene groups with Fisher's exact test.↳ Could also: A chi-square test of proportions, reported together with an odds ratio and its confidence interval — An odds ratio with a CI would convey the size and precision of the association in addition to statistical significance.
-
P-values for the size and proportion comparisons are reported as upper-bound thresholds (e.g., p<10^-20, p<10^-23, p<0.0003).↳ Could also: Reporting the exact computed p-value — Exact values let readers gauge the precise strength of evidence, though at these extreme magnitudes the practical conclusion is unlikely to change.
-
Three separate tests (one Mann-Whitney, two Fisher's exact) were each evaluated at their own significance threshold without a stated adjustment across the set.↳ Could also: A multiple-comparison correction such as Bonferroni or Benjamini-Hochberg FDR applied across the family of comparisons — This would formally control the family-wise error rate or false discovery rate when several related comparisons are drawn from the same dataset, though the very small p-values here make the qualitative conclusions robust to such an adjustment.
-
Reproducibility between ChIP-seq replicates was quantified using Pearson correlation.↳ Could also: Spearman rank correlation or a concordance measure such as the intraclass correlation coefficient (ICC) — Spearman or ICC can be less sensitive to non-normality or outliers in read-count data, which is a common feature of sequencing-based signal, and ICC additionally partitions variance attributable to replicate vs true signal.
-
Gene-length and pausing-index distributions were summarized primarily via medians and violin plots.↳ Could also: Reporting explicit dispersion statistics such as the interquartile range (IQR) alongside the median — An explicit IQR gives readers a numeric spread value to accompany the visual distribution shown in the violin plot.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a strong reproduction. Working from raw GSE41703 FASTQ through fresh bowtie1/dm3 alignment, the room hit the paper's second-wave count exactly (4007 = 4007), landed the first-wave count at 107 vs 117 (91.5%), and the pre-MBT median gene length at 1127 bp vs 1228 bp (~8%); the central pausing claim reproduces unambiguously (mean stalling index 0.239 for pre-MBT/first-wave vs 1.674 for MBT/second-wave, ~7x; genome-wide 0.052 → 0.575, ~11x). Every residual deviation sits on our side, not the authors': an approximate duplicate-read filter, an independently rebuilt FlyBase r5.47 annotation, and — decisively — the deliberate decision to port only steps 1 and 6 of a 7-step pipeline, which meant Figure 2A used plain annotation instead of the custom transcript models and no maternal/zygotic split existed to match the paper's 844-gene 6024 bp comparator (ours, 3585 bp, is a blended figure). A notable positive-control signal: the pipeline independently regenerated a 23-gene 'no confirmed TBP peak' candidate set that exactly matches the 23 rows in the authors' own shipped manual-curation spreadsheet, showing we reach the same intermediate state the authors curated by hand. Nothing here is fabrication- or derivability-suspect.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.