Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A global change in RNA polymerase II pausing during the Drosophila midblastula transition.

Elife · 2013
L1 71/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -3
✓ What held up
  • Same input data as the authors
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the Chen/Zeitlinger 2013 eLife paper's two headline gene-list numbers with strong fidelity -- 4007/4007 second-wave (MBT) genes (exact match) and 107/117 first-wave (pre-MBT) genes (within-tolerance, corroborated by an exact 23/23 match against the paper's own shipped manual-curation spreadsheet) -- built end-to-end from raw GSE41703 FASTQ through a from-scratch Python port of the paper's bowtie1-alignment-to-gene-selection pipeline. The paper's title-level claim of a global increase in Pol II pausing during the MBT was also qualitatively and directionally confirmed via a simplified stalling-index comparison (~7-11x higher pausing at/after the MBT), though graded only 'partial' because the original figure's custom-transcript-redefinition steps (2/3) were out of scope for this room. A supporting gene-length claim partially reproduced. Roughly 15 of 26 GEO samples (TBP/histone ChIP-seq, MNase-seq) and 4 of 7 pipeline steps were not attempted -- an explicit, honestly-documented scope limitation rather than a silent gap.

💻 Code ↗ 🗄 Data: GSE41703

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

How the differential transcription of the small set of pre-MBT genes versus the massive zygotic gene set is set up globally in the early embryo, and specifically when RNA polymerase II recruitment and pausing are first established — i.e., whether Pol II can be recruited and paused during the rapid nuclear cycles before the midblastula transition or is recruited 'just in time'.

Core claims
  • Massive de novo recruitment of Pol II (and TBP) with widespread pausing occurs during the Drosophila midblastula transition, at 4007 promoters (~one third of all genes). finding
  • Only ~100 genes (117 after curation) are strongly occupied by Pol II before the MBT, and most of them show no apparent Pol II pausing. finding
  • The global change in Pol II pausing correlates with distinct core promoter elements, associating a TATA-enriched promoter with rapid early (pre-MBT) transcription, suggesting promoters with distinct dynamic properties are differentially used during zygotic genome activation. mechanism
  • Pre-MBT genes are shorter, more often intronless, and more often use the shortest available transcript than MBT-zygotic genes, consistent with a requirement for rapid transcription during the short pre-MBT nuclear cycles making lack of pausing advantageous. finding
  • Pre-MBT genes fall into three classes by Pol II occupancy over development: not-paused (n = 77), dual (initially non-paused then paused during/after MBT, n = 30), and paused already at pre-MBT stages (n = 10); thus pausing exists pre-MBT but is rare. finding
  • Among MBT-zygotic genes, 251 are actively transcribed during the MBT while 593 are paused and poised for activation at later embryonic stages (in situ first detection mostly stages 9–10). finding
  • H3K4me3 is not detectable before the MBT and H3K27me3 is not established at genes prior to transcription, so these histone modifications do not pre-mark genes before zygotic gene activation in Drosophila (unlike reports in zebrafish/ESCs/sperm). finding
  • Manual DAPI/DIC-guided hand-sorting of embryo collections to remove out-of-stage embryos yields tightly staged pre-MBT and MBT embryo material suitable for reproducible ChIP-seq. method
Experimental setups
Assay System Perturbation Readout Platform
ChIP-seq (Pol II) hand-sorted Drosophila melanogaster embryos: pre-MBT (nuclear cycles 8–12), MBT (nc 13–14), post-MBT control none Pol II occupancy/enrichment over input at TSS regions (first 200 bp), pause site (+30–50 bp) and gene bodies; pausing index
ChIP-seq (TBP) hand-sorted Drosophila embryos, pre-MBT / MBT / post-MBT none TBP occupancy relative to TSS (average −20 bp upstream of TSS)
ChIP-seq (histone modifications H3K4me3, H3K27me3) hand-sorted Drosophila embryos, pre-MBT / MBT / post-MBT none normalized enrichment of H3K4me3 and H3K27me3 at genes/promoters
Immunostaining (immunofluorescence) Drosophila embryos at pre-blastoderm (nc 1–7), pre-MBT (nc 8–12) and MBT (nc 13–14) stages none nuclear detection of Ser5-phosphorylated Pol II CTD, unphosphorylated Pol II, TBP, H3K4me3, H3K27me3, with Lamin 0 marking nuclei (scale 20 μm)
DAPI staining with DIC/UV microscopy for embryo staging and hand-sorting Drosophila embryo collections (1–2 hr for pre-MBT, 2–3 hr for MBT) none nuclear number/DAPI intensity and presence of cellularization or gastrulation furrow, used to remove out-of-stage embryos with a pipette
Public RNA-seq expression data reanalysis (Lott et al., 2011) Drosophila embryos across nuclear cycles (e.g. nc 10, nc 14D) none median transcript levels (RPKM; RPKM > 1 at nc 10 for maternal, RPKM > 5 at nc 14D for active) per gene group
In situ hybridization pattern analysis (ImaGO database) Drosophila embryos, all embryonic stages none stage of first detected expression for MBT active vs MBT poised genes
Computational sequence analysis (core promoter motif enrichment; phastCons conservation scoring; gene architecture from FlyBase annotation) Drosophila melanogaster genome/promoter sequences, insect genome alignments none core promoter element/motif enrichment per gene class, mean phastCons conservation score per transcript, transcript length, intron content, TSS usage
Key results
  • Pol II and TBP are recruited de novo to 4007 promoters during the MBT, roughly a third of all genes, whereas Pol II occupies only ~100 genes before the MBT 4007 promoters (~1/3 of genes)
  • Curated list of pre-MBT Pol II-occupied genes: 117 genes (after removing 12 false positives from read-through and adding 10 with un-annotated alternative start sites), including 14 non-coding RNA precursors and genes in sex determination, cellularization, A-P and D-V patterning 117 genes
  • Most pre-MBT genes show no notable Pol II enrichment at the pause site and a much lower pausing index than MBT genes, while TBP is positioned upstream of the TSS
  • Pre-MBT genes are substantially shorter than MBT-zygotic genes median 1228 bp vs 6024 bp (Table 2: 6042 bp)
  • Pre-MBT genes are far more often intronless and more often use the shortest transcript isoform than MBT-zygotic genes intronless 54.6% vs 9.2%; shortest-transcript usage 62.9% vs 30.6%
  • MBT-zygotic genes (844 after subtracting 3163 maternally expressed MBT genes) frequently show high Pol II occupancy at the pause site and high pausing index; 251 (30%) are expressed at late nc 14 (MBT active) and 593 are MBT poised 251/844 = 30% active; 593 poised
  • Ser5-P Pol II and TBP are first detectable in nuclei only at nuclear cycles 8–12, while unphosphorylated Pol II is nuclear from the earliest cleavage stages, indicating de novo promoter recruitment at zygotic genome activation
  • H3K4me3 signal (immunostaining and ChIP-seq) only becomes detectable during the MBT; H3K27me3 is seen in nuclei and polar bodies at earliest cleavage stages but not as a pre-transcription gene mark
Key statistics
  • count 4007 promoters (Pol II/TBP promoters recruited de novo during MBT (~1/3 of all genes))
  • count 117 (curated pre-MBT genes occupied by Pol II before the MBT)
  • pvalue p<10−20 (Mann–Whitney) (pre-MBT vs MBT-zygotic transcript size, median 1228 bp vs 6024 bp)
  • pvalue p<10−23 (Fisher) (intronless genes, pre-MBT 54.6% (53/97) vs MBT-zygotic 9.2% (68/736))
  • pvalue p<0.0003 (Fisher) (use of shortest transcript, pre-MBT 62.9% (22/35) vs MBT-zygotic 30.6% (87/284))
  • count 77 not-paused, 30 dual, 10 paused (three pre-MBT gene classes by Pol II occupancy over development)
  • count 3163 MBT-maternal, 844 MBT-zygotic (251 active, 593 poised) (classification of genes newly bound by Pol II during MBT)
  • other at least twofold enrichment of Pol II over input at the TSS across four Pol II ChIP-seq replicates; nuclear cycle length 8 min (nc 10) to 13 min (nc 12); 5–20% out-of-stage embryos in conventional collections (pre-MBT gene calling threshold, cycle timing motivating rapid transcription, contamination removed by hand-sorting)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study compares gene groups defined by ChIP-seq/RNA-seq occupancy and expression status (pre-MBT vs MBT-zygotic genes, and further subgroups such as not-paused/dual/paused and active/poised) using nonparametric tests for gene-length and categorical-proportion differences, and assesses replicate reproducibility with Pearson correlation. Results for the length and proportion comparisons are reported as p-value thresholds (e.g., p<10^-20) alongside descriptive medians/percentages and a computed 'pausing index' summarized in violin plots.

Replicationbiological Sample sizeGene counts underlying each comparison are stated (e.g., 117 pre-MBT genes; 844 MBT-zygotic genes; subgroups of 77/30/10 and 251/593), but no formal sample-size or power calculation is described Groupspre-MBT genes vs MBT-zygotic genes, and subgroups (not-paused/dual/paused; active/poised) Pairingunpaired Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Mann-Whitney test Comparison of transcript (gene body) size between pre-MBT and MBT-zygotic genes (Table 2) 117 pre-MBT genes (97 protein-coding) vs 844 MBT-zygotic genes (736 protein-coding), as listed in Table 2 not stated
Fisher's exact test Comparison of intronless gene proportion between pre-MBT and MBT-zygotic genes (54.6% vs 9.2%, Table 2) 97 protein-coding pre-MBT genes vs 736 protein-coding MBT-zygotic genes, as listed in Table 2 not stated
Fisher's exact test Comparison of shortest-transcript TSS usage between pre-MBT and MBT-zygotic genes (62.9% vs 30.6%, Table 2) 35 pre-MBT genes with multiple TSS vs 284 MBT-zygotic genes with multiple TSS, as listed in Table 2 not stated
Pearson correlation Reproducibility between Pol II ChIP-seq replicates across pre-MBT and MBT stages (Figure 1—figure supplement 2B) not stated
Approaches that could also have been used
  • Transcript size differences between pre-MBT and MBT-zygotic genes were assessed with a Mann-Whitney test and reported as a p-value threshold (p<10^-20).
    Could also: A Kolmogorov-Smirnov test comparing the full length distributions, or reporting an exact p-value with an accompanying effect-size measure such as a rank-biserial correlation — This would let readers see both the shape of the distributional difference and the magnitude of the effect alongside the significance threshold.
  • Categorical proportions (intronless genes, shortest-transcript TSS usage) were compared between gene groups with Fisher's exact test.
    Could also: A chi-square test of proportions, reported together with an odds ratio and its confidence interval — An odds ratio with a CI would convey the size and precision of the association in addition to statistical significance.
  • P-values for the size and proportion comparisons are reported as upper-bound thresholds (e.g., p<10^-20, p<10^-23, p<0.0003).
    Could also: Reporting the exact computed p-value — Exact values let readers gauge the precise strength of evidence, though at these extreme magnitudes the practical conclusion is unlikely to change.
  • Three separate tests (one Mann-Whitney, two Fisher's exact) were each evaluated at their own significance threshold without a stated adjustment across the set.
    Could also: A multiple-comparison correction such as Bonferroni or Benjamini-Hochberg FDR applied across the family of comparisons — This would formally control the family-wise error rate or false discovery rate when several related comparisons are drawn from the same dataset, though the very small p-values here make the qualitative conclusions robust to such an adjustment.
  • Reproducibility between ChIP-seq replicates was quantified using Pearson correlation.
    Could also: Spearman rank correlation or a concordance measure such as the intraclass correlation coefficient (ICC) — Spearman or ICC can be less sensitive to non-normality or outliers in read-count data, which is a common feature of sequencing-based signal, and ICC additionally partitions variance attributable to replicate vs true signal.
  • Gene-length and pausing-index distributions were summarized primarily via medians and violin plots.
    Could also: Reporting explicit dispersion statistics such as the interquartile range (IQR) alongside the median — An explicit IQR gives readers a numeric spread value to accompany the visual distribution shown in the violin plot.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

first_wave_gene_count
Reported
117
Reproduced
107
within tolerance
second_wave_gene_count
Reported
4007
Reproduced
4007
exact
global_pausing_increase_fig2a
Reported
Paper's central/title claim, visualized in Figure 2A: 'pre-MBT genes are indeed much less paused than MBT genes' -- a global increase in Pol II promoter-proximal pausing (stalling index si = max(tss.ratio,0) - max(dst.ratio,0)) occurring de novo during the MBT.
Reproduced
{'fw_genes_preMBT_mean_si': 0.2389, 'fw_genes_preMBT_median_si': 0.2201, 'n_fw_genes_with_si': 107, 'sw_genes_MBT_mean_si': 1.6739, 'sw_genes_MBT_median_si': 1.5968, 'n_sw_genes_with_si': 4002, 'all_transcripts_preMBT_mean_si': 0.052, 'all_transcripts_MBT_mean_si': 0.5748}
partial
gene_length_difference
Reported
{'preMBT_median_bp': 1228, 'comparison_group_median_bp': 6024}
Reproduced
{'fw_median_tx_width_bp': 1127.0, 'sw_median_tx_width_bp': 3585.0}
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -3

This is a strong reproduction. Working from raw GSE41703 FASTQ through fresh bowtie1/dm3 alignment, the room hit the paper's second-wave count exactly (4007 = 4007), landed the first-wave count at 107 vs 117 (91.5%), and the pre-MBT median gene length at 1127 bp vs 1228 bp (~8%); the central pausing claim reproduces unambiguously (mean stalling index 0.239 for pre-MBT/first-wave vs 1.674 for MBT/second-wave, ~7x; genome-wide 0.052 → 0.575, ~11x). Every residual deviation sits on our side, not the authors': an approximate duplicate-read filter, an independently rebuilt FlyBase r5.47 annotation, and — decisively — the deliberate decision to port only steps 1 and 6 of a 7-step pipeline, which meant Figure 2A used plain annotation instead of the custom transcript models and no maternal/zygotic split existed to match the paper's 844-gene 6024 bp comparator (ours, 3585 bp, is a blended figure). A notable positive-control signal: the pipeline independently regenerated a 23-gene 'no confirmed TBP peak' candidate set that exactly matches the 23 rows in the authors' own shipped manual-curation spreadsheet, showing we reach the same intermediate state the authors curated by hand. Nothing here is fabrication- or derivability-suspect.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.