Differentiating Drosophila female germ cells initiate Polycomb silencing by regulating PRC2-interacting proteins.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for full reproduction; result is 1:1 (different code from authors' is not needed — used their own shipped DESeq2 script PcGRNAseq_Methods.R on their own deposited GEO counts). The repo is a single fully-specified DESeq2 pipeline over per-sample HTSeq counts (GSE145282_RAW.tar) with a shipped gene->chromatin-class table. All 32 referenced RNA-seq count files were present with a clean 1:1 name mapping; all 15 DESeq2 comparisons ran and regenerated the Figure-7 median-log2FC-per-class values. Headline claims reproduce: E(z) germline KD derepresses inactive (1.37x) and PcG (1.40x) genes by a median ~1.4-fold while active genes stay flat; five other PcG-component knockdowns (Pcl/Scm/Sce/Pc/Jarid2) show no widespread change; the developmental decline of inactive/PcG relative to active genes is ~2-fold by stage 6 (bam->young ovary) and ~6-fold by stage 14 (bam->whole ovary), matching the paper's 'twofold'/'sixfold' (2^2.58=5.98). NOT attempted (out of scope): Hisat2 re-alignment from SRA FASTQ (started from deposited counts as the repo expects), ChIP-seq H3K27me3 4-state domain calling that DEFINES the Active/Inactive/PcG labels (consumed the shipped classification table as input), and all wet-lab/imaging/genetics results. Caveat: paper gives rounded prose values, not a decimal table, so 'exact' means the precise reproduced median rounds to the published figure; human reviewer confirms vs the Figure-7 panels. No possible-fabrication flag raised.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 87assessed: 2026-06-18 ⛓ ec7cc3b24917
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper asks how Polycomb (PRC2-dependent H3K27me3) silencing is developmentally initiated in naive precursor cells, using Drosophila female germ cell differentiation (GSCs to nurse cells) as a tractable model for the transition from a non-canonical to canonical Polycomb state.
- ★ Drosophila female germline stem cells lack canonical Polycomb silencing and have a non-canonical H3K27me3 distribution resembling early embryos. finding
- ★ PRC2-dependent silencing initiates during nurse cell differentiation in both inactive and PcG domains and strengthens as nurse cells grow. finding
- ★ Developmentally controlled expression of two PRC2-interacting proteins, Pcl and Scm, initiates silencing during differentiation; abundant Pcl in GSCs globally inhibits PRC2-dependent silencing, while declining Pcl plus newly induced Scm in nurse cells concentrates PRC2 activity on traditional Polycomb domains. mechanism
- ★ Nurse cells acquire a highly similar collection of H3K27me3-enriched PcG domains as embryonic somatic cells, indicating canonical Polycomb silencing. finding
- ★ A heat-shock-inducible GFP (hsGFP) reporter integrated into MiMIC sites genome-wide measures developmental gene silencing at thousands of loci with single-cell resolution. method
- Accessory proteins regulate PRC2 either by increasing PRC2 concentration at target sites or by inhibiting the rate that PRC2 samples chromatin. mechanism
- FACS purification of fixed, fluorescently labeled ovarian nuclei across DNA-content stages enables ChIPseq of defined germline cell types and stages. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| hsGFP heat-shock-inducible silencing reporter (fluorescence imaging) | Drosophila ovarian germline (nurse cells, oocytes, GSC/precursors) across developmental stages | heat shock induction; reporter integrated in active/inactive/PcG/Hp1 chromatin domains | GFP fluorescence induction ([GFP]+hs – [GFP]-hs) per cell type/stage | MiMIC transposons with FLP and phiC31 site-specific recombination |
| Germline-specific RNAi knockdown (GLKD) | Drosophila ovarian nurse cells/oocytes with hsGFP reporters | GLKD of E(z), Scm, or w (control) | effect on reporter GFP induction (de-repression) in PcG/inactive/active domains | — |
| ChIPseq (H3K27me3, H3K27me2, H3K27ac, total H3, input) | FACS-purified ovarian nuclei: GSCs (bam mutant), nurse cells (2c-512c), follicle cells (somatic control) | none (cell-type/stage comparison); bam mutation to block differentiation for GSC purification | genome-wide RPM-normalized read depth and IP/Input enrichment over chromatin domains | FACS (FACSDiva) purification of fixed GFP/tdTomato-labeled nuclei with DAPI DNA content |
- ▼ Reporters in PcG (Antp) and inactive (OR67D) domains were non-inducible in stage 9-10 nurse cells, while the active-domain (Dak1) reporter was strongly induced.
- ▲ E(z) GLKD relieved repression of reporters near Antp and OR67D but had no effect at Dak1.
- – Reporters showed little difference between active and repressed loci in pre-meiotic germ cells; E(z)-dependent silencing appeared by stage 1-2 in inactive/most PcG domains and in all PcG domains by stage 6, strengthening over time.
- ▲ H3K27me3 was highly enriched on PcG domains in late (512c) nurse cells, with lower but significant enrichment on inactive domains co-enriched for H3K27me2.
- ▲ 117 of 118 PcG domains were highly H3K27me3-enriched in nurse cells and 112 of 118 in follicle cells. 117/118 and 112/118
- – 111 of 130 PcG domains (85%) were shared among nurse cells, follicle cells, and S2 cells. 85% (111/130)
- ▼ A few PcG domains depleted of H3K27me3 in nurse/follicle cells (fusilli, eyes absent, tramtrack, 18 wheeler, apontic) have known oogenesis functions.
- count 117 and 112 of 118 PcG domains H3K27me3-enriched (PcG domains enriched in nurse cells (117) and follicle cells (112))
- count 111 out of 130 (85%) (PcG domains shared between nurse cells, follicle cells, and S2 cells)
- count 109 of 300 sites (successful reporter integrations from targeted MiMIC sites)
- count 12 PcG, 3 active, 5 inactive reporter lines (quantitative reporter lines presented in the paper)
- count 7 additional H3K27me3-enriched domains each (two shared) (domains in nurse/follicle cells annotated as active/inactive in S2 cells)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study uses a reporter gene system (hsGFP inserted at pre-selected MiMIC loci in distinct chromatin domains) to quantify Polycomb silencing in Drosophila female germline cells across developmental stages; fluorescence after heat shock is expressed as mean GFP induction ([GFP]+hs – [GFP]-hs) with standard-deviation shading for visual comparison across chromatin domain types and genotypes. Genome-wide chromatin state is assessed by ChIPseq on FACS-purified nuclei from defined cell types, with enrichment reported as RPM-normalized IP/Input ratios compared visually via genome-browser tracks, scatter plots, and histograms. The provided text does not describe a formal inferential statistics section or named statistical tests; conclusions are drawn from directional differences in reporter induction and ChIPseq enrichment between domain types, genotypes, and developmental stages.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Mean GFP induction difference ([GFP]+hs – [GFP]-hs) plotted with SD shading — descriptive summary, no formal test named | Reporter induction across 12 developmental stages for 20 reporter lines grouped by chromatin domain type (Figure 1G–H) | — | not stated |
| ChIPseq IP/Input enrichment ratio (RPM-normalized) compared visually between cell types and chromatin domain classes | H3K27me3, H3K27me2, H3K27ac enrichment across PcG, inactive, and active domains in GSCs, nurse cells, and follicle cells (Figure 2C–I) | — | not stated |
-
Differences in reporter induction between chromatin domain types and genotypes are assessed visually from mean ± SD plots without a formal inferential test↳ Could also: A linear mixed-effects model (e.g., in R/lme4) with domain type and genotype as fixed effects and reporter line as a random effect could formally test these differences — A mixed-effects approach would account for the repeated-measures structure (same reporter lines measured across stages and genotypes), quantify effect sizes, and yield FDR-controlled p-values for domain-type × genotype contrasts
-
ChIPseq enrichment differences between cell types are assessed by visual inspection of genome-browser tracks and IP/Input scatter plots, without stated differential-enrichment testing↳ Could also: Tools such as DiffBind, DESeq2-on-counts, or csaw applied to ChIPseq peak read counts could also formally quantify differential H3K27me3 enrichment between cell types with FDR control — Quantitative differential ChIPseq analysis provides effect-size estimates and FDR-corrected statistics for each domain, complementing visual inspection and enabling replication benchmarking
-
The number of biological replicates for fluorescence measurements and ChIPseq experiments is not stated in the provided text↳ Could also: Explicitly reporting the number of independent biological replicates per cell type, along with replicate-level concordance metrics (e.g., Pearson r of IP/Input scores, or IDR for ChIPseq peaks), is standard practice — Replicate counts and concordance metrics allow readers to assess reproducibility, are required by ENCODE and many journals, and are needed to apply statistical models appropriately
-
Dispersion for GFP fluorescence data is displayed as SD↳ Could also: 95% confidence intervals or SEM could also convey the precision of mean estimates across stages — SD describes variability among individual measurements; CI or SEM describes uncertainty in the mean and is often preferred when the primary goal is comparing group means, especially with small n
-
No sample-size rationale or power calculation is described in the provided text↳ Could also: A prospective power calculation based on the expected effect size (e.g., fold-change in reporter induction between domain types) and estimated within-group SD could be stated — Power calculations help readers evaluate whether the study is sized to detect biologically meaningful differences and are increasingly required by journals for quantitative imaging and genomics studies
-
Multiple t-test-like pairwise comparisons (multiple domain types × multiple genotypes × multiple stages) are implicit in the visual analysis without stated family-wise error control↳ Could also: A single two-way ANOVA (domain type × genotype) with a post-hoc correction such as Tukey HSD or Benjamini-Hochberg FDR could also control the family-wise error rate across the full comparison family — When many pairwise visual comparisons are made simultaneously, the probability of at least one false positive increases; a correction method applied to the whole family of comparisons maintains the stated error rate
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32773039
Paper: DeLuca SZ, Ghildiyal M, Pang LY, Spradling AC. Differentiating Drosophila female germ cells initiate Polycomb silencing by regulating PRC2-interacting proteins. eLife 2020;9:e56922. PMID 32773039 · PMCID PMC7438113.
Code: https://github.com/ciwemb/polycomb-development (single R script
PcGRNAseq_Methods.R + metadata/ sample sheets + gene→ME_type table).
Data: GEO GSE145282 (73 samples; RNA-seq + ChIP-seq of Drosophila ovaries).
The shipped computational pipeline (in scope)
PcGRNAseq_Methods.R is a fully-specified DESeq2 differential-expression
analysis driven by HTSeq raw-count files (*.htseq) deposited at GEO:
getResults(sample, reference):DESeqDataSetFromHTSeqCount(sampleTable, dir, ~condition)→relevel(condition, ref=reference)→DESeq()→results()→ merge withpolyATranscripts_MEtype.txtto label each gene'sME_type∈ {Active, Inactive, PcG} (classification from the H3K27me3 4-state chromatin model).- Per comparison it produces the values plotted in Figure 7:
- 7A scatter (log2FC vs log10 baseMean, colored by ME_type) + boxplots.
- 7B bar graph data = median
log2FoldChangeand a 95% CI (1.58*IQR/sqrt(n)) for each ME_type class — these are the printedbarDatatables (concrete numeric reproduction targets). - Ribbon plots (staged median FC), and
summary()reference stats.
15 DESeq2 comparisons are defined (Ezbam, Ez, scm, sce, Pc, pcl, Jarid, and the bam-vs-staged contrasts bamYO/bamEzYO/bamWO/bamPcWO/bampclWO/bamsceWO/bamscmWO/ bamjarid2WO).
IN SCOPE (pipeline-derived, attempted)
- Re-run the DESeq2 pipeline on the GEO HTSeq counts.
- Reproduce the median log2FoldChange per ME_type (the Fig 7B bar-graph values) for each comparison — the quantitative core.
- Reproduce the per-comparison gene-class counts (n Active/Inactive/PcG).
Reported claims targeted
- C1: "E(z)GLKD upregulated the majority of genes in both inactive and PcG
domains by a median 1.4-fold" (log2 ≈ +0.485) — the
Ez(young ovary, Ez vs luc) comparison, Inactive & PcG classes. - C2: bam-ovary E(z) knockdown (
Ezbam, EzRNAi vs none) — same direction. - C3: "germline knockdown of Pcl, Scm, Pc, Sce, or Jarid2 did not cause widespread gene-expression changes in nurse cells" → median log2FC of all classes ≈ 0 for those whole-ovary comparisons.
OUT OF SCOPE (not attempted)
- Wet-lab: immunostaining, antibody work, FACS germ-cell purification, fly genetics.
- ChIP-seq peak/domain calling (the upstream H3K27me3 4-state model that defines
ME_type) — we consume the shipped
polyATranscripts_MEtype.txtlabels as given rather than re-deriving domains from the ChIP-seq. - Hisat2 alignment + StringTie/HTSeq counting from FASTQ — the deposited count matrices are the pipeline input the repo expects; re-aligning 73 samples from SRA is a separate (out-of-scope) upstream step. We start from the deposited counts.
- Microscopy/quantification figures, models/cartoons.
Pipeline name per result
All Figure-7 numeric results: DESeq2 (R/Bioconductor) on HTSeq-count inputs,
exactly as in PcGRNAseq_Methods.R.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean 1:1 reproduction: the authors deposited per-sample HTSeq counts (GSE145282, complete, grade A) and shipped a fully-specified DESeq2 script, and re-running it regenerates every Figure-7 value within rounding (E(z) KD inactive 1.366x / PcG 1.397x ≈ 1.4-fold; stage-14 6.08x/5.97x ≈ sixfold; five control KDs flat). The only deviations are sub-precision decimal differences attributable to rounding, all on the technical/expected side. The one unverified element — the Active/Inactive/PcG labels — was consumed as a shipped input because the upstream ChIP-seq domain-calling code was not deposited; this is an availability caveat on the class definition, not a defect in the reproduced RNA-seq claims, and the central conclusion holds fully.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.