Sequencing mRNA from cryo-sliced Drosophila embryos to determine genome-wide spatial patterns of gene expression.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- 🟡Reported values were only indirectly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the WT Replicate 1 (6-slice) portion of the SliceSeq pipeline (Combs & Eisen 2013, PMID 23951250) end-to-end on «our HPC»: downloaded FASTQs, built the combined dmel+D.persimilis bowtie2/TopHat reference, aligned all 6 slices, ran the paper's own shipped AssignReads2.py species-classification script, and quantified dmel-assigned reads with Cufflinks. Table 1's per-slice Total Reads reproduce EXACTLY for all 6 slices (cross-checked independently via ENA). Table 1's per-slice species-classification percentages (%dmel, %ambiguous) do NOT reproduce in absolute terms (dmel% ~2-6x too high, ambiguous% ~125-170x too low, consistently across all 6 slices), but the relative spatial RANK ORDER of dmel% across slices matches Table 1 almost perfectly (5/6 slices in identical order). This discrepancy is well-explained by a genuine gap between the paper's stated methods (a discordant-read-end override rule) and the actually shipped AssignReads2.py code, which does not implement that rule; reimplementing it gets closer but does not exactly reproduce Table 1 either -- graded 'partial' with full methodological documentation, not asserted as a pipeline failure. Figures 1-3 (spatial expression pattern claims via BDTNP atlas matching / clustering) depend on a separate, un-downloaded 87-sample 'Fine Sliced' dataset and external BDTNP reference data; these were not attempted in this pass (explicit scope decision, not a drop) beyond a coarse 2-of-6-slice FPKM sanity check on classic AP-patterning genes showing directionally plausible (but statistically thin) anterior-posterior gradients. The GSE43506 dataset (122-123 samples) was profiled: only the 6-sample WT-Rep1 subset actually processed; Bcd6x (17) and Fine Sliced (87) subsets confirmed to exist in GEO/ENA metadata but not downloaded or further verified.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusBecause individual Drosophila embryos yield more RNA than is needed for accurate expression measurement, the authors ask whether RNA-seq of physically cryosectioned single embryos can recover reliable, genome-wide spatial gene expression information along the anterior-posterior axis.
- ★ Cryosectioning single blastoderm-stage D. melanogaster embryos along the A–P axis and sequencing mRNA from each slice yields reliable genome-wide spatial expression patterns. method
- ★ Adding total RNA from a distantly related, fully sequenced Drosophila species (or yeast) as carrier makes library preparation from ~15 ng slice RNA robust, and carrier reads are computationally separable rather than wasted. method
- ★ Reconstructed spatial patterns from slice RNA-seq closely match patterns previously determined by in situ hybridization and by a cellular-resolution expression atlas. finding
- ★ Slice-seq identifies numerous genes with clear spatial patterns not described by ongoing systematic in situ projects, including many with expression restricted to the embryonic poles. finding
- ★ 25 µm slices resolve individual gap-gene domains and at least three of the seven eve pair-rule stripes, including inter-stripe repression. finding
- Sequencing-based spatial methods are simpler, cheaper, more portable across species/genotypes, and apparently more sensitive than in situ-based systematic surveys. finding
- ★ A genome-wide timecourse of spatial expression spanning stage 2 through gastrulation (seven morphology-staged timepoints) is provided as a community resource (GEO GSE43506, eisenlab.org/sliceseq, github.com/eisenlab/SliceSeq). resource
- A single embryo contains enough RNA in principle to sequence over 700 samples at 20 million reads each, making 20 µm three-dimensional dicing theoretically feasible. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| mRNA-seq (paired-end, 50 bp) of 60 µm cryosections | single Canton-S D. melanogaster embryos, cellular blastoderm (cell cycle 14, stage 5); 3 embryos, 6 slices each | none (wild type); cryosectioning plus spike-in of carrier total RNA from single D. persimilis, D. willistoni, or D. mojavensis embryos | transcript abundance (FPKM) per slice for all annotated mRNAs | Illumina HiSeq 2000; modified Illumina TruSeq RNA kit; Microm HM 550 cryotome; TRIzol extraction; Kapa Library Quantification kit on Roche LC480 |
| mRNA-seq of 25 µm cryosections (developmental timecourse) | single D. melanogaster embryos from 7 morphology-defined timepoints (stage 2, stage 4, five timepoints within stage 5); 10–15 usable slices per embryo | none (wild type); cryosectioning plus carrier total RNA from Saccharomyces cerevisiae and Torulaspora delbruckii | per-slice expression of patterning genes (hb, kni, gt, eve) and genome-wide transcript levels normalized to peak | Illumina HiSeq 2000, 50 bp paired-end (Vincent Coates Genome Sequencing Laboratory) |
| Light-microscopy imaging and morphological staging | methanol-fixed, dechorionated Canton-S D. melanogaster embryos in halocarbon oil | methanol cracking fixation; bromophenol blue staining for slice visualization | depth of membrane invagination and other morphological features for stage selection | Nikon 80i microscope with DS-5M camera |
| Read mapping and transcript quantification (computational) | reads from sliced D. melanogaster embryos plus carrier species genomes (FlyBase FB2012_05) | none | species assignment of reads (≥4 discriminating positions per read), uniquely mapped percentage, FPKM transcript abundances | TopHat v2.0.6 (max 6 mismatches), Cufflinks; HiSeq Control Software v1.8, Real Time Analysis v2.8 |
| Differential expression testing between slices | 60 µm slice RNA-seq data, 3 replicate embryos | none | genes with statistically significant expression differences between slices | Cuffdiff |
| k-means clustering of spatial expression profiles | 60 µm slice RNA-seq data filtered to FPKM >10 in at least one slice and inter-embryo Pearson r >0.5 | none | clusters of normalized spatial expression profiles (k = 20, centroid linkage) | — |
| Bayesian positional matching against a cellular-resolution 3D expression atlas | in silico 60 µm sections of BDTNP 'virtual embryos' (95 genes, single-nucleus resolution) compared to averaged 60 µm slice data | none | posterior probability that each slice corresponds to a given start position in the atlas | — |
| Comparison to published RNA in situ hybridization annotations/images | BDGP in situ image and annotation database for D. melanogaster embryos | none | qualitative concordance of spatial patterns; annotation status ('present in a subset of the embryo') | — |
- – Between 1.7% and 31.4% of read ends per slice+carrier sample mapped unambiguously to D. melanogaster 1.7–31.4%
- – Reconstructed spatial patterns for 33 genes qualitatively agree with published BDGP in situ patterns across all three sliced 60 µm embryos 33 genes
- – Bayesian matching of slice data to atlas sections gave a clear posterior peak per slice, with peaks ordered sequentially along the embryo and spaced ~60 µm, matching slice thickness ~60 µm spacing
- – Cuffdiff identified 85 genes with statistically significant between-slice expression differences; 21 lacked BDGP imaging data, 33 were annotated as 'present in a subset of the embryo', and 31 showed unannotated patterns or no clear staining 85 genes (21/33/31)
- – 194 genes annotated as patterned by BDGP were not called statistically significant in the slice data, mostly due to dorsal-ventral patterns, faint staining, later staging, or falling just above the significance cutoff 194 genes
- – k-means clustering of 745 filtered genes yielded 13 of 20 clusters with non-uniform patterns, including anterior and posterior pole localization and five gap-gene-like bands; only 349 of the 745 had BDGP images 745 genes; 13/20 clusters; 349 with images
- – 25 µm slices resolve low-expression gaps between the multiple hb, kni, and gt domains and at least three eve pair-rule stripes with inter-stripe repression ≥3 of 7 eve stripes
- ▼ Yeast carrier RNA (S. cerevisiae, T. delbruckii) reduced ambiguously mapping reads to under 0.003%, versus ~0.9–4.7% ambiguous with Drosophila carriers <0.003% vs 0.9–4.7%
- count approximately 15 ng total RNA per slice; ~100 ng carrier total RNA added (input material per 60 µm slice before and after carrier addition)
- count approximately 40 million 50 bp paired-end reads per slice+carrier sample (sequencing depth for 60 µm slice libraries (3 embryos × 6 slices))
- fold_change 1.7% to 31.4% uniquely mapped D. melanogaster read ends (Table 1, range across 60 µm slice samples)
- count 85 genes significantly differentially expressed between slices (Cuffdiff on 60 µm data from 3 embryos; described as a very conservative estimate)
- correlation average Pearson correlation between embryos >0.5 (replicate-agreement filter applied before k-means clustering)
- count 745 genes passing filters; k = 20 clusters; 349 with BDGP images (k-means clustering of spatial patterns)
- other fewer than 0.003% of reads ambiguously mapping (yeast carrier RNA in the 25 µm timecourse)
- count ~350 µm embryo cut into six 60 µm pieces; 10–15 usable 25 µm slices per embryo; 7 timepoints (sectioning geometry and timecourse design)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used RNA-seq (TopHat alignment, Cufflinks/Cuffdiff quantification) on cryosectioned Drosophila embryo slices to infer spatial gene-expression patterns, comparing biological replicate embryos (three for the 60 µm dataset, one per timepoint for the 25 µm timeseries). Reproducibility was assessed with Pearson correlation between slices/embryos, differential expression across slices was identified with Cuffdiff, and pattern discovery used k-means clustering (k=20) after correlation-based filtering. Validation against known patterns relied on qualitative comparison to published in-situ images and a custom Bayesian posterior-probability procedure matching slice profiles to a reference spatial atlas, with results reported mainly as percentages, correlation coefficients, and qualitative concordance rather than through p-values, effect sizes, or confidence intervals.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Cuffdiff differential expression test | Identifying genes with statistically significant expression differences between slices (85 genes identified) | — | not stated |
| Pearson correlation | Assessing reproducibility between slices within an embryo (Figure S1) and between equivalent slices of different embryos (Figure S2); also used as a filtering criterion (>0.5) before clustering | three CaS embryos (60 µm dataset) | not stated |
| Custom Bayesian posterior-probability procedure | Comparing slice expression levels to sections of the BDTNP virtual-embryo atlas to estimate slice position (Figure 1B) | 98 genes, averaged across three embryos | not stated |
| k-means clustering (k=20, centroid linkage) | Grouping genes into spatial expression pattern classes (Figure 2, Dataset S3) | 745 genes passing expression-level and cross-embryo correlation filters | not stated |
-
Differential expression between slices was identified using Cuffdiff.↳ Could also: Negative-binomial count-based methods such as DESeq2 or edgeR — These tools model gene-wise dispersion from replicate counts and are widely used for RNA-seq differential expression; they could offer an alternative estimate of significant genes alongside Cuffdiff's.
-
The paper reports a specific number of statistically significant genes (85) without stating an explicit p-value or FDR threshold in the presented text.↳ Could also: Explicitly reporting the significance threshold and multiple-testing correction method (e.g., Benjamini-Hochberg FDR with a stated q-value cutoff) — Stating the exact threshold and correction approach would make it more transparent how the family-wise or false-discovery rate was controlled across the many genes tested.
-
Pearson correlation was used to assess reproducibility between slices and embryos and as a filtering criterion before clustering.↳ Could also: Spearman's rank correlation — A rank-based correlation is less sensitive to outliers and does not assume a linear relationship, which can be useful for expression data spanning a wide dynamic range.
-
Genes were grouped into spatial pattern classes using k-means clustering with k=20 chosen empirically.↳ Could also: Hierarchical clustering with a data-driven cut height, or model-based clustering (e.g., Gaussian mixture models) selected via BIC/AIC — These approaches can provide a more formal, data-driven criterion for choosing the number of clusters rather than empirical selection.
-
Reproducibility and pattern agreement with published in-situ data were assessed largely through qualitative visual comparison and correlation.↳ Could also: A quantitative concordance statistic (e.g., Cohen's kappa for categorical pattern calls, or a formal regression-based comparison) — A quantitative concordance metric could complement the visual and correlation-based comparisons with a single summary statistic of agreement.
-
Slice position was inferred using a custom Bayesian posterior-probability procedure comparing gene levels to the BDTNP atlas.↳ Could also: A standard spatial regression or maximum-likelihood alignment method — Such methods are commonly used for matching expression profiles to reference spatial coordinates and could serve as an alternative or complementary cross-check to the custom Bayesian approach.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
What deviates: Table 1's Total Reads column reproduces exactly (all six WT-Rep1 slices, integer-identical, cross-checked against ENA), but the adjacent species-classification columns do not: running the authors' own shipped AssignReads2.py on the authors' own FASTQs gives dmel% of 12.64-18.17 against Table 1's 1.7-8.7, and ambiguous% of ~0.012-0.017 against the reported ~2.0-2.3 — a systematic ~3x inflation and ~150x deflation, reproducible in every slice. Whose side: authors' — the deposited code classifies read1/read2 independently, contradicting the paper's own Methods footnote about discordant ends, and neither the literal code nor a reimplementation of the stated rule (slice1: 6.89% / 11.51%) lands on the published values, so Table 1 is not derivable from the shared data + shared code. Severity: large on that table, but the biological signal survives (slice rank order of dmel% matches 5/6, bcd falls and cad rises anteroposteriorly), so this reads as a code-deposit/documentation gap rather than a fabrication signal. Caveat on our side: only WT-Rep1 (6 of 122-123 samples) was processed and the 87-sample fine-sliced dataset behind Figures 1-3 was never downloaded, so the paper's headline claim is supported directionally but not actually re-tested.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.