Toward a Consensus in the Repertoire of Hemocytes Identified in Drosophila.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Meta-analysis paper; named third-party code (pySCENIC) applied to the named/primary data (Cattenoz E-MTAB-8698 wasp-infested hemocytes) per rule P16. PARTIAL reproduction. Data-level claims reproduce: C1 exact (deposited WI annotation = 13PL+1CC+2LM), C2 within-tol (lamellocytes 8->1342 with wasp infestation, 168x). SCENIC regulon-recovery (Fig 3A) executed fully on «our HPC» with the paper's exact resources (dm6 motifs v8, 5kb upstream, seed 777): GRNBoost2 -> ctx -> AUCell, 118 regulons over 1342 lamellocytes. All 8 named lamellocyte-regulon TFs (kay, Jra, CrebB, foxo, REPTOR-BP, pnr, Maf-S, CHES-1-like) are recovered AS regulons -> no sign of fabricated TFs. Lamellocyte-SPECIFIC enrichment reproduces for 4/8 (Jra, REPTOR-BP, pnr, CHES-1-like); kay/CrebB/foxo/Maf-S form regulons not LM-enriched in our run -> partial, expected from GRNBoost2 stochasticity + pySCENIC 0.12.1 vs paper 0.9.19 + a cruder mean-AUC enrichment test than the paper's heatmap/RSS. Corrected an earlier provisional error: Maf-S is NOT absent from the v8 5kb DB (present as 'maf-S'; recovered). NOT attempted: marker-overlap counts (threshold-sensitive), tissue datasets (<100 hemocytes, authors' own caveat), wet-lab results. The pyscenic ctx CLI hangs at multiprocessing teardown in 0.12.1 (workers finish but never write output); replaced with a Python-API driver (reproduction/code/ctxauc.py) that writes outputs then os._exit(0). «infra» workdir: «path»
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-18 ⛓ 53642bdf5429
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether three independently generated scRNAseq datasets of Drosophila larval hemocytes, despite using different experimental and analytical approaches, converge on a common, robust repertoire of hemocyte subgroups.
- ★ Comparative analysis of three scRNAseq studies identifies eight common, robust hemocyte subgroups associated with distinct functions (proliferation, immune response, phagocytosis, secretion) finding
- ★ The three published scRNAseq studies used different experimental and analytical parameters yet report overlapping hemocyte subgroup diversity method
- ★ Some larval immune cells resemble embryonic hemocyte progenitors when compared to stage 6 embryo scRNAseq data finding
- ★ Larval immune cells associated with peripheral tissues (eye disc, brain) express tissue-specific properties finding
- Tattikota et al. identified twelve plasmatocyte subgroups, two crystal cell subgroups, and two lamellocyte subgroups from feeding/wandering 3rd instar larvae under steady state and challenged conditions resource
- Fu et al. identified four plasmatocyte subgroups, one crystal cell subgroup, one lamellocyte subgroup, and two minor populations (primocytes, thanacytes) from wandering 3rd instar larvae in steady state resource
- Cattenoz et al. identified thirteen plasmatocyte subgroups, one crystal cell subgroup, and two lamellocyte subgroups (the latter specific to wasp-challenged condition) from wandering 3rd instar larvae resource
- ★ Developmental trajectories among hemocyte subgroups are preserved across the three independent datasets, supporting the marker-based subgroup correspondence finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single cell RNA sequencing (10x Genomics) | wandering 3rd instar larval hemocytes (female, Drosophila) | wasp infestation (Leptopilina boulardi) vs steady state | transcriptional profile / subgroup identification | 10x Genomics |
| single cell RNA sequencing (10x Genomics) | wandering 3rd instar larval hemocytes (Drosophila) | none (steady state) | transcriptional profile / subgroup identification | 10x Genomics |
| single cell RNA sequencing (3 technologies) | feeding and wandering 3rd instar larval hemocytes (Drosophila, two genotypes) | clean wounding, wasp infestation, or steady state | transcriptional profile / subgroup identification | — |
| single cell RNA-seq comparative reanalysis (Seurat pipeline) | stage 6 Drosophila embryo | none | hemocyte progenitor marker expression, pseudo-transcriptome correlation | — |
| single cell RNA-seq comparative reanalysis (Seurat pipeline) | wild type larval eye disc (Drosophila) | none | hemocyte subgroup identification via markers (Srp, Hml, Pxn, NimC1, Crq, Sn) | — |
| single cell RNA-seq comparative reanalysis (Seurat pipeline, integrated) | 1st, 2nd, and 3rd instar larval brains (Drosophila) | none | hemocyte subgroup identification via markers (Srp, Hml, Pxn, NimC1, He, Nplp2) | — |
| SCENIC/pySCENIC regulon (gene regulatory network) analysis | wandering 3rd instar larval hemocytes, wasp-infested (Drosophila) | wasp infestation | differential regulon activity (AUC score, z-score) between lamellocyte subgroups LM1 and LM2 | pySCENIC v0.9.19 |
| immunolabeling and confocal microscopy | hemocytes, embryos, lymph gland, and filet preparations (Drosophila larvae/embryos) | none/genetic reporter lines (srp(hemo)>RFP, BAC-gcm-Flag) | marker protein localization (Srp, Flag, GFP, RFP, Pxn, Hemese) and phalloidin/DAPI staining | Leica Spinning Disk and Leica SP8 confocal microscopes |
- – Eight common hemocyte subgroups robustly identified across the three independent scRNAseq datasets
- – Plasmatocytes constitute the majority of hemocytes ~95% of hemocytes
- – Developmental trajectories of hemocyte subgroups are conserved across the three datasets, corroborating the marker-based subgroup matching
- – Comparison with stage 6 embryo data yields a single embryonic hemocyte subgroup enriched for progenitor markers (Gcm, Ham, Ttk, CrebA, Shep, RhoL, Fok, Knrl, Kni, Zfh1, CG33099, Srp, Btd, NetB)
- – Hemocyte subgroups unambiguously identified in larval eye disc and brain datasets using core hemocyte markers
- – Regulons with differential activity between lamellocyte subgroups LM1 and LM2 identified via SCENIC z-score >2 or <-2
- – PM12 subgroup in Tattikota et al. appears exclusively under wounding condition, not assessed by the other two studies
- other Log2(enrichment) > 0.25, adjusted p < 0.01 (threshold used to define subgroup markers in Cattenoz et al. and Tattikota et al. datasets)
- other z-score above 2 or below -2 (threshold for selecting regulons with differential activity in lamellocyte subgroups (Mann-Whitney U-test on SCENIC AUC scores))
- count ~95% (proportion of hemocytes that are plasmatocytes)
- count ~10 μm diameter (plasmatocyte cell size)
- count >60 μm diameter (lamellocyte cell size)
- count 12 plasmatocyte, 2 crystal cell, 2 lamellocyte subgroups (subgroups identified by Tattikota et al. (2020))
- count 4 plasmatocyte, 1 crystal cell, 1 lamellocyte, 2 minor (primocytes, thanacytes) subgroups (subgroups identified by Fu et al. (2020))
- count 13 plasmatocyte, 1 crystal cell, 2 lamellocyte subgroups (subgroups identified by Cattenoz et al. (2020))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This computational meta-analysis re-analyzed three published Drosophila larval hemocyte scRNAseq datasets to identify consensus immune cell subgroups. Cross-dataset marker overlap was assessed qualitatively and via enrichment metrics (Log2 enrichment >0.25, adjusted p < 0.01) generated by Seurat's FindMarkers. Transcriptome similarity across datasets and developmental stages was quantified with Pearson correlation on pseudo-transcriptomes. Regulon activity differences between lamellocyte clusters and all other clusters were assessed with Mann-Whitney U-tests, with differentially active regulons selected by z-score threshold (|z| > 2).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Mann-Whitney U-test | Comparing SCENIC AUC scores for each regulon between lamellocyte subgroups and all remaining clusters (Figure 3A heatmap) | — | not stated |
| Pearson correlation coefficient | Comparing pseudo-transcriptomes of hemocyte subgroups across datasets (larval vs. embryo, eye disc, brain; Supplementary Figures S3C, S3E, S3G) | — | not stated |
| Seurat FindMarkers (test type not named; Wilcoxon rank-sum is Seurat default) | Defining subgroup marker genes in Cattenoz et al. and Tattikota et al. datasets used for cross-dataset comparison (Figure 2A) | — | not stated |
| z-score thresholding (|z| > 2) | Selecting differentially active regulons in lamellocyte subgroups for heatmap display (Figure 3A) | — | na |
-
Pseudo-transcriptome similarity across datasets was measured with Pearson correlation↳ Could also: Spearman rank correlation could also be used for pseudo-transcriptome comparisons — Spearman correlation is less sensitive to highly expressed outlier genes and does not assume a linear relationship, properties that may be advantageous given the skewed, zero-inflated nature of aggregated scRNAseq expression values
-
Differentially active regulons were selected using a z-score threshold (|z| > 2) applied after Mann-Whitney U-tests across all regulons↳ Could also: Benjamini-Hochberg FDR correction applied to the Mann-Whitney p-values could also be used to select significant regulons — When many regulons are tested simultaneously, an FDR-based approach provides a principled control of the false discovery rate across the entire family of tests, complementing or replacing the z-score threshold
-
Cross-dataset integration relied on separately analyzing each dataset and then comparing pseudo-transcriptomes↳ Could also: Batch-aware integration methods such as Seurat CCA/RPCA, Harmony, or scVI could also be applied to jointly embed cells from all three datasets — Joint embedding would allow direct single-cell-level comparison across datasets and could reveal subgroup correspondences that are obscured by dataset-specific technical variation when using pseudo-transcriptome averaging
-
Marker genes were defined with fixed thresholds (Log2 enrichment >0.25, adjusted p < 0.01) derived from FindMarkers↳ Could also: ROC-based marker scoring (also available in Seurat) or an AUC classifier per cluster could also be used — An AUC-based approach directly quantifies how well a gene separates one cluster from others, providing an effect-size measure that is independent of the statistical significance threshold and easier to interpret as a classification metric
-
Cluster-vs-rest Mann-Whitney U comparisons were performed for each regulon independently without a stated multiplicity correction↳ Could also: A multi-group non-parametric test (e.g., Kruskal-Wallis followed by Dunn's post-hoc test with FDR correction) could also be applied when comparing regulon activity across all clusters simultaneously — A global test followed by corrected pairwise tests would control the family-wise error rate across clusters and regulons simultaneously, making the threshold for claiming differential activity explicit and reproducible
-
Subgroup correspondence across the three datasets was inferred by overlapping published marker lists and visual inspection of enrichment levels↳ Could also: Label transfer algorithms (e.g., Seurat's TransferData or SingleR) could also be used to map cluster identities from one dataset onto another at the single-cell level — Automated label transfer produces a quantitative confidence score for each cell's assignment to a reference cluster, reducing subjectivity in cross-dataset subgroup matching and allowing uncertainty to be reported
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33748138
Title: Toward a Consensus in the Repertoire of Hemocytes Identified in Drosophila. Cattenoz, Monticelli, Pavlidaki, Giangrande. Front Cell Dev Biol 2021. DOI 10.3389/fcell.2021.643712. Registry code: https://github.com/aertslab/pySCENIC (third-party tool — valid under rule P16). Registry data: GEO GSE134722.
What kind of paper this is
This is a meta-analysis / consensus reanalysis paper. It does not generate new sequencing data. It re-examines several previously published single-cell RNA-seq datasets of Drosophila larval hemocytes (blood cells) to argue for a consensus set of hemocyte subtypes. The computational pipeline is, per dataset: Seurat (NormalizeData → FindVariableFeatures vst 2000 → ScaleData → RunPCA → ElbowPlot → FindNeighbors/FindClusters → UMAP/tSNE), marker detection, pseudo-transcriptome correlation (Pearson), and pySCENIC (v0.9.19) regulon inference on one dataset.
Datasets the paper relies on
| dataset | accession | role | access |
|---|---|---|---|
| Cattenoz et al. 2020 (own prior atlas; wasp-infested + non-infested larval hemocytes) | E-MTAB-8698 (ArrayExpress) | primary; SCENIC substrate; cluster structure | open |
| Tattikota et al. 2020 (larval blood scRNA-seq) | GSE146596 | primary (marker overlap) | open |
| Fu et al. 2020 (larval hemocytes; thanacytes/primocytes) | (GEO; not stated in this paper) | primary | open |
| Karaiskos et al. 2017 (stage-6 embryo) | DVEX web app | secondary tissue | web |
| Ariss et al. 2018 (larval eye disc) | E-MTAB-7195 | secondary tissue | open |
| Brunet Avalos et al. 2019 (1st-instar larval brain) | GSE134722 (registry accession) | secondary tissue (<100 brain hemocytes) | open |
| Cocanougher et al. 2019 (2nd/3rd-instar brain) | GSE135810 | secondary tissue | open |
Note: GSE134722 (the registry accession) is a minor secondary dataset — the 1st-instar brain atlas (from the Sprecher lab, PMID 31746739), used by this paper only to extract <100 brain-associated hemocytes. The paper's quantitative pipeline results come from the primary hemocyte datasets, above all Cattenoz E-MTAB-8698, which is also the named code's (pySCENIC) input.
In scope (pipeline-derived, attempted)
- SCENIC regulon recovery in lamellocytes (Fig 3A) — PRIMARY TARGET. Run pySCENIC
(GRNBoost2 → cisTarget → AUCell) on the wasp-infested Cattenoz matrix (E-MTAB-8698,
RRCZ23) using the paper's resources (Drosophila dm6, motif collection v8, 5 kb
upstream), with the deposited cell→cluster annotation (
WI_cell_cluster_ID.txt) defining lamellocyte cells. Reported claim: 7 novel lamellocyte regulons (CrebB, foxo, REPTOR-BP, pnr/Pannier, Maf-S, a zinc-finger, CHES-1-like) + 2 known JNK regulons (kay/Kayak, Jra/Jun-related antigen). Pipeline = pySCENIC (named code). - Cattenoz cluster structure (data-level). Reported: 13 plasmatocyte + 1 crystal cell + 2 lamellocyte subgroups. Checkable directly from the deposited annotation.
- Dataset profiling of E-MTAB-8698 (WI+NI) and GSE134722 (named accession).
Out of scope / not attempted (and why)
- Seurat re-clustering cluster counts of the primary datasets: the paper re-uses the source papers' published annotations; no crisp recomputed N to match for Cattenoz.
- Marker-overlap counts (53 crystal-cell, 265 lamellocyte common markers; Fig 2): require Tattikota + Cattenoz markers and are highly threshold-sensitive (known to mismatch); deprioritised, may attempt if time.
- Tissue datasets (embryo via DVEX web app, eye disc, brains): <100 hemocytes each; authors themselves flag this as the analysis's main caveat. Qualitative only.
- All wet-lab / in-situ / antibody results: out of scope (non-pipeline).
- pySCENIC version: paper used 0.9.19; reproduced with 0.12.1 (same 3-step CLI; GRNBoost2 is stochastic — not bit-reproducible by design). Documented.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Data-level claims reproduce exactly from public, fully-deposited Cattenoz E-MTAB-8698 annotations: the consensus cluster repertoire (13 PL + 1 CC + 2 LM, C1) and the wasp-induced lamellocyte expansion (WI=1342 vs NI=8, ~170x, C2). The central SCENIC regulon claims (Fig3A: kay/Jra + 7 novel TFs, C3/C4) are still pending the «our HPC» job and one TF (Maf-S) is provably unrecoverable from the v8 5kb cisTarget DB — a database constraint on our side / the field, not an authors' defect. Remaining caveats are technical (pySCENIC v0.12.1 vs v0.9.19, GRNBoost2 stochasticity). Overall: a clean, partial reproduction with no discrepancy on anything checked, but incomplete, so q5/q7/q8 sit at yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.