Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Toward a Consensus in the Repertoire of Hemocytes Identified in Drosophila.

Front Cell Dev Biol · 2021
L1 74/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
74/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 43% of all assessed papers rank 644 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Meta-analysis paper; named third-party code (pySCENIC) applied to the named/primary data (Cattenoz E-MTAB-8698 wasp-infested hemocytes) per rule P16. PARTIAL reproduction. Data-level claims reproduce: C1 exact (deposited WI annotation = 13PL+1CC+2LM), C2 within-tol (lamellocytes 8->1342 with wasp infestation, 168x). SCENIC regulon-recovery (Fig 3A) executed fully on «our HPC» with the paper's exact resources (dm6 motifs v8, 5kb upstream, seed 777): GRNBoost2 -> ctx -> AUCell, 118 regulons over 1342 lamellocytes. All 8 named lamellocyte-regulon TFs (kay, Jra, CrebB, foxo, REPTOR-BP, pnr, Maf-S, CHES-1-like) are recovered AS regulons -> no sign of fabricated TFs. Lamellocyte-SPECIFIC enrichment reproduces for 4/8 (Jra, REPTOR-BP, pnr, CHES-1-like); kay/CrebB/foxo/Maf-S form regulons not LM-enriched in our run -> partial, expected from GRNBoost2 stochasticity + pySCENIC 0.12.1 vs paper 0.9.19 + a cruder mean-AUC enrichment test than the paper's heatmap/RSS. Corrected an earlier provisional error: Maf-S is NOT absent from the v8 5kb DB (present as 'maf-S'; recovered). NOT attempted: marker-overlap counts (threshold-sensitive), tissue datasets (<100 hemocytes, authors' own caveat), wet-lab results. The pyscenic ctx CLI hangs at multiprocessing teardown in 0.12.1 (workers finish but never write output); replaced with a Python-API driver (reproduction/code/ctxauc.py) that writes outputs then os._exit(0). «infra» workdir: «path»

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 71
    assessed: 2026-06-18 ⛓ 53642bdf5429
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether three independently generated scRNAseq datasets of Drosophila larval hemocytes, despite using different experimental and analytical approaches, converge on a common, robust repertoire of hemocyte subgroups.

Core claims
  • Comparative analysis of three scRNAseq studies identifies eight common, robust hemocyte subgroups associated with distinct functions (proliferation, immune response, phagocytosis, secretion) finding
  • The three published scRNAseq studies used different experimental and analytical parameters yet report overlapping hemocyte subgroup diversity method
  • Some larval immune cells resemble embryonic hemocyte progenitors when compared to stage 6 embryo scRNAseq data finding
  • Larval immune cells associated with peripheral tissues (eye disc, brain) express tissue-specific properties finding
  • Tattikota et al. identified twelve plasmatocyte subgroups, two crystal cell subgroups, and two lamellocyte subgroups from feeding/wandering 3rd instar larvae under steady state and challenged conditions resource
  • Fu et al. identified four plasmatocyte subgroups, one crystal cell subgroup, one lamellocyte subgroup, and two minor populations (primocytes, thanacytes) from wandering 3rd instar larvae in steady state resource
  • Cattenoz et al. identified thirteen plasmatocyte subgroups, one crystal cell subgroup, and two lamellocyte subgroups (the latter specific to wasp-challenged condition) from wandering 3rd instar larvae resource
  • Developmental trajectories among hemocyte subgroups are preserved across the three independent datasets, supporting the marker-based subgroup correspondence finding
Experimental setups
Assay System Perturbation Readout Platform
single cell RNA sequencing (10x Genomics) wandering 3rd instar larval hemocytes (female, Drosophila) wasp infestation (Leptopilina boulardi) vs steady state transcriptional profile / subgroup identification 10x Genomics
single cell RNA sequencing (10x Genomics) wandering 3rd instar larval hemocytes (Drosophila) none (steady state) transcriptional profile / subgroup identification 10x Genomics
single cell RNA sequencing (3 technologies) feeding and wandering 3rd instar larval hemocytes (Drosophila, two genotypes) clean wounding, wasp infestation, or steady state transcriptional profile / subgroup identification
single cell RNA-seq comparative reanalysis (Seurat pipeline) stage 6 Drosophila embryo none hemocyte progenitor marker expression, pseudo-transcriptome correlation
single cell RNA-seq comparative reanalysis (Seurat pipeline) wild type larval eye disc (Drosophila) none hemocyte subgroup identification via markers (Srp, Hml, Pxn, NimC1, Crq, Sn)
single cell RNA-seq comparative reanalysis (Seurat pipeline, integrated) 1st, 2nd, and 3rd instar larval brains (Drosophila) none hemocyte subgroup identification via markers (Srp, Hml, Pxn, NimC1, He, Nplp2)
SCENIC/pySCENIC regulon (gene regulatory network) analysis wandering 3rd instar larval hemocytes, wasp-infested (Drosophila) wasp infestation differential regulon activity (AUC score, z-score) between lamellocyte subgroups LM1 and LM2 pySCENIC v0.9.19
immunolabeling and confocal microscopy hemocytes, embryos, lymph gland, and filet preparations (Drosophila larvae/embryos) none/genetic reporter lines (srp(hemo)>RFP, BAC-gcm-Flag) marker protein localization (Srp, Flag, GFP, RFP, Pxn, Hemese) and phalloidin/DAPI staining Leica Spinning Disk and Leica SP8 confocal microscopes
Key results
  • Eight common hemocyte subgroups robustly identified across the three independent scRNAseq datasets
  • Plasmatocytes constitute the majority of hemocytes ~95% of hemocytes
  • Developmental trajectories of hemocyte subgroups are conserved across the three datasets, corroborating the marker-based subgroup matching
  • Comparison with stage 6 embryo data yields a single embryonic hemocyte subgroup enriched for progenitor markers (Gcm, Ham, Ttk, CrebA, Shep, RhoL, Fok, Knrl, Kni, Zfh1, CG33099, Srp, Btd, NetB)
  • Hemocyte subgroups unambiguously identified in larval eye disc and brain datasets using core hemocyte markers
  • Regulons with differential activity between lamellocyte subgroups LM1 and LM2 identified via SCENIC z-score >2 or <-2
  • PM12 subgroup in Tattikota et al. appears exclusively under wounding condition, not assessed by the other two studies
Key statistics
  • other Log2(enrichment) > 0.25, adjusted p < 0.01 (threshold used to define subgroup markers in Cattenoz et al. and Tattikota et al. datasets)
  • other z-score above 2 or below -2 (threshold for selecting regulons with differential activity in lamellocyte subgroups (Mann-Whitney U-test on SCENIC AUC scores))
  • count ~95% (proportion of hemocytes that are plasmatocytes)
  • count ~10 μm diameter (plasmatocyte cell size)
  • count >60 μm diameter (lamellocyte cell size)
  • count 12 plasmatocyte, 2 crystal cell, 2 lamellocyte subgroups (subgroups identified by Tattikota et al. (2020))
  • count 4 plasmatocyte, 1 crystal cell, 1 lamellocyte, 2 minor (primocytes, thanacytes) subgroups (subgroups identified by Fu et al. (2020))
  • count 13 plasmatocyte, 1 crystal cell, 2 lamellocyte subgroups (subgroups identified by Cattenoz et al. (2020))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This computational meta-analysis re-analyzed three published Drosophila larval hemocyte scRNAseq datasets to identify consensus immune cell subgroups. Cross-dataset marker overlap was assessed qualitatively and via enrichment metrics (Log2 enrichment >0.25, adjusted p < 0.01) generated by Seurat's FindMarkers. Transcriptome similarity across datasets and developmental stages was quantified with Pearson correlation on pseudo-transcriptomes. Regulon activity differences between lamellocyte clusters and all other clusters were assessed with Mann-Whitney U-tests, with differentially active regulons selected by z-score threshold (|z| > 2).

Replicationunclear Sample size10 wandering 3rd instar larvae used for immunolabeling; cell counts for scRNAseq analyses not restated here (derived from three previously published datasets) Groupshemocyte transcriptional subgroups vs. each other; larval hemocytes vs. embryonic, eye disc, and brain-associated hemocytes Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionadjusted p-value < 0.01 (correction method not named in this paper; inherited from Cattenoz et al. 2020 and Tattikota et al. 2020 marker tables)
Statistical tests used
Test Applied to n Assumptions
Mann-Whitney U-test Comparing SCENIC AUC scores for each regulon between lamellocyte subgroups and all remaining clusters (Figure 3A heatmap) not stated
Pearson correlation coefficient Comparing pseudo-transcriptomes of hemocyte subgroups across datasets (larval vs. embryo, eye disc, brain; Supplementary Figures S3C, S3E, S3G) not stated
Seurat FindMarkers (test type not named; Wilcoxon rank-sum is Seurat default) Defining subgroup marker genes in Cattenoz et al. and Tattikota et al. datasets used for cross-dataset comparison (Figure 2A) not stated
z-score thresholding (|z| > 2) Selecting differentially active regulons in lamellocyte subgroups for heatmap display (Figure 3A) na
Approaches that could also have been used
  • Pseudo-transcriptome similarity across datasets was measured with Pearson correlation
    Could also: Spearman rank correlation could also be used for pseudo-transcriptome comparisons — Spearman correlation is less sensitive to highly expressed outlier genes and does not assume a linear relationship, properties that may be advantageous given the skewed, zero-inflated nature of aggregated scRNAseq expression values
  • Differentially active regulons were selected using a z-score threshold (|z| > 2) applied after Mann-Whitney U-tests across all regulons
    Could also: Benjamini-Hochberg FDR correction applied to the Mann-Whitney p-values could also be used to select significant regulons — When many regulons are tested simultaneously, an FDR-based approach provides a principled control of the false discovery rate across the entire family of tests, complementing or replacing the z-score threshold
  • Cross-dataset integration relied on separately analyzing each dataset and then comparing pseudo-transcriptomes
    Could also: Batch-aware integration methods such as Seurat CCA/RPCA, Harmony, or scVI could also be applied to jointly embed cells from all three datasets — Joint embedding would allow direct single-cell-level comparison across datasets and could reveal subgroup correspondences that are obscured by dataset-specific technical variation when using pseudo-transcriptome averaging
  • Marker genes were defined with fixed thresholds (Log2 enrichment >0.25, adjusted p < 0.01) derived from FindMarkers
    Could also: ROC-based marker scoring (also available in Seurat) or an AUC classifier per cluster could also be used — An AUC-based approach directly quantifies how well a gene separates one cluster from others, providing an effect-size measure that is independent of the statistical significance threshold and easier to interpret as a classification metric
  • Cluster-vs-rest Mann-Whitney U comparisons were performed for each regulon independently without a stated multiplicity correction
    Could also: A multi-group non-parametric test (e.g., Kruskal-Wallis followed by Dunn's post-hoc test with FDR correction) could also be applied when comparing regulon activity across all clusters simultaneously — A global test followed by corrected pairwise tests would control the family-wise error rate across clusters and regulons simultaneously, making the threshold for claiming differential activity explicit and reproducible
  • Subgroup correspondence across the three datasets was inferred by overlapping published marker lists and visual inspection of enrichment levels
    Could also: Label transfer algorithms (e.g., Seurat's TransferData or SingleR) could also be used to map cluster identities from one dataset onto another at the single-cell level — Automated label transfer produces a quantitative confidence score for each cell's assignment to a reference cluster, reducing subjectivity in cross-dataset subgroup matching and allowing uncertainty to be reported
Software: R/Seurat · Python/pySCENIC 0.9.19 · R/ggplot2 · R/pheatmap · Fiji · Adobe Illustrator CS6 CS6

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33748138

Title: Toward a Consensus in the Repertoire of Hemocytes Identified in Drosophila. Cattenoz, Monticelli, Pavlidaki, Giangrande. Front Cell Dev Biol 2021. DOI 10.3389/fcell.2021.643712. Registry code: https://github.com/aertslab/pySCENIC (third-party tool — valid under rule P16). Registry data: GEO GSE134722.

What kind of paper this is

This is a meta-analysis / consensus reanalysis paper. It does not generate new sequencing data. It re-examines several previously published single-cell RNA-seq datasets of Drosophila larval hemocytes (blood cells) to argue for a consensus set of hemocyte subtypes. The computational pipeline is, per dataset: Seurat (NormalizeData → FindVariableFeatures vst 2000 → ScaleData → RunPCA → ElbowPlot → FindNeighbors/FindClusters → UMAP/tSNE), marker detection, pseudo-transcriptome correlation (Pearson), and pySCENIC (v0.9.19) regulon inference on one dataset.

Datasets the paper relies on

dataset accession role access
Cattenoz et al. 2020 (own prior atlas; wasp-infested + non-infested larval hemocytes) E-MTAB-8698 (ArrayExpress) primary; SCENIC substrate; cluster structure open
Tattikota et al. 2020 (larval blood scRNA-seq) GSE146596 primary (marker overlap) open
Fu et al. 2020 (larval hemocytes; thanacytes/primocytes) (GEO; not stated in this paper) primary open
Karaiskos et al. 2017 (stage-6 embryo) DVEX web app secondary tissue web
Ariss et al. 2018 (larval eye disc) E-MTAB-7195 secondary tissue open
Brunet Avalos et al. 2019 (1st-instar larval brain) GSE134722 (registry accession) secondary tissue (<100 brain hemocytes) open
Cocanougher et al. 2019 (2nd/3rd-instar brain) GSE135810 secondary tissue open

Note: GSE134722 (the registry accession) is a minor secondary dataset — the 1st-instar brain atlas (from the Sprecher lab, PMID 31746739), used by this paper only to extract <100 brain-associated hemocytes. The paper's quantitative pipeline results come from the primary hemocyte datasets, above all Cattenoz E-MTAB-8698, which is also the named code's (pySCENIC) input.

In scope (pipeline-derived, attempted)

  1. SCENIC regulon recovery in lamellocytes (Fig 3A) — PRIMARY TARGET. Run pySCENIC (GRNBoost2 → cisTarget → AUCell) on the wasp-infested Cattenoz matrix (E-MTAB-8698, RRCZ23) using the paper's resources (Drosophila dm6, motif collection v8, 5 kb upstream), with the deposited cell→cluster annotation (WI_cell_cluster_ID.txt) defining lamellocyte cells. Reported claim: 7 novel lamellocyte regulons (CrebB, foxo, REPTOR-BP, pnr/Pannier, Maf-S, a zinc-finger, CHES-1-like) + 2 known JNK regulons (kay/Kayak, Jra/Jun-related antigen). Pipeline = pySCENIC (named code).
  2. Cattenoz cluster structure (data-level). Reported: 13 plasmatocyte + 1 crystal cell + 2 lamellocyte subgroups. Checkable directly from the deposited annotation.
  3. Dataset profiling of E-MTAB-8698 (WI+NI) and GSE134722 (named accession).

Out of scope / not attempted (and why)

  • Seurat re-clustering cluster counts of the primary datasets: the paper re-uses the source papers' published annotations; no crisp recomputed N to match for Cattenoz.
  • Marker-overlap counts (53 crystal-cell, 265 lamellocyte common markers; Fig 2): require Tattikota + Cattenoz markers and are highly threshold-sensitive (known to mismatch); deprioritised, may attempt if time.
  • Tissue datasets (embryo via DVEX web app, eye disc, brains): <100 hemocytes each; authors themselves flag this as the analysis's main caveat. Qualitative only.
  • All wet-lab / in-situ / antibody results: out of scope (non-pipeline).
  • pySCENIC version: paper used 0.9.19; reproduced with 0.12.1 (same 3-step CLI; GRNBoost2 is stochastic — not bit-reproducible by design). Documented.
Figures / tables: Fig1Fig 3A
C1
Reported
Cattenoz hemocyte cluster structure: 13 plasmatocyte + 1 crystal-cell + 2 lamellocyte subgroups
Reproduced
WI & NI deposited annotation each = 16 clusters (13 PL + 1 CC + 2 LM)
exact
C2
Reported
wasp infestation induces lamellocytes
Reproduced
WI lamellocytes(LM-1+LM-2)=1342 vs NI=8 (~168x)
within tolerance
C3
Reported
2 known JNK regulons (kay, Jra) active in lamellocytes (Fig3A)
Reproduced
both regulons recovered by pySCENIC; Jra LM-enriched (AUCell fold 2.19), kay recovered but NOT LM-enriched (0.88)
partial
C4
Reported
7 novel lamellocyte regulons (CrebB, foxo, REPTOR-BP, pnr, Maf-S, zinc-finger, CHES-1-like) (Fig3A)
Reproduced
6/6 named TFs recovered as regulons; LM-enriched: REPTOR-BP(1.69), pnr(1.74), CHES-1-like(2.62); NOT enriched: CrebB(0.68), foxo(0.40), Maf-S(0.53); 7th unnamed zinc-finger not scored
partial
C5
Reported
pySCENIC (v0.9.19, motifs v8, 5kb upstream) runs on the paper's data
Reproduced
ran end-to-end with pyscenic 0.12.1, dm6 motifs v8, 5kb upstream, seed 777: 763913 GRN edges -> 2971 modules -> 118 regulons over 1342 lamellocytes
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 74/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Data-level claims reproduce exactly from public, fully-deposited Cattenoz E-MTAB-8698 annotations: the consensus cluster repertoire (13 PL + 1 CC + 2 LM, C1) and the wasp-induced lamellocyte expansion (WI=1342 vs NI=8, ~170x, C2). The central SCENIC regulon claims (Fig3A: kay/Jra + 7 novel TFs, C3/C4) are still pending the «our HPC» job and one TF (Maf-S) is provably unrecoverable from the v8 5kb cisTarget DB — a database constraint on our side / the field, not an authors' defect. Remaining caveats are technical (pySCENIC v0.12.1 vs v0.9.19, GRNBoost2 stochasticity). Overall: a clean, partial reproduction with no discrepancy on anything checked, but incomplete, so q5/q7/q8 sit at yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

286.7 k
tokens (I/O) · 18 M incl. cache
163 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.