Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A single-cell survey of Drosophila blood.

Elife · 2020
L1 94/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
94/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 87% of all assessed papers rank 133 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

salvaged by watchdog from agreement.json (agent omitted ROOM_RESULT.json)

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-07-29

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

How diverse is the Drosophila larval hemocyte (blood cell) repertoire at the molecular level, and can single-cell RNA sequencing across steady-state and inflammatory conditions resolve mature hemocyte types into distinct subtypes, intermediate states, and activated states to reveal new immune functions?

Core claims
  • scRNA-seq of Drosophila larval hemocytes across unwounded, wounded, and wasp-infested conditions resolves 17 clusters spanning plasmatocytes, crystal cells, lamellocytes, and a non-hemocyte population. finding
  • Plasmatocytes are resolved into multiple states defined by cell cycle (PM2), immune activation/antimicrobial peptide expression (PM3-7), metabolism, and several minor subpopulations (PM8-12). finding
  • Crystal cells and lamellocytes each split into two clusters representing a mature state and an immature/intermediate state (CC1 immature vs CC2 mature; LM1 intermediate vs LM2 mature). finding
  • Rare subsets of crystal cells express the FGF ligand branchless (bnl) and rare lamellocytes express the FGF receptor breathless (btl), revealing FGF signaling components within hemocytes. finding
  • FGF signaling (bnl/btl) mediates inter-hemocyte crosstalk required for effective immune responses against parasitoid wasp eggs in vivo. mechanism
  • A novel Metchnikowin-like antimicrobial peptide gene (CG43236/Mtk-like, Mtkl) is identified enriched in plasmatocyte cluster PM7. finding
  • A searchable Drosophila blood scRNA-seq web portal provides a resource of hemocyte gene expression profiles across conditions. resource
  • Harmony batch correction integrated into Seurat was used to merge inDrops, 10X Chromium, and Drop-seq datasets and remove batch effects. method
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA sequencing (inDrops) Drosophila melanogaster larval hemocytes wounding / wasp infestation / unwounded control per-cell gene expression (UMIs, transcriptomes), cell clustering inDrops
single-cell RNA sequencing (10X Chromium) Drosophila melanogaster larval hemocytes wounding / wasp infestation / unwounded control per-cell gene expression / cell clustering 10X Chromium
single-cell RNA sequencing (Drop-seq) Drosophila melanogaster larval hemocytes wounding / wasp infestation / unwounded control per-cell gene expression / cell clustering Drop-seq
quantitative real-time PCR (qRT-PCR) Drosophila larval sessile and circulating hemocytes (unwounded) none (steady state) expression of plasmatocyte cluster marker genes in sessile vs circulating compartments
bulk RNA-seq comparison (pseudobulk vs published bulk) Drosophila larval whole hemocytes and lz+ crystal cells none Spearman correlation of pseudobulk scRNA-seq vs published bulk RNA-seq
in vivo functional immune assay Drosophila larvae parasitoid wasp (Leptopilina boulardi) infestation; FGF (bnl/btl) manipulation immune response against wasp eggs
Key results
  • 19,458 hemocytes profiled across conditions yielding 17 clusters after Harmony integration 19,458 cells; 17 clusters
  • PM6 cecropin-enriched cluster proportion matches previously observed fraction of cecropin-expressing hemocytes upon infection ~5.5% (vs prior ~5-10%)
  • PM AMP clusters (PM6, PM7) and lamellocyte LM2 emerge mainly upon wounding or wasp infestation
  • Crystal cell clusters CC1 and CC2 are underrepresented in wasp inf. 24 hr, consistent with crystal cell loss after oviposition
  • Pseudobulk scRNA-seq correlates strongly with published bulk RNA-seq for whole hemocytes and crystal cells, indicating high data quality Spearman r ~0.79
  • A non-hemocyte cluster lacks pan-hemocyte markers He and srp ~0.2% of cells
  • Novel Mtk-like AMP (Mtkl/CG43236) identified among top expressed genes of cluster PM7
Key statistics
  • count 19,458 cells (total hemocytes profiled across all conditions/platforms)
  • count median 1010 genes per cell (median genes detected per cell across conditions)
  • count median 2883 UMIs per cell (median unique molecular identifiers per cell)
  • correlation Spearman ~0.79 (pseudobulk scRNA-seq vs published bulk RNA-seq for hemocytes and crystal cells)
  • count 17 clusters (total clusters identified after Harmony batch correction)
  • other ~5.5% (proportion of cecropin-enriched PM6 cluster)
  • other ~90-95% (plasmatocyte fraction of total hemocytes)
  • other ~0.2% (fraction of profiled cells in non-hemocyte cluster)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is a descriptive single-cell RNA-sequencing survey of Drosophila hemocytes across unwounded, wounded, and wasp-infested larvae, profiling 19,458 cells with 3–4 replicates per condition on three platforms (inDrops, 10X Chromium, Drop-seq). Data sets were merged, batch-corrected with Harmony within the Seurat pipeline, and unsupervised clustering identified 17 clusters annotated by known and de novo marker genes (ranked by average log fold-change). Validation relied on Spearman correlation of pseudobulk scRNA-seq against published bulk RNA-seq, plus qRT-PCR of selected markers; results were reported largely as cluster identities, marker dot/heat maps, and cell-fraction changes rather than formal hypothesis tests.

Replicationbiological Sample size3–4 replicates per condition; 19,458 total cells; n=4 unwounded, n=4 wounded, n=3 wasp inf. 24 hr, n=3 wasp inf. 48 hr; no formal power/sample-size calculation described Groupsunwounded vs wounded vs wasp-infested larval hemocytes Pairingunpaired Randomization/blindingnot stated Dispersionnone Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Spearman rank correlation pseudobulk scRNA-seq vs published bulk RNA-seq for all hemocytes and for crystal cells (Figure 1—figure supplement 2A,B) not stated
Cluster marker / differential expression ranking by average log fold-change (avg_logFC) top genes per cluster, dot plot (Figure 1D) and marker tables (Supplementary file 2) not stated
qRT-PCR expression comparison (sessile vs circulating hemocytes) plasmatocyte marker genes in unwounded larvae (Figure 1—figure supplement 1G) na
Approaches that could also have been used
  • Pseudobulk-to-bulk concordance was quantified with a Spearman correlation coefficient (~0.79).
    Could also: A Pearson correlation on log-transformed counts, or a concordance correlation coefficient, could also be reported, optionally with a confidence interval. — Pearson on log-scale captures linear agreement and a confidence interval would also convey the precision of the correlation estimate; reporting both gives a fuller picture of agreement.
  • Cluster marker genes were prioritized by average log fold-change.
    Could also: Adjusted p-values from a differential-expression test (e.g. Wilcoxon rank-sum with Benjamini-Hochberg FDR, or model-based methods such as MAST or DESeq2 on pseudobulk) could also accompany the fold-change ranking. — Pairing effect size with an FDR-controlled significance value would also communicate statistical confidence in each marker alongside its magnitude.
  • Cell-fraction changes across conditions were summarized as proportions/percentages per cluster (Figure 2D).
    Could also: A compositional-analysis framework (e.g. scCODA, or a Dirichlet-multinomial / proportion test across replicates) could also be applied. — Such approaches would also incorporate replicate-to-replicate variability and the compositional nature of the data when describing fraction shifts.
  • Batch effects across condition, replicate, and technology were addressed with Harmony.
    Could also: Alternative integration methods such as Seurat CCA/RPCA anchors, scVI, or fastMNN could also be used. — Comparing integration methods would also let one describe how robust the cluster structure is to the choice of batch-correction algorithm.
  • Marker expression in sessile vs circulating compartments was assessed by qRT-PCR (Figure 1—figure supplement 1G).
    Could also: Replicate-level qRT-PCR values could also be summarized with SD or a 95% CI and compared with a t-test or Mann-Whitney U, depending on n and distribution. — Reporting spread and an accompanying test would also convey the variability and the strength of the sessile-vs-circulating difference.
Software: Seurat (R) · Harmony (batch correction, integrated in Seurat) · inDrops platform · 10X Chromium platform · Drop-seq platform · Jalview (protein alignment)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
212
Impact: very high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

BDSC:30140 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BDSC:32210 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BDSC:33042 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BDSC:43544 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BDSC:46517 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BDSC:5 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BDSC:5137 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BDSC:52271 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BDSC:60013 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BDSC:6314 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BDSC:66782 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BL# 32210 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BL# 33042 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BL# 34572 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BL# 6314 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BL# 66782 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
BL#60013 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE146596 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig1Fig1B
C1
Reported
19458
Reproduced
19458
exact
C2
Reported
17
Reproduced
17
exact
C3
Reported
17 total
Reproduced
17 distinct clusters (count exact; PM/CC/LM identity = manual annotation, not re-derived)
exact
C4
Reported
1010
Reproduced
1011
within tolerance
C5
Reported
2883
Reproduced
2884
within tolerance
C6
Reported
derivable
Reproduced
pending (see reclust_result.json)
m.public.grade.uncheckable
C7
Reported
~17
Reproduced
pending (see reclust_result.json)
m.public.grade.uncheckable

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 94/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

All examined headline values reproduce from the authors' own deposited GSE146596 metadata: cell count (19,458) and cluster count (17) are exact, and median genes/UMIs differ only by 1 (rounding/even-N median, ~0.1%). No value was found non-derivable from the shipped data — no fabrication signal. The main scope caveat is on our side, not the authors': this is an internal-consistency audit of the deposited output vs the manuscript, not a from-FASTQ re-derivation, and cell-type identities (C3) plus independent re-clustering (C7) remain manual/pending.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

111.4 k
tokens (I/O) · 8.1 M incl. cache
40 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.