Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Single-cell transcriptome maps of myeloid blood cell lineages in Drosophila.

Nat Commun · 2020
L1 86/100 3/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
86/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 70% of all assessed papers rank 334 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

salvaged by watchdog from agreement.json (agent omitted ROOM_RESULT.json)

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator headless) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-07-29

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

How Drosophila lymph gland hemocytes develop and are regulated at single-cell resolution is unclear; the authors test whether single-cell RNA sequencing can comprehensively resolve the heterogeneity, developmental trajectories, and immune responses of developing myeloid-like hemocytes and distinguish embryonic- versus lymph gland-derived lineages.

Core claims
  • Single-cell RNA-seq of developing Drosophila lymph glands resolves heterogeneity of hemocytes and identifies major and sub cell types. resource
  • Previously undescribed hemocyte types exist, including adipohemocytes, stem-like prohemocytes, and intermediate prohemocytes. finding
  • A GST-rich hemocyte cluster and an adipohemocyte cluster (lipid-metabolism/phagocytosis gene-enriched) are present in wild-type lymph glands. finding
  • Developmental trajectories of hemocytes can be reconstructed, with PH1 as the start point and emergence of the lamellocyte lineage upon wasp infestation. finding
  • Crystal cells and lamellocytes each split into premature (CC1/LM1) and mature (CC2/LM2) states. finding
  • SCENIC analysis delineates cell-type-specific transcriptional regulators (e.g., DREAM/Dpp factors in prohemocytes; ecdysone pathway factors in plasmatocytes). mechanism
  • New genetic reporter tools/markers were generated and validated (e.g., Ance-MiMiC, dome-LexA; PSC markers Ilp6, tau, mthl7, chrb; crystal cell markers Men, Numb). resource
  • Embryonically derived and larval lymph gland hemocytes share similarities and differences. finding
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-seq (Drop-seq) Drosophila larval lymph gland hemocytes at 72, 96, 120 h AEL none (normal development) single-cell transcriptomes; cell-type clustering and proportions Drop-seq; Seurat3 integration; Scrublet QC
single-cell RNA-seq (Drop-seq) Drosophila lymph gland hemocytes after parasitic wasp infestation wasp (Leptopilina) infestation / active cellular immunity emergence of lamellocyte lineage; transcriptomes Drop-seq
bulk RNA-seq Drosophila wild-type lymph glands none validation of gene detection and cluster signature gene expression
gene regulatory network inference (SCENIC) lymph gland scRNA-seq dataset none transcription factor regulons per cell type SCENIC
trajectory/pseudotime analysis (Monocle3) lymph gland hemocyte scRNA-seq (PSC excluded) none developmental trajectory/pseudotime ordering Monocle3
fluorescence in situ hybridization / immunostaining (in vivo validation) Drosophila lymph gland; reporter lines (gal4, MiMiC, LexA, GTRACE) none marker gene/protein localization (e.g., CG18547, CG3397, Sirup, Lsd-2, vir-1, Antp)
lipid droplet staining Drosophila lymph gland hemocytes (cortical zone) none neutral lipid droplet detection in adipohemocytes BODIPY; Nile Red; Phalloidin
DAPI cell counting / clustering (Louvain) single lymph gland lobe none DAPI-positive cell counts and cluster identity
Key results
  • 22,645 high-quality cells retained across three timepoints (72h:2321; 96h:9400; 120h:10,924) 22,645 cells
  • Eight major cell types identified including prohemocytes, plasmatocytes, crystal cells, lamellocytes, PSC, GST-rich, adipohemocyte, dorsal vessel PH 36.2%, PM 57.6%, CC 1.3%, LM 1.5%, PSC 0.9%, DV 0.1%, GST-rich 1.0%, adipohemocyte 1.4%
  • At 72h AEL prohemocytes and plasmatocytes are nearly equal; by 120h AEL plasmatocytes dominate and only ~30% retain prohemocyte signature 49.8% PH and 46.1% PM at 72h; ~30% PH at 120h
  • Crystal cells and GST-rich cells first appear at 96h AEL; lamellocytes and adipohemocytes appear at 120h AEL
  • 17 transcriptionally distinct hemocyte subclusters identified (6 PH, 4 PM, 2 LM, 2 CC subclusters) 17 subtypes
  • Crystal cells split into CC1 (early, low lz with MZ/CZ markers) and CC2 (mature, high PPO1/PPO2)
  • Adipohemocytes contain neutral lipid droplets and express Sirup/Lsd-2 within Hml+ cortical zone
  • Cell coverage of one lymph gland lobe was 5.5X, 6.8X, and 2.4X at 72, 96, 120h AEL respectively 5.5X / 6.8X / 2.4X
Key statistics
  • count 22,645 cells total (high-quality cells retained after QC)
  • count median 6361 transcripts (UMIs) and 1477 genes per cell (per-cell complexity)
  • count median 397, 1392, 4557 (DAPI+ cell counts per lobe at 72, 96, 120h AEL (n=30 each))
  • count plasmatocytes ~95% of hemocytes (plasmatocyte proportion of total hemocytes (background))
  • count crystal cells ~5% (crystal cell proportion of blood population (background))
  • count 49.8% and 46.1% (prohemocyte and plasmatocyte proportions at 72h AEL)
  • count 14 libraries (5 for 72h, 5 for 96h, 4 for 120h AEL) (independent sequencing libraries)
  • count ~30% (fraction retaining prohemocyte signature at 120h AEL)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study used single-cell RNA sequencing (Drop-seq) of Drosophila lymph gland hemocytes across three developmental timepoints (72, 96, and 120 h AEL), integrating 14 libraries with Seurat3 batch correction and Louvain-based unsupervised clustering. Subclusters were characterised by Wilcoxon Rank-Sum tests for differentially expressed marker genes. Developmental trajectories were reconstructed with Monocle3, and transcriptional regulatory networks were inferred with SCENIC. Results were primarily reported as cell counts, proportions, and median values, with in vivo validation of key marker genes by fluorescence imaging.

Replicationbiological Sample size14 independent sequencing libraries (5 × 72 h AEL, 5 × 96 h AEL, 4 × 120 h AEL); n = 30 lymph gland lobes per timepoint for cell-count quantification (Fig. 1b); total 22,645 cells retained after QC GroupsSeven major hemocyte types and 17 subclusters across three developmental timepoints (72, 96, 120 h AEL); additionally embryonic- vs. lymph gland-derived hemocytes and wasp-infested vs. uninfested conditions mentioned Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Wilcoxon Rank-Sum test Identification of significant marker genes across 17 hemocyte subclusters (Fig. 2b dot plot) 22,645 cells across subclusters (cell counts per subcluster listed in Supplementary Table 4) not stated
Louvain community-detection algorithm Unsupervised clustering of all cells into major cell types and subsequent subclusters (Fig. 1c, Fig. 2a) 22,645 cells not stated
Scrublet doublet prediction Quality-control pipeline to remove predicted doublets prior to downstream analysis not stated
SCENIC regulon activity scoring Transcription factor regulatory network inference per cell type (Supplementary Fig. 1k) 22,645 cells not stated
Monocle3 pseudotime trajectory reconstruction Developmental ordering of hemocyte subclusters (Supplementary Fig. 3a onward) 22,645 cells minus PSC not stated
Seurat3 integration / batch correction (canonical correlation + mutual nearest neighbours) Integration of 14 Drop-seq libraries across timepoints prior to clustering 14 libraries; 22,645 cells retained post-QC not stated
Approaches that could also have been used
  • The Louvain algorithm was used for graph-based clustering of single cells
    Could also: The Leiden algorithm (Traag et al. 2019) is an alternative graph-based clustering method that avoids the disconnected-community artefact inherent to Louvain and is now the default in many scRNA-seq workflows — Leiden guarantees well-connected communities and is generally considered a methodological refinement of Louvain; reporting which algorithm and resolution parameter was used also aids reproducibility
  • t-SNE was used for two-dimensional visualisation of cell type relationships (Figs. 1c, 2a)
    Could also: UMAP (McInnes et al. 2018) is a widely adopted alternative for single-cell visualisation — UMAP tends to better preserve global structure (inter-cluster distances) while remaining competitive on local structure, and has become the de facto standard in scRNA-seq publications; showing both would allow readers to assess robustness of cluster separation
  • Wilcoxon Rank-Sum tests were applied across many genes and all 17 subclusters without a stated multiple-testing correction
    Could also: Applying Benjamini–Hochberg FDR correction across all gene–subcluster comparisons is a standard companion to Wilcoxon tests in differential-expression analyses (and is an option within Seurat's FindMarkers function) — With thousands of genes tested across 17 groups simultaneously, FDR correction explicitly controls the expected proportion of false discoveries, making the threshold for 'significant' markers more interpretable to readers
  • Monocle3 was used to reconstruct developmental pseudotime trajectories
    Could also: RNA velocity (scVelo; Bergen et al. 2020) or PAGA (Wolf et al. 2019) are complementary trajectory methods; RNA velocity uses spliced/unspliced RNA ratios to infer directionality independently of pseudotime ordering — RNA velocity provides an orthogonal, data-driven estimate of differentiation direction that can corroborate or refine Monocle3 trajectories, particularly for identifying branch points
  • Cell count distributions (Fig. 1b) were summarised with median values only
    Could also: Reporting an IQR or range alongside the median, or showing individual data points as an overlay, would also convey the spread of counts across the n = 30 lymph gland lobes per timepoint — With n = 30 per group, variability metrics help readers judge biological consistency and are useful for power estimation in follow-up experiments
  • Batch effects across 14 libraries were corrected using Seurat3's integration (CCA + MNN anchors)
    Could also: Harmony (Korsunsky et al. 2019) or scVI (Lopez et al. 2018) are alternative batch-correction approaches that operate in latent space and can be applied to the same data — Benchmarking studies suggest that no single integration method outperforms all others across dataset types; cross-method concordance (e.g., consistent cluster assignments after Harmony integration) would strengthen confidence that observed cell types are not integration artefacts
Software: Seurat3 3 (exact patch version not stated) · Scrublet · SCENIC · Monocle3 · Drop-seq (platform/pipeline)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig 1cFig 2aFig 1a
C1
Reported
15,540
Reproduced
15,540 (all 14 matrices identical)
exact
C2
Reported
14 (5/5/4)
Reproduced
14 (5/5/4)
exact
C3
Reported
authors' per-cell metrics
Reproduced
exact for 19,439/19,439 cells (pearson 1.0, max|diff|=0)
exact
C4a
Reported
6,361
Reproduced
6,306
within tolerance
C4b
Reported
1,477
Reproduced
1,473
within tolerance
C5
Reported
19 clusters / 8 types
Reproduced
m.public.grade.uncheckable
C6
Reported
PH 36.2/PM 57.6/CC 1.3/LM 1.5/PSC 0.9/GST 1.0/Adipo 1.4/DV 0.1
Reproduced
PH 39.1/PM 54.1/CC 1.44/LM 1.75/PSC 0.97/GST 1.20/Adipo 0.92/DV 0.16
within tolerance
C7
Reported
17
Reproduced
20 (17 + DV/RG/Neurons)
within tolerance
HEADLINE
Reported
22,645 (2321/9400/10924)
Reproduced
19,439 (2227/9383/7829) in shipped revised labels
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator headless) · v1.0 L1 86/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This is a strong, honest reproduction: deterministic outputs (15,540-gene union, 14 libraries, per-cell QC) match 1:1 and the C3 anti-fabrication check passes decisively (deposited matrices regenerate the authors' per-cell metrics exactly, pearson 1.0, max|diff|=0), so there is no fabrication concern. The substantive deviation is the headline cell count (22,645 vs 19,439, with 120h dropping 10,924->7,829) plus minor Fig 1c proportion shifts (within 1-3.5 pts, rank preserved) and subcluster count (17->20). These all trace to the deposit being a later revised re-analysis rather than the published version — an authors'-side versioning/data-availability gap (the original 22,645 set is not recoverable from shared data), not a methodological or computational error on our side. The central conclusion — the 8-cell-type myeloid blood lineage atlas — holds robustly, so overall quality is solid with cleanly explainable deviations.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

178.5 k
tokens (I/O) · 13.4 M incl. cache
40 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.