Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Single-cell transcriptomics identifies Keap1-Nrf2 regulated collective invasion in a Drosophila tumor model.

Elife · 2022
L1 72/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
72/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 41% of all assessed papers rank 688 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL, honest 1:1 where the data allows. Chatterjee 2022 eLife e80956 (Drosophila Keap1-Nrf2 collective invasion). Reproduced the authors' OWN repo code on the deposited GSE175435 processed data (no FASTQ realignment needed). BULK: C2 PCA PC1=77% EXACT; C1 DEGs 319 vs 477 (same edgeR glmTreat pipeline, version drift, paper pins no edgeR version). scRNA (Seurat v4): C3a w1118 follicle cells 17835 vs 17875 (0.2%, within-tol); C3b Lgl-KD 13840 vs 14537 (4.8%); C4 21 vs 20 clusters (within-tol); C5 cluster-7 size NOT comparable (Seurat cluster IDs not preserved across versions). Deterministic QC cell counts all factual (25144/19986/19313/16060). Limits: cell_cycle_genes.txt missing from deposit (substituted tinyatlas Drosophila CC list); non-epithelial removal uses author-run-specific hard-coded cluster IDs. NOT attempted: wet-lab assays, STAR/cellranger/velocyto preprocessing (deposit ships processed outputs), scvelo trajectory, GO/marker biology. No values fabricated; all from repo code @39a75f4 on GSE175435.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 72
    assessed: 2026-06-22 ⛓ 203687bb781b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

How does pathological apical-basal polarity loss (via Lgl knockdown in Drosophila follicle cells) link to cellular plasticity and invasive behavior, and what gene expression programs and pathways mediate this transition?

Core claims
  • Follicle cell-specific Lgl knockdown (Lgl-KD) causes loss of apical-basal polarity and invasive multilayering of the follicular epithelium without long-distance metastasis finding
  • Keap1-Nrf2 oxidative stress signaling is genetically required for multilayer formation in Lgl-KD follicle cells mechanism
  • Ectopic Keap1 expression increases the volume of delaminated follicle cells and enhances invasive behavior with significant cytoskeletal (F-actin) remodeling finding
  • Integrated single-cell transcriptomes of Lgl-KD and wildtype (w1118) follicle cells identify clusters of cells unique to the Lgl-KD multilayering phenotype finding
  • A comprehensive single-cell transcriptomic atlas of the follicle cell tumor model is generated as a resource at single-cell resolution resource
  • Whole-tissue bulk RNA-seq shows the 96hr-Lgl-KD transcriptome is significantly divergent from control and shorter-induction samples finding
  • Shg (DE-Cad) and Arm (α-catenin) enrichment progressively decreases along the apical-basal axis of Lgl-KD multilayers while F-actin becomes mildly elevated at the apical invasive front finding
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-seq (scRNA-seq) Drosophila ovary follicle cells (tjTS>lglRNAi, 72hr induction) Lgl RNAi knockdown transcriptomic clustering, cell-type/phenotype-specific gene expression
bulk whole-tissue RNA-seq Drosophila ovarian tissue Lgl RNAi knockdown (24hr and 96hr induction) vs experimental control differential gene expression, PCA, GO term enrichment
immunofluorescence/confocal microscopy Drosophila ovary follicle cells, egg chamber cross-sections Lgl-KD Hnt, Cut, Shg (DE-Cad), Arm, F-actin (Phalloidin), pH3 staining intensity and localization along apical-basal axis confocal microscope
mitotic clonal analysis (MARCM) Drosophila follicle cells lgl RNAi MARCM clones (GFP+) and homozygous lgl4 mutant clones (GFP-) apical invasive clonal movement
phenotype quantification (imaging-based scoring) Drosophila ovarioles (tjTS>lglRNAi, 72hr) Lgl-KD percentage of ovarioles/egg chambers with multilayering, fusion, or degeneration phenotypes
Key results
  • Multilayering was the most prevalent phenotype at midoogenesis after 72hr Lgl-KD 83.25% (n=280)
  • Degenerated egg chambers observed at late oogenesis stages beyond stage 9/10 51.86% (n=280)
  • Fused egg chambers observed at early oogenesis 6.05% (n=280)
  • PCA separates 96h-Lgl-KD sample from other samples along PC1 77% variance explained by PC1
  • Genes for actin binding, locomotion, and cell periphery GO terms elevated in 96h-Lgl-KD
  • Apoptotic clearance genes croquemort (crq) and draper (drpr) upregulated in 96h-Lgl-KD samples crq: 0.66 log2FC; drpr: 0.642 log2FC
  • 14,537 Lgl-KD follicle cells integrated with 17,875 w1118 follicle cells into 20 clusters, with clusters 7, 8, 13, 16, and 17 unique to Lgl-KD dataset
  • Shg and Arm enrichment declines along apical-basal axis of multilayers while F-actin is mildly elevated at the apical-most invading front
Key statistics
  • pvalue p=2.904 × 10^-19 (GO term cell periphery (GO:0071944) enrichment in 96h-Lgl-KD upregulated genes)
  • pvalue p=3.748 × 10^-11 (GO term locomotion (GO:0040011) enrichment in 96h-Lgl-KD upregulated genes)
  • pvalue p=2.843 × 10^-3 (GO term actin binding (GO:0003779) enrichment in 96h-Lgl-KD upregulated genes)
  • fold_change 0.66 log2FC, p=0.00015 (croquemort (crq) upregulation in 96h-Lgl-KD vs other samples)
  • fold_change 0.642 log2FC, p=0.00027 (draper (drpr) upregulation in 96h-Lgl-KD vs other samples)
  • count 83.25% (proportion of egg chambers showing multilayering phenotype at midoogenesis after 72h-Lgl-KD)
  • count 51.86% (proportion of egg chambers showing degeneration at late oogenesis after Lgl-KD)
  • count 6.05% (proportion of egg chambers showing fusion at early oogenesis after Lgl-KD)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper combines whole-tissue and single-cell RNA sequencing of Drosophila ovarian tissue to compare Lgl-knockdown follicle cells against wild-type/experimental controls at several induction time points. Sample-level relationships were visualized with PCA, differential gene expression between conditions was reported with log2 fold-changes and p-values, Gene Ontology enrichment was reported with adjusted p-values, and phenotype frequencies across independent replicate trials were summarized as percentages and box-and-whisker plots.

Replicationmixed Sample sizePhenotype quantification: n=280 egg chambers pooled from five independent replicate trials across 1250 intact ovarioles from 165 flies (Figure 1D); RNA-seq: two replicates per condition, four samples total (Figure 1E); scRNA-seq: 14,537 Lgl-KD and 17,875 w1118 follicle cells integrated (Figure 2A) GroupstjTS experimental control vs. tjTS>lglRNAi (Lgl-knockdown) follicle cells, at different induction durations (24h, 72h, 96h) Pairingunclear Randomization/blindingnot stated DispersionIQR Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionyes
Statistical tests used
Test Applied to n Assumptions
Differential gene expression analysis (specific statistical model not named in text) Whole-tissue RNA-seq comparison of 96h-Lgl-KD samples vs. other samples (Figure 1—figure supplement 2, e.g. crq and drpr log2FC/p-values) Two replicates per condition (four RNA-seq samples total, as shown in Figure 1E PCA) not stated
Gene Ontology (GO) term enrichment analysis 477 differentially expressed genes from the 96h-Lgl-KD vs. others comparison (Figure 1—figure supplement 2B) 477 differentially expressed genes not stated
Approaches that could also have been used
  • Phenotype percentages from five independent replicate trials (color-coded) were pooled and shown as a single box-and-whisker plot (n=280 total egg chambers).
    Could also: A mixed-effects/hierarchical model treating replicate trial (or fly) as a random effect — This would explicitly account for possible non-independence of egg chambers sampled from the same trial or fly, alongside the pooled descriptive summary already shown.
  • Differential expression between 96h-Lgl-KD and other RNA-seq samples is reported with log2 fold-change and p-values without naming the underlying statistical model.
    Could also: A count-based generalized linear model framework such as DESeq2 or edgeR with a Wald or likelihood-ratio test — These tools model RNA-seq count dispersion explicitly and are a standard way to generate the fold-change/p-value pairs described, and stating the tool and correction method (e.g., Benjamini-Hochberg FDR) would make the multiple-testing control explicit alongside the reported values.
  • GO term enrichment was performed on a fixed list of 477 differentially expressed genes, using an unspecified adjusted p-value method.
    Could also: Gene set enrichment analysis (GSEA) using the full ranked gene list rather than a thresholded DE gene set — GSEA can capture coordinated but sub-threshold expression changes across a pathway, complementing the threshold-based enrichment approach already used.
  • PCA was used to visualize separation between whole-tissue RNA-seq samples (Figure 1E).
    Could also: Hierarchical clustering or a sample-to-sample distance/correlation heatmap — This offers a complementary view of sample similarity structure alongside PCA, which can be useful when checking whether replicate samples cluster together beyond the first two principal components.
  • Effect sizes for individual genes are reported as log2 fold-change with a paired exact p-value (e.g., crq: 0.66 log2FC, p=0.00015).
    Could also: Reporting an accompanying confidence interval for the fold-change estimate — A CI would convey the precision of the fold-change estimate in addition to the point estimate and significance test, which can be informative when replicate numbers are small.
  • Large-scale phenotype categories (multilayering, fused, degenerated) were quantified as percentages of total egg chambers across pooled trials.
    Could also: A chi-square or logistic regression model comparing phenotype proportions between conditions/time points — This would allow a formal statistical comparison of proportions between groups (e.g., 24h vs 96h induction) in addition to the descriptive percentages and plot already provided.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36321803 (Chatterjee 2022, eLife e80956)

Authors' own repo (chatterjee89/eLife2022-11-e80956, CC0) = 4 scripts: preprocessing_code.sh (STAR + featureCounts + cellranger + velocyto — needs FASTQ, OUT OF SCOPE: deposit ships processed outputs so re-alignment is unnecessary), bulkSeq_code.R (edgeR DE + DESeq2 PCA — IN SCOPE), seurat_code.R (Seurat QC/cluster/integration — IN SCOPE), scvelo_code.py (RNA velocity of cluster 7 — trajectory plot, OUT OF SCOPE as a numeric claim).

IN SCOPE (pipeline-derived, attempted)

  • C1 DEGs 96h-LglRNAi vs others -> edgeR glmTreat (bulkSeq_code.R)
  • C2 PCA PC1 variance (Fig 1E) -> DESeq2 VST plotPCA (bulkSeq_code.R)
  • C3a/C3b follicle cell counts -> Seurat QC + non-epithelial cluster removal (seurat_code.R)
  • C4 clusters after integration -> Seurat two-round integration (seurat_code.R)
  • C5 cluster 7 size -> Seurat clustering (seurat_code.R)
  • QC1-4 deterministic cell counts -> Read10X / loom + explicit QC thresholds

OUT OF SCOPE (not attempted)

  • All wet-lab: fly genetics, immunostaining, invasion assays, lifespan, GFP imaging.
  • STAR/cellranger/velocyto pre-processing (deposit provides the processed outputs).
  • scvelo RNA-velocity vector field / trajectory (qualitative plot, not a number).
  • GO enrichment and marker-gene biology interpretation.

Pipelines named: edgeR, DESeq2 (bulk); Seurat v4 (scRNA). Method-card hits: DESeq2/edgeR (Diff-expr), Seurat (scRNA-seq).

Figures / tables: Fig 1EFig 2AFig 4A
C1
Reported
477 DEGs (96h-LglRNAi vs others, edgeR glmTreat, Suppl Data 1)
Reproduced
319 DEGs (FDR<0.05; 248 up / 71 down)
partial
C2
Reported
PCA PC1 = 77% variance (Fig 1E)
Reproduced
77% (PC2=13%)
exact
C3a
Reported
17875 w1118 follicle cells (Fig 2A)
Reproduced
17835
within tolerance
C3b
Reported
14537 Lgl-KD follicle cells (Fig 2A)
Reproduced
13840
partial
C4
Reported
20 clusters after integration (Fig 2A)
Reproduced
21
within tolerance
C5
Reported
cluster 7 = 1144 cells (Fig 4A)
Reproduced
1779 (NOT like-for-like: Seurat cluster numbering not preserved across versions)
did not match
QC1
Reported
control raw barcodes (deposited 10x)
Reproduced
25144
within tolerance
QC2
Reported
control cells post-QC (deterministic)
Reproduced
19986
within tolerance
QC3
Reported
LglIR raw cells (deposited loom)
Reproduced
19313
within tolerance
QC4
Reported
LglIR cells post-QC (deterministic)
Reproduced
16060
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 72/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This is a solid partial reproduction of the authors' own repo run on the open GSE175435 deposit: the most-specified number, PC1=77% (Fig 1E), matched exactly, and w1118 cells (17835 vs 17875, 0.2%) and cluster count (21 vs 20) landed within tolerance. The notable deviations — C1 DEGs 319 vs 477 and C3b 13840 vs 14537 (4.8%) — are explainable by edgeR version drift (no version pinned by the paper) and a missing input file (cell_cycle_genes.txt absent, substituted with tinyatlas). C5 (1779 vs 1144) is not a like-for-like comparison because Seurat cluster numbering is not preserved across versions. No values were fabricated; deviations sit on the technical/our-method and data-availability side, and the central collective-invasion conclusion is not contradicted.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

262.6 k
tokens (I/O) · 14.4 M incl. cache
68 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.