Pathway-targeting gene matrix for Drosophila gene set enrichment analysis.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (1:1, third-party-tool P16). The authors' GeneMatrix repo ships two GMT pathway matrices (Drosophila KEGG + Reactome) plus a ready-to-run 13615x10 microarray validation matrix derived from GSE45344 (ELAV neurons vs REPO glia). Ran GSEA 4.2.3 on «our HPC» (paper used 4.1.0, no longer downloadable; algorithm unchanged). EXACT: KEGG=131 pathways, total genes=13615, size-filter pass 615 Reactome / 98 KEGG, gene-association split ELAV 5616 (41.2%) / REPO 7999 (58.8%). WITHIN-TOL: significant-set counts Reactome ELAV 216/154 (vs 214/148), REPO 68/49 (vs 68/52); KEGG ELAV 26/18 (vs 27/20), REPO 32/20 (vs 31/19) -- all within +-6 (REPO p<=0.05 exact), deltas explained by permutation stochasticity + GSEA version. Qualitative claims confirmed (mRNA splicing tops neurons; fatty-acid degradation + glutathione metabolism top glia). KEY METHOD NOTE: the paper says 'GSEA default parameters' but the reported counts are only reproducible with gene_set permutation, not GSEA's default phenotype permutation (which is degenerate at 5v5 and gives ~4x fewer sig sets, e.g. 53 vs 214). gene_set permutation is GSEA's own documented recommendation for n<7 per class -- derivable from shipped data, not a fabrication concern, but an under-specified Methods detail. NOT attempted (out of scope): manual web-curation of the matrices from KEGG/Reactome (no runnable pipeline; the GMT files are the curated result, which IS what we tested) and the wet-lab FACS+microarray generation of GSE45344 (DeSalvo et al., external).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 95assessed: 2026-06-21 ⛓ eb992c668a03
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether newly generated Reactome- and KEGG-based gene matrix (gmt) files for Drosophila melanogaster can be validated by GSEA to correctly distinguish and characterize known biological differences between neurons and glia.
- Gene matrix files for GSEA are largely unavailable for Drosophila, limiting pathway-level enrichment analysis in this model organism finding
- ★ Generated a Drosophila-specific Reactome gene matrix file covering 1450 annotated pathways resource
- ★ Generated a Drosophila-specific KEGG gene matrix file covering 131 annotated pathways with genetic information resource
- ★ The Reactome gmt file successfully separates neuron (ELAV) and glia (REPO) expression profiles via GSEA finding
- ★ The KEGG gmt file successfully separates neuron (ELAV) and glia (REPO) expression profiles via GSEA finding
- ★ Neuron-enriched pathways include mRNA splicing/spliceosome, transport of mature transcript, PIP synthesis, RAC1 GTPase cycle, WNT signaling, mTOR signaling, and endocytosis, consistent with known neuronal biology finding
- ★ Glia-enriched pathways include peptide hormone metabolism, mitochondrial fatty acid beta-oxidation, lipase complex assembly, matrix metalloproteinase activation, glutathione metabolism, and ABC transporters, consistent with known glial biology finding
- ★ A manual, documented procedure for curating pathway-gene lists from the KEGG and Reactome websites and converting them into GSEA-compatible gmt files method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Gene Set Enrichment Analysis (GSEA 4.1.0) | Drosophila melanogaster brain, FACS-sorted neurons (Elav-GAL4>UAS-GFP) vs glia (Repo-GAL4>UAS-GFP) | none (cell-type/phenotype comparison) | pathway enrichment scores, gene set correlation with phenotype class, FDR/nominal p-values | GSEA 4.1.0 software; GeneChip microarray dataset GSE45344 |
| Fluorescence Activated Cell Sorting (FACS) | Drosophila brain cells expressing GFP under Repo or Elav driver | none | isolation of GFP-positive neuron or glia populations for expression profiling | — |
| Manual pathway-gene curation / database mining | Reactome database, Drosophila-specific pathways | none | gene set membership (FlyBase IDs) per pathway, compiled into gmt file | — |
| Manual pathway-gene curation / database mining | KEGG database, Drosophila-specific pathway maps | none | gene set membership (FlyBase IDs) per pathway, compiled into gmt file | — |
- – Using Reactome gmt, 5616 genes (41.2%) associated with ELAV (neuron) class with correlation area 40.7%, and 7999 genes (58.8%) associated with REPO (glia) class with correlation area 59.3% 58.8% vs 41.2%
- ▲ For ELAV group with Reactome gmt, 214 or 148 gene sets significantly upregulated at FDR<25% with nominal p<5% or p<1% respectively 214/148 gene sets
- ▲ For REPO group with Reactome gmt, 68 or 52 gene sets significantly upregulated at FDR<25% with nominal p<5% or p<1% respectively 68/52 gene sets
- – With KEGG gmt, 98 of 131 gene sets passed filtering; ELAV group had 27 or 20 gene sets upregulated (FDR<25%, p<5%/1%), REPO group had 31 or 19 gene sets upregulated 27/20 (ELAV) vs 31/19 (REPO) gene sets
- ▲ Representative ELAV (neuron) enrichment plots highlighted mRNA splicing, transport of mature transcript to cytoplasm, PIP synthesis at plasma membrane, RAC1 GTPase cycle (Reactome), and spliceosome, endocytosis, WNT signaling, mTOR signaling (KEGG)
- ▲ Representative REPO (glia) enrichment plots highlighted peptide hormone metabolism, mitochondrial fatty acid beta-oxidation, LPL & LIPC lipase complex assembly, matrix metalloproteinase activation (Reactome), and fatty acid degradation, glutathione metabolism, ABC transporters, glycine/serine/threonine metabolism (KEGG)
- – Global enrichment score histograms (ES) showed a typical two-peak separation between ELAV and REPO groups for both Reactome and KEGG gmt files
- count 1450 Reactome pathways annotated with genetic information (Drosophila Reactome gene matrix file content)
- count 131 of 137 KEGG pathways supplied with genetic information (Drosophila KEGG gene matrix file content)
- count 13615 genes included in profiling dataset (input expression dataset size for GSEA)
- count 615 of 1450 Reactome gene sets used after default size filter (min15/max500) (Reactome GSEA filtering)
- correlation correlation area 40.7% (ELAV) vs 59.3% (REPO) (Reactome gmt phenotype correlation)
- count 5616 (41.2%) genes associated with ELAV; 7999 (58.8%) with REPO (Reactome gmt gene-phenotype association)
- count 98 of 131 KEGG gene sets used after default size filter (KEGG GSEA filtering)
- pvalue FDR<25% with nominal p<5% or p<1% thresholds used to call significant gene sets (significance criteria for enriched gene sets in both Reactome and KEGG analyses)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper describes a resource-generation and feasibility-validation study: Drosophila-specific gene matrix (GMT) files were compiled from Reactome (1450 gene sets) and KEGG (131 gene sets) pathway databases, then validated using GSEA 4.1.0 on a publicly available Drosophila brain microarray dataset (GEO accession GSE45344) comparing ELAV-GFP-labeled neurons versus REPO-GFP-labeled glia. Enrichment significance was evaluated using GSEA-native FDR (<25%) combined with nominal p-value thresholds (5% and 1%), and results were presented as counts of significantly enriched gene sets together with representative enrichment score plots. The primary aim was to demonstrate that the generated GMT files successfully separate the two biologically distinct cell types and recover known pathway signatures.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Gene Set Enrichment Analysis (GSEA) with permutation-based enrichment scoring and FDR estimation | ELAV-GFP (neuron) vs REPO-GFP (glia) comparison across 615 retained Reactome gene sets and 98 retained KEGG gene sets | 13615 genes included; biological replicates per group not explicitly stated — phenotype labels 'Elav_1' and 'Repo_1' suggest potentially one sample per group | not stated |
-
Validation used what appears to be a single labeled sample per group ('Elav_1', 'Repo_1'), making gene-set permutation the operative null-distribution strategy in GSEA↳ Could also: Phenotype permutation with multiple biological replicates per cell type is the GSEA-recommended approach for two-group comparisons — Phenotype permutation preserves gene–gene correlation structure and yields better-calibrated FDR estimates; with biological replicates, it more closely models the actual sampling variability of the experiment
-
GSEA 4.1.0 desktop software was used for the enrichment analysis↳ Could also: R/Bioconductor tools such as fgsea, camera (limma), or clusterProfiler could also accept the same GMT files and perform gene set enrichment — These tools integrate naturally with R-based microarray preprocessing pipelines, support scripted reproducible workflows, and offer additional statistical options — for example, camera adjusts for inter-gene correlation, which can improve FDR calibration
-
When multiple microarray probes mapped to the same CG_ID, only the probe with the highest average intensity was retained↳ Could also: Averaging intensities across all probes for the same gene, or summarizing at the probe-set level with RMA or median polish during preprocessing, are also widely used approaches — Model-based or average summarization can reduce the influence of individual noisy probes; selecting the maximum-intensity probe may introduce an upward bias in estimated expression for genes with many probes
-
Enrichment significance was reported against a combined FDR <25% and nominal p-value threshold↳ Could also: A more conservative FDR threshold (e.g., <5% or <10%) is commonly applied, particularly in studies with larger sample sizes where the null distribution is well estimated — The GSEA default of FDR <25% is recommended for exploratory resource-validation contexts; a stricter threshold reduces the expected false-discovery proportion when the analysis shifts toward identifying specific pathways with higher confidence
-
Validation was performed on a single publicly available microarray dataset from one laboratory↳ Could also: Cross-validation on an independent Drosophila neuron/glia dataset — or on an RNA-seq dataset covering the same cell types — would also provide evidence of generalizability — A second independent validation dataset, especially from a different platform (e.g., bulk RNA-seq), would demonstrate that the GMT files recover biologically expected pathway signatures across experimental conditions and measurement technologies
-
The gene matrix files were built from Reactome and KEGG only, and validated without comparison to GO-based gene sets↳ Could also: Including GO Biological Process or WikiPathways GMT files as a comparator in the same GSEA run would also illustrate the complementary coverage of different database types — A side-by-side comparison with GO-based gene sets in the same validation experiment would concretely illustrate the added detail that curated pathway databases provide over ontology-based approaches, strengthening the motivation for the new resource
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34710184
Paper: Cheng J, Hsu LF, Juan YH, Liu HP, Lin WY. Pathway-targeting gene matrix for Drosophila gene set enrichment analysis. PLoS One 2021. PMID 34710184 · PMCID PMC8553153 · DOI 10.1371/journal.pone.0259201
Code: https://github.com/JackCheng-TW/GeneMatrix
- pinned commit
a8438bbdf7e77503a63865b4ff8cf4fbfc9d42de(2021-08-24, branchmain) - license: none stated · not archived
- Ships 4 files:
Drosophila_KEGG.gmt(71 KB),Drosophila_REACTOME.gmt(2.0 MB),SampleData.tsv(1.7 MB),README.md. This is the authors' own deposit of the gene-matrix artifact + a ready-to-run validation expression matrix.
Data: GEO GSE45344 — DeSalvo et al. 2014, "The Drosophila surface glia
transcriptome." Microarray (Affymetrix Drosophila Genome 2.0, platform GPL1322),
Drosophila melanogaster, 20 samples (GSM1102716–GSM1102735). The paper's
validation uses a 10-sample subset: 5 ELAV-GFP (neurons) + 5 REPO-GFP (glia),
already processed into the repo's SampleData.tsv (13,615 genes × 10 samples).
What the paper reports (candidate claims)
This paper has two kinds of results:
A. Gene-matrix composition (deterministic — from the shipped .gmt files)
- Reactome matrix: 1450 pathways annotated with genetic information.
- KEGG matrix: 131 pathways with genetic information (137 annotated total).
- After GSEA size filter (min=15, max=500), 615 Reactome / 98 KEGG gene sets are analyzed (computed by GSEA after intersecting each gene set with the 13,615 dataset genes — so this is a GSEA-run output, not a raw line count).
B. GSEA enrichment of GSE45344 (ELAV neurons vs REPO glia)
Tool: GSEA 4.1.0; Collapse = No_Collapse; gene-set size 15–500; significance
at nominal p ≤ 0.05 and ≤ 0.01, FDR < 25%.
- Gene–phenotype association: 5616 (41.2%) genes correlate with ELAV (Class A, correlation area 40.7%); 7999 (58.8%) with REPO (Class B, area 59.3%). (5616 + 7999 = 13,615 ✓.)
- Reactome up-regulated gene sets: ELAV 214 (p<0.05) / 148 (p<0.01); REPO 68 / 52.
- KEGG up-regulated gene sets: ELAV 27 / 20; REPO 31 / 19.
- Representative enriched pathways named (mRNA splicing, spliceosome, endocytosis, WNT/mTOR for neurons; fatty-acid degradation, glutathione metabolism, ABC transporters for glia) — qualitative, checked against top-of-list output.
In scope (pipeline-derived → attempted)
- A. Matrix composition — count pathways in the two
.gmtfiles; reproduce the GSEA effective size-filter pass counts (615 / 98). Pure counting + one GSEA run. - B. GSEA enrichment — run GSEA 4.1.0 on
SampleData.tsvwith each.gmt, phenotype ELAV vs REPO, default params, and compare: gene-association split, size-filter counts, and the four pairs of significant-gene-set counts.
Out of scope (not pipeline / not attempted)
- Manual construction of the gene matrices from the live KEGG/Reactome web sites
(manual curation, "June 4th 2021" snapshot — not a runnable pipeline; the result
of that curation, the
.gmtfiles, IS what we test). - Wet-lab FACS isolation / microarray generation underlying GSE45344 (external, DeSalvo et al.).
- Qualitative biological interpretation of enriched pathways (judgement, not compute) — we only check that the named pathways appear among the significant sets.
Reproduction strategy
Third-party-tool reproduction (P16): run the authors' deposited matrices + data through GSEA 4.1.0 exactly as the Methods describe. Heavy compute (GSEA Java runs, phenotype permutation) on «our HPC»/«infra». Deterministic counts done by direct inspection of the shipped files.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.