Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Pathway-targeting gene matrix for Drosophila gene set enrichment analysis.

PLoS One · 2021
L1 95/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1, third-party-tool P16). The authors' GeneMatrix repo ships two GMT pathway matrices (Drosophila KEGG + Reactome) plus a ready-to-run 13615x10 microarray validation matrix derived from GSE45344 (ELAV neurons vs REPO glia). Ran GSEA 4.2.3 on «our HPC» (paper used 4.1.0, no longer downloadable; algorithm unchanged). EXACT: KEGG=131 pathways, total genes=13615, size-filter pass 615 Reactome / 98 KEGG, gene-association split ELAV 5616 (41.2%) / REPO 7999 (58.8%). WITHIN-TOL: significant-set counts Reactome ELAV 216/154 (vs 214/148), REPO 68/49 (vs 68/52); KEGG ELAV 26/18 (vs 27/20), REPO 32/20 (vs 31/19) -- all within +-6 (REPO p<=0.05 exact), deltas explained by permutation stochasticity + GSEA version. Qualitative claims confirmed (mRNA splicing tops neurons; fatty-acid degradation + glutathione metabolism top glia). KEY METHOD NOTE: the paper says 'GSEA default parameters' but the reported counts are only reproducible with gene_set permutation, not GSEA's default phenotype permutation (which is degenerate at 5v5 and gives ~4x fewer sig sets, e.g. 53 vs 214). gene_set permutation is GSEA's own documented recommendation for n<7 per class -- derivable from shipped data, not a fabrication concern, but an under-specified Methods detail. NOT attempted (out of scope): manual web-curation of the matrices from KEGG/Reactome (no runnable pipeline; the GMT files are the curated result, which IS what we tested) and the wet-lab FACS+microarray generation of GSE45344 (DeSalvo et al., external).

💻 Code ↗ 🗄 Data: GSE45344

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 95
    assessed: 2026-06-21 ⛓ eb992c668a03
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether newly generated Reactome- and KEGG-based gene matrix (gmt) files for Drosophila melanogaster can be validated by GSEA to correctly distinguish and characterize known biological differences between neurons and glia.

Core claims
  • Gene matrix files for GSEA are largely unavailable for Drosophila, limiting pathway-level enrichment analysis in this model organism finding
  • Generated a Drosophila-specific Reactome gene matrix file covering 1450 annotated pathways resource
  • Generated a Drosophila-specific KEGG gene matrix file covering 131 annotated pathways with genetic information resource
  • The Reactome gmt file successfully separates neuron (ELAV) and glia (REPO) expression profiles via GSEA finding
  • The KEGG gmt file successfully separates neuron (ELAV) and glia (REPO) expression profiles via GSEA finding
  • Neuron-enriched pathways include mRNA splicing/spliceosome, transport of mature transcript, PIP synthesis, RAC1 GTPase cycle, WNT signaling, mTOR signaling, and endocytosis, consistent with known neuronal biology finding
  • Glia-enriched pathways include peptide hormone metabolism, mitochondrial fatty acid beta-oxidation, lipase complex assembly, matrix metalloproteinase activation, glutathione metabolism, and ABC transporters, consistent with known glial biology finding
  • A manual, documented procedure for curating pathway-gene lists from the KEGG and Reactome websites and converting them into GSEA-compatible gmt files method
Experimental setups
Assay System Perturbation Readout Platform
Gene Set Enrichment Analysis (GSEA 4.1.0) Drosophila melanogaster brain, FACS-sorted neurons (Elav-GAL4>UAS-GFP) vs glia (Repo-GAL4>UAS-GFP) none (cell-type/phenotype comparison) pathway enrichment scores, gene set correlation with phenotype class, FDR/nominal p-values GSEA 4.1.0 software; GeneChip microarray dataset GSE45344
Fluorescence Activated Cell Sorting (FACS) Drosophila brain cells expressing GFP under Repo or Elav driver none isolation of GFP-positive neuron or glia populations for expression profiling
Manual pathway-gene curation / database mining Reactome database, Drosophila-specific pathways none gene set membership (FlyBase IDs) per pathway, compiled into gmt file
Manual pathway-gene curation / database mining KEGG database, Drosophila-specific pathway maps none gene set membership (FlyBase IDs) per pathway, compiled into gmt file
Key results
  • Using Reactome gmt, 5616 genes (41.2%) associated with ELAV (neuron) class with correlation area 40.7%, and 7999 genes (58.8%) associated with REPO (glia) class with correlation area 59.3% 58.8% vs 41.2%
  • For ELAV group with Reactome gmt, 214 or 148 gene sets significantly upregulated at FDR<25% with nominal p<5% or p<1% respectively 214/148 gene sets
  • For REPO group with Reactome gmt, 68 or 52 gene sets significantly upregulated at FDR<25% with nominal p<5% or p<1% respectively 68/52 gene sets
  • With KEGG gmt, 98 of 131 gene sets passed filtering; ELAV group had 27 or 20 gene sets upregulated (FDR<25%, p<5%/1%), REPO group had 31 or 19 gene sets upregulated 27/20 (ELAV) vs 31/19 (REPO) gene sets
  • Representative ELAV (neuron) enrichment plots highlighted mRNA splicing, transport of mature transcript to cytoplasm, PIP synthesis at plasma membrane, RAC1 GTPase cycle (Reactome), and spliceosome, endocytosis, WNT signaling, mTOR signaling (KEGG)
  • Representative REPO (glia) enrichment plots highlighted peptide hormone metabolism, mitochondrial fatty acid beta-oxidation, LPL & LIPC lipase complex assembly, matrix metalloproteinase activation (Reactome), and fatty acid degradation, glutathione metabolism, ABC transporters, glycine/serine/threonine metabolism (KEGG)
  • Global enrichment score histograms (ES) showed a typical two-peak separation between ELAV and REPO groups for both Reactome and KEGG gmt files
Key statistics
  • count 1450 Reactome pathways annotated with genetic information (Drosophila Reactome gene matrix file content)
  • count 131 of 137 KEGG pathways supplied with genetic information (Drosophila KEGG gene matrix file content)
  • count 13615 genes included in profiling dataset (input expression dataset size for GSEA)
  • count 615 of 1450 Reactome gene sets used after default size filter (min15/max500) (Reactome GSEA filtering)
  • correlation correlation area 40.7% (ELAV) vs 59.3% (REPO) (Reactome gmt phenotype correlation)
  • count 5616 (41.2%) genes associated with ELAV; 7999 (58.8%) with REPO (Reactome gmt gene-phenotype association)
  • count 98 of 131 KEGG gene sets used after default size filter (KEGG GSEA filtering)
  • pvalue FDR<25% with nominal p<5% or p<1% thresholds used to call significant gene sets (significance criteria for enriched gene sets in both Reactome and KEGG analyses)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper describes a resource-generation and feasibility-validation study: Drosophila-specific gene matrix (GMT) files were compiled from Reactome (1450 gene sets) and KEGG (131 gene sets) pathway databases, then validated using GSEA 4.1.0 on a publicly available Drosophila brain microarray dataset (GEO accession GSE45344) comparing ELAV-GFP-labeled neurons versus REPO-GFP-labeled glia. Enrichment significance was evaluated using GSEA-native FDR (<25%) combined with nominal p-value thresholds (5% and 1%), and results were presented as counts of significantly enriched gene sets together with representative enrichment score plots. The primary aim was to demonstrate that the generated GMT files successfully separate the two biologically distinct cell types and recover known pathway signatures.

Replicationunclear Sample sizeNot described in terms of biological replicates or power; the validation sample size per group is not stated and is inferable only from the manually entered phenotype labels ('Elav_1', 'Repo_1'), which suggest potentially n=1 per group GroupsELAV-GFP FACS-sorted neurons vs REPO-GFP FACS-sorted glia from Drosophila brains Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionGSEA-native FDR (Benjamini-Hochberg-based) with threshold <25%
Statistical tests used
Test Applied to n Assumptions
Gene Set Enrichment Analysis (GSEA) with permutation-based enrichment scoring and FDR estimation ELAV-GFP (neuron) vs REPO-GFP (glia) comparison across 615 retained Reactome gene sets and 98 retained KEGG gene sets 13615 genes included; biological replicates per group not explicitly stated — phenotype labels 'Elav_1' and 'Repo_1' suggest potentially one sample per group not stated
Approaches that could also have been used
  • Validation used what appears to be a single labeled sample per group ('Elav_1', 'Repo_1'), making gene-set permutation the operative null-distribution strategy in GSEA
    Could also: Phenotype permutation with multiple biological replicates per cell type is the GSEA-recommended approach for two-group comparisons — Phenotype permutation preserves gene–gene correlation structure and yields better-calibrated FDR estimates; with biological replicates, it more closely models the actual sampling variability of the experiment
  • GSEA 4.1.0 desktop software was used for the enrichment analysis
    Could also: R/Bioconductor tools such as fgsea, camera (limma), or clusterProfiler could also accept the same GMT files and perform gene set enrichment — These tools integrate naturally with R-based microarray preprocessing pipelines, support scripted reproducible workflows, and offer additional statistical options — for example, camera adjusts for inter-gene correlation, which can improve FDR calibration
  • When multiple microarray probes mapped to the same CG_ID, only the probe with the highest average intensity was retained
    Could also: Averaging intensities across all probes for the same gene, or summarizing at the probe-set level with RMA or median polish during preprocessing, are also widely used approaches — Model-based or average summarization can reduce the influence of individual noisy probes; selecting the maximum-intensity probe may introduce an upward bias in estimated expression for genes with many probes
  • Enrichment significance was reported against a combined FDR <25% and nominal p-value threshold
    Could also: A more conservative FDR threshold (e.g., <5% or <10%) is commonly applied, particularly in studies with larger sample sizes where the null distribution is well estimated — The GSEA default of FDR <25% is recommended for exploratory resource-validation contexts; a stricter threshold reduces the expected false-discovery proportion when the analysis shifts toward identifying specific pathways with higher confidence
  • Validation was performed on a single publicly available microarray dataset from one laboratory
    Could also: Cross-validation on an independent Drosophila neuron/glia dataset — or on an RNA-seq dataset covering the same cell types — would also provide evidence of generalizability — A second independent validation dataset, especially from a different platform (e.g., bulk RNA-seq), would demonstrate that the GMT files recover biologically expected pathway signatures across experimental conditions and measurement technologies
  • The gene matrix files were built from Reactome and KEGG only, and validated without comparison to GO-based gene sets
    Could also: Including GO Biological Process or WikiPathways GMT files as a comparator in the same GSEA run would also illustrate the complementary coverage of different database types — A side-by-side comparison with GO-based gene sets in the same validation experiment would concretely illustrate the added detail that curated pathway databases provide over ontology-based approaches, strengthening the motivation for the new resource
Software: GSEA 4.1.0 · Microsoft Excel (implied for GMT file construction)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34710184

Paper: Cheng J, Hsu LF, Juan YH, Liu HP, Lin WY. Pathway-targeting gene matrix for Drosophila gene set enrichment analysis. PLoS One 2021. PMID 34710184 · PMCID PMC8553153 · DOI 10.1371/journal.pone.0259201

Code: https://github.com/JackCheng-TW/GeneMatrix

  • pinned commit a8438bbdf7e77503a63865b4ff8cf4fbfc9d42de (2021-08-24, branch main)
  • license: none stated · not archived
  • Ships 4 files: Drosophila_KEGG.gmt (71 KB), Drosophila_REACTOME.gmt (2.0 MB), SampleData.tsv (1.7 MB), README.md. This is the authors' own deposit of the gene-matrix artifact + a ready-to-run validation expression matrix.

Data: GEO GSE45344 — DeSalvo et al. 2014, "The Drosophila surface glia transcriptome." Microarray (Affymetrix Drosophila Genome 2.0, platform GPL1322), Drosophila melanogaster, 20 samples (GSM1102716–GSM1102735). The paper's validation uses a 10-sample subset: 5 ELAV-GFP (neurons) + 5 REPO-GFP (glia), already processed into the repo's SampleData.tsv (13,615 genes × 10 samples).


What the paper reports (candidate claims)

This paper has two kinds of results:

A. Gene-matrix composition (deterministic — from the shipped .gmt files)

  • Reactome matrix: 1450 pathways annotated with genetic information.
  • KEGG matrix: 131 pathways with genetic information (137 annotated total).
  • After GSEA size filter (min=15, max=500), 615 Reactome / 98 KEGG gene sets are analyzed (computed by GSEA after intersecting each gene set with the 13,615 dataset genes — so this is a GSEA-run output, not a raw line count).

B. GSEA enrichment of GSE45344 (ELAV neurons vs REPO glia)

Tool: GSEA 4.1.0; Collapse = No_Collapse; gene-set size 15–500; significance at nominal p ≤ 0.05 and ≤ 0.01, FDR < 25%.

  • Gene–phenotype association: 5616 (41.2%) genes correlate with ELAV (Class A, correlation area 40.7%); 7999 (58.8%) with REPO (Class B, area 59.3%). (5616 + 7999 = 13,615 ✓.)
  • Reactome up-regulated gene sets: ELAV 214 (p<0.05) / 148 (p<0.01); REPO 68 / 52.
  • KEGG up-regulated gene sets: ELAV 27 / 20; REPO 31 / 19.
  • Representative enriched pathways named (mRNA splicing, spliceosome, endocytosis, WNT/mTOR for neurons; fatty-acid degradation, glutathione metabolism, ABC transporters for glia) — qualitative, checked against top-of-list output.

In scope (pipeline-derived → attempted)

  • A. Matrix composition — count pathways in the two .gmt files; reproduce the GSEA effective size-filter pass counts (615 / 98). Pure counting + one GSEA run.
  • B. GSEA enrichment — run GSEA 4.1.0 on SampleData.tsv with each .gmt, phenotype ELAV vs REPO, default params, and compare: gene-association split, size-filter counts, and the four pairs of significant-gene-set counts.

Out of scope (not pipeline / not attempted)

  • Manual construction of the gene matrices from the live KEGG/Reactome web sites (manual curation, "June 4th 2021" snapshot — not a runnable pipeline; the result of that curation, the .gmt files, IS what we test).
  • Wet-lab FACS isolation / microarray generation underlying GSE45344 (external, DeSalvo et al.).
  • Qualitative biological interpretation of enriched pathways (judgement, not compute) — we only check that the named pathways appear among the significant sets.

Reproduction strategy

Third-party-tool reproduction (P16): run the authors' deposited matrices + data through GSEA 4.1.0 exactly as the Methods describe. Heavy compute (GSEA Java runs, phenotype permutation) on «our HPC»/«infra». Deterministic counts done by direct inspection of the shipped files.

kegg_n_pathways
Reported
131
Reproduced
131
exact
reactome_n_pathways
Reported
1450
Reproduced
1450 lines (1449 real pathways + 1 header row)
within tolerance
total_genes
Reported
13615
Reproduced
13615
exact
reactome_size_filter
Reported
615
Reproduced
615
exact
kegg_size_filter
Reported
98
Reproduced
98
exact
gene_assoc
Reported
ELAV 5616 (41.2%) / REPO 7999 (58.8%)
Reproduced
ELAV 5616 (41.2%) / REPO 7999 (58.8%)
exact
reactome_sig
Reported
ELAV 214/148, REPO 68/52 (p<=0.05/<=0.01)
Reproduced
ELAV 216/154, REPO 68/49
within tolerance
kegg_sig
Reported
ELAV 27/20, REPO 31/19
Reproduced
ELAV 26/18, REPO 32/20
within tolerance
qualitative_pathways
Reported
neurons: mRNA splicing; glia: fatty-acid degradation, glutathione metabolism
Reproduced
ELAV top Reactome=MRNA SPLICING (#1-5); REPO top KEGG=FATTY ACID DEGRADATION (#1), GLUTATHIONE METABOLISM (#5)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 95/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

181.3 k
tokens (I/O) · 9.8 M incl. cache
42 min
runtime · 0.03 CPU-h
1.6 GB
peak RAM
2
HPC jobs
hummel
machine