Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Developing a novel immune infiltration-associated mitophagy prediction model for amyotrophic lateral sclerosis using bioinformatics strategies.

Front Immunol · 2024
L1 71/100 PQI 90
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough at the DATA layer, partially at the DEG layer, not at the model layer. The brief's 'code' link is the generic R survival package, not authors' code; no authors' repo exists, so per P16 we ran the named standard Bioconductor pipeline (limma) on the paper's own GEO data (GSE112676/GSE112680) on «our HPC». RESULT: sample composition reproduces 1:1 EXACTLY (233+508=741; 164+137=301 with 75 MIM mimics correctly dropped), confirming correct data+grouping. DEG counts reproduce to the same order of magnitude and same up>down direction but ~2x the reported values (10571 vs 5256 for GSE112676; 3025 vs 2379 for GSE112680) using the paper's stated threshold (adj.P.Val<0.05, no logFC) -> graded partial; the discrepancy is best explained by an unstated additional filter in the paper (pre-filter / logFC cutoff / probe collapse / covariate adjustment), NOT by fabrication. All 22 named mitophagy genes are real and present on the array (exact), but only 14/22 are DEGs in our run (partial). NOT ATTEMPTED (hard 20%): the 18-gene LASSO signature + risk score + time-dependent ROC/AUC and CIBERSORT immune infiltration. Survival data (survival_yr+censored) IS present so the Cox/LASSO model is technically feasible, but we deliberately did not chase it because (a) the LASSO gene names in PMC are OCR-garbled/non-existent as printed, (b) LASSO is seed-dependent, (c) the reported training AUC of exactly 1.000 is a strong overfitting/leakage flag raised for the human auditor. RT-qPCR validation is wet-lab (out of scope).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 71
    assessed: 2026-06-15 ⛓ 7c3764a7df1f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether mitophagy-related genes associated with immune infiltration can be used to construct a diagnostic and prognostic prediction model for amyotrophic lateral sclerosis (ALS).

Core claims
  • An 18-gene mitophagy-related prognostic signature was identified by machine learning and used to build a prognostic risk score model for ALS. resource
  • 22 mitophagy-related differentially expressed genes and 40 prognostic genes were identified in ALS. finding
  • Four mitophagy-related immune infiltration genes (BCKDHA, JTB, KYNU, GTF2H5) show expression in ALS mouse models and patients consistent with bioinformatics predictions, supporting prognostic potential. finding
  • High-risk and low-risk groups differ in innate and adaptive immune cell infiltration, particularly T lymphocyte subgroups. finding
  • Functional enrichment shows involvement of oxidative phosphorylation, unfolded protein response, KRAS, mTOR signaling, and immune-related pathways. mechanism
  • A LASSO-Cox regression pipeline combined with univariate Cox analysis was used to derive prognostic feature genes and risk scores. method
  • CIBERSORT was used to estimate proportions of 22 immune cell subsets to link the risk model with immune infiltration. method
Experimental setups
Assay System Perturbation Readout Platform
Bulk gene expression microarray (DEG/limma analysis) Human peripheral blood, GSE112676 (233 ALS, 508 control) and GSE112680 (164 ALS, 137 control) none (ALS vs control comparison) differentially expressed genes (logFC, adjusted p-value) Illumina HumanHT-12 V3.0 and V4.0 expression bead chip arrays
CIBERSORT immune cell deconvolution Human ALS gene expression data (GSE112676) none (high-risk vs low-risk groups) proportions of 22 immune cell subsets
GSEA hallmark enrichment / GO / KEGG enrichment analysis Human ALS gene expression (GSE112676 high- vs low-risk groups) none pathway enrichment scores MSigDB 50 hallmark gene sets; clusterProfiler
LASSO-Cox regression prognostic modeling / ROC analysis Human ALS cohorts (GSE112676 training, GSE112680 validation) none risk score, survival prognosis, AUC values glmnet, survival, pROC R packages
RT-qPCR Lumbar spinal cord of B6SJL-Tg(SOD1 G93A) ALS mouse model SOD1 G93A transgene (ALS model) relative gene expression (2^-ΔΔCt) FastKing One-Step RT-PCR Kit; GAPDH internal control
RT-qPCR Peripheral blood of 10 ALS patients vs 10 age/sex-matched healthy individuals none (disease vs control) relative gene expression (2^-ΔΔCt) FastKing One-Step RT-PCR Kit; GAPDH internal control
Key results
  • 5256 DEGs identified between ALS and control in training set GSE112676 (2822 up, 2434 down) 5256 genes
  • 34 ALS-associated mitophagy genes obtained from GeneCards; 25 overlapping genes identified after comparison with training set 34 / 25 genes
  • Eight mitophagy genes (ATG12, ATG5, MAP1LC3B, MFN1, OPTN, SRC, TOMM20, TOMM7) upregulated in ALS
  • Four mitophagy genes (CDC37, MFN2, SQSTM1, TOMM40) downregulated in ALS
  • Validation set GSE112680 DEGs: 1300 upregulated, 1079 downregulated 2379 genes
  • 4405 genes identified by Spearman correlation (P<0.05, |r|>0.3) between DEGs and mitophagy genes 4405 genes
  • BCKDHA, JTB, KYNU, GTF2H5 expression consistent between ALS models/patients and bioinformatics analysis
Key statistics
  • count 397 ALS patients with survival information (combined ALS samples with follow-up across both datasets)
  • count 233 ALS / 508 control (GSE112676 training cohort composition)
  • count 164 ALS / 137 control (GSE112680 validation cohort composition)
  • other median survival time 2.42 years [IQR 1.59, 3.52] (survival from onset to death/tracheostomy/NIV)
  • count 342 (86.15%) dead, 55 (13.85%) survival (status of overall cohort)
  • other C9orf72 repeat expansion 12.8% vs 5.2% (GSE112680 vs GSE112676 cohorts)
  • mean mean ages 63.92 and 63.58 (ALS and control groups age of onset)
  • count 10 ALS patients and 10 matched healthy controls (RT-qPCR validation cohort)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics study analyzed two GEO microarray datasets (GSE112676, n=741; GSE112680, n=301) to identify mitophagy-related differentially expressed genes in ALS. The main workflow used limma-based DEG analysis with Benjamini-Hochberg FDR correction, Spearman correlation for gene filtering, and univariate Cox followed by LASSO-Cox regression to build an 18-gene prognostic risk score model. Log-rank tests compared survival between median-split high- and low-risk groups; immune infiltration was estimated via CIBERSORT and compared across groups using Wilcoxon tests; and four candidate genes were validated by RT-qPCR in an ALS mouse model and 10 ALS patient blood samples.

Replicationbiological Sample sizeTraining n=233 ALS + 508 CON (GSE112676); validation n=164 ALS + 137 CON (GSE112680); RT-qPCR n=10 ALS + 10 age/sex-matched controls; no formal power calculation stated GroupsALS vs CON; high-risk vs low-risk ALS; training cohort vs validation cohort Pairingunpaired Randomization/blindingnot stated DispersionIQR Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
limma moderated t-test with Benjamini-Hochberg FDR correction Differential expression between ALS and CON groups, training set GSE112676 741 (233 ALS + 508 CON) not stated
limma moderated t-test with Benjamini-Hochberg FDR correction Differential expression between high-risk and low-risk ALS groups, GSE112676 233 ALS patients not stated
Spearman correlation Correlation between 25 mitophagy genes and 5256 DEGs to identify mitophagy-related gene set 741 (GSE112676 samples) not stated
Univariate Cox proportional hazards regression Identification of prognostic mitophagy-related candidate genes from training set 233 ALS patients with survival information not stated
LASSO-Cox regression (penalized Cox proportional hazards) Feature gene selection and risk score model construction 233 ALS patients (training set GSE112676) not stated
Log-rank test Survival comparison between high-risk and low-risk groups (training and validation) 233 ALS training; 164 ALS validation not stated
Wilcoxon rank sum test Immune cell infiltration proportions and TMB scores between high- and low-risk groups; feature gene expression ALS vs CON in both datasets; checkpoint gene and HLA family gene expression between risk groups 233 ALS (training); 164 ALS (validation) not stated
Time-dependent ROC / AUC Prognostic model accuracy at 5, 7, 10 years (training) and 1, 2, 3 years (validation) 233 ALS (training); 164 ALS (validation) na
GSEA (gene set enrichment analysis) Hallmark pathway enrichment between high- and low-risk groups using 50 MSigDB hallmark gene sets 233 ALS (training set) na
Pearson Chi-squared test; Wilcoxon rank sum test Baseline characteristic comparisons between training and validation cohorts (Table 1) 397 total (233 training + 164 validation) not stated
2^-ΔΔCt method (RT-qPCR relative quantification) Validation of four feature genes in ALS patient blood and SOD1-G93A mouse lumbar spinal cord 10 ALS patients + 10 matched healthy controls; mouse n not stated not stated
Approaches that could also have been used
  • Multiple Wilcoxon tests were applied simultaneously to 22 immune cell subset proportions, a panel of checkpoint genes, and 17 HLA family genes between risk groups, without a stated multiplicity correction
    Could also: Apply Benjamini-Hochberg FDR correction across the family of simultaneous comparisons within each panel — When many tests are conducted on the same samples in parallel, FDR adjustment would also control the expected proportion of false discoveries; this is standard practice in high-dimensional immune-profiling comparisons and would align these tests with the correction already applied in the DEG analyses
  • The median Risk score of the training cohort was used to dichotomize patients into high- and low-risk groups, and this same threshold was applied to the validation cohort
    Could also: Use a cutpoint derived by cross-validation within the training set (e.g., maximally selected rank statistics or pre-specified quartile) and report the cutpoint value explicitly — A median split on training data optimizes apparent group separation on that set; cross-validated or pre-specified cutpoints would also partition patients while providing an estimate of how well the threshold generalizes, reducing the risk of optimistic apparent separation
  • Univariate Cox regression at P<0.05 was used to pre-screen the mitophagy-related gene candidates before entering them into LASSO-Cox
    Could also: Apply LASSO-Cox regularization directly to all mitophagy-related candidate genes without a univariate pre-filtering step — Univariate pre-filtering can exclude genes whose marginal effects are small but jointly informative; applying penalized regression to the full candidate set would also allow the regularization itself to determine relevance, which is the conventional approach when the candidate set is not excessively large
  • LASSO-Cox was the sole machine learning method used for feature selection and risk model construction
    Could also: Apply elastic net Cox regression (mixing parameter alpha between 0 and 1) or a random survival forest as an alternative or complementary approach — LASSO tends to select one gene from a correlated cluster; elastic net handles correlated predictors differently, and random survival forests could also capture non-linear gene–survival relationships — comparing concordant features across methods is a common approach to increase confidence in selected biomarkers
  • Immune cell deconvolution was performed exclusively with CIBERSORT, estimating 22 immune cell subsets
    Could also: Apply additional deconvolution methods such as xCell, TIMER2.0, or EPIC in parallel and report concordance across methods — Different deconvolution algorithms use different reference gene sets and assumptions; findings that are consistent across multiple methods would also provide stronger evidence for immune infiltration differences than a single tool alone
  • The RT-qPCR small-sample validation (n=10 ALS, n=10 controls) reports relative expression via 2^-ΔΔCt for four genes; no formal between-group statistical test or multiplicity correction is explicitly described for this comparison
    Could also: Apply a Mann-Whitney U test (or paired test exploiting the age/sex matching) to the ΔCt values for each gene, report exact p-values, and apply a Bonferroni or FDR correction across the four genes — Formal hypothesis testing with stated statistics and multiplicity adjustment would also quantify uncertainty in the small-sample validation and make the validation directly comparable to the bioinformatics significance thresholds
Software: R / limma R 4.1.0 and 4.3.0; limma V-3.84.3 · R / glmnet (LASSO-Cox) V-4.1-2 · R / survival V3.2.13 · R / clusterProfiler (GSEA, GO, KEGG) V-4.6.2 · R / pROC V-1.18.2 · R / corrplot v-0.90 · CIBERSORTx (online tool) · RStudio 3.84.3

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
7
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE112676 GEO in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet
GSE112680 GEO in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet

Downstream reach in the literature

12 downstream papers · 2 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: TableFig 1Fig 2Fig ROC
C0a
Reported
GSE112676: 233 ALS + 508 control = 741
Reproduced
233 ALS + 508 CON = 741
exact
C0b
Reported
GSE112680: 164 ALS + 137 control = 301
Reproduced
164 ALS + 137 CON = 301 (75 MIM excluded)
exact
C1
Reported
5256 DEGs in GSE112676 (2822 up/2434 down), limma adj.P.Val<0.05
Reproduced
10571 probe-level / 7607 unique-gene DEGs (5386 up/5185 down)
partial
C2
Reported
2379 DEGs in GSE112680 (1300 up/1079 down)
Reproduced
3025 probe-level / 2517 unique-gene DEGs (1697 up/1328 down)
partial
C3a
Reported
22 named mitophagy genes retained
Reproduced
22/22 present on Illumina HT-12 array
exact
C3b
Reported
all 22 mitophagy genes are DEGs (DEG-intersect-mitophagy)
Reproduced
14/22 are DEGs in our limma run
partial
C6
Reported
prognostic model ROC AUC train 5/7/10-yr = 0.933/0.966/1.000; valid 0.643/0.709/0.630
Reproduced
not attempted (hard 20%)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The data layer reproduces 1:1 (sample counts 741 and 301 exact, MIM mimics correctly dropped), confirming correct data and grouping. The DEG counts deviate ~2x (10571 vs 5256; 3025 vs 2379) with direction/order-of-magnitude preserved — most plausibly an unstated filtering step on the authors' side, an under-specification rather than fabrication. The central prognostic model was not reproducible: the 18-gene LASSO names are OCR-garbled and the reported training AUC=1.000 (vs validation 0.63–0.71) is a clear overfit/leakage flag, so the core claim is only weakly supported. Overall a solid-but-deviating reproduction with the main risk concentrated in an unverifiable, over-optimistic model.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

177.7 k
tokens (I/O) · 16.5 M incl. cache
31 min
runtime · 0.02 CPU-h
1.7 GB
peak RAM
2
HPC jobs
hummel
machine