The motor neuron m6A repertoire governs neuronal homeostasis and FTO inhibition mitigates ALS symptom manifestation.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Paper described well enough to run; data + code are public (MIT repo + 3.58GB figshare Seurat object). Ran the documented Seurat FindMarkers(bimod, p_adj<0.05) DEG pipeline on the shipped integrated multiome object on «our HPC» («job», Seurat 5.1.0). SANITY claims reproduce 1:1: replicate structure (3 WT + 3 KO), 7 annotated cell types, and genotype-stable cell-type proportions all EXACT. PRIMARY DEG-count claim does NOT reproduce: observed 3301/1903/768 DEGs (Skeletal/Visceral/Cholinergic) vs reported 652/500/604 — ~5x/3.8x/1.3x too many. Root cause is honest and documented: the published Figures.R FindMarkers() call is TRUNCATED (trailing comma, continuation line with the threshold arguments lost during the single 'Add files via upload' commit; ggsave() truncated identically), so the exact min.pct/logfc.threshold are unrecoverable; the authors' derived DEGs_bimod_KOCtrl_subtypes.csv is NOT deposited on figshare; and the bimod test + Bonferroni denominator are Seurat-version sensitive (authors' version unpinned). A threshold sweep (|log2FC| 0.25/0.5/1.0, min.pct 0.1/0.25/0.5) found NO single consistent post-filter reproducing all three reported counts, so no value was forced. The up>down direction and subtype DEG-burden ranking are preserved. NOT attempted: nanopore m6A site count (Fig 5d, wet-lab, out of scope), GO/KEGG term lists (downstream of unmatched DEGs).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 59assessed: 2026-06-20 ⛓ 3c9b43fce53d
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-20
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether N6-methyladenosine (m6A) RNA hypomethylation, driven by reduced METTL3/METTL14 methyltransferase activity, is a direct causative driver of motor neuron degeneration in both familial and sporadic ALS, and whether restoring m6A levels can mitigate ALS-associated MN degeneration and motor symptoms.
- ★ m6A hypomethylation (not hypermethylation) is consistently associated with ALS across patient iPSC-MNs, postmortem tissue, and independent transcriptomic datasets finding
- ★ Conditional knockout of Mettl14 in motor neurons (ChAT-Cre;Mettl14floxed mice) recapitulates the major molecular, cellular, and behavioral hallmarks of ALS finding
- ★ Reduced METTL3/METTL14 expression and consequent hypo-m6A precedes and causes motor neuron degeneration in familial ALS iPSC-MN lines (SOD1+/L144F, C9ORF72exp~800G4C2, TDP43G298S) finding
- ★ Pharmacological inhibition of METTL3 (STM2457) in wild-type iPSC-MNs is sufficient to reduce m6A levels and induce neurite degeneration finding
- ★ shRNA knockdown of METTL3 or METTL14 reduces m6A-mRNA levels and causes neurite degeneration finding
- ★ Pervasive m6A hypomethylation dysregulates high-risk ALS-associated genes and reduces chromatin accessibility in motor neurons mechanism
- ★ Intrathecal Fto-shRNA knockdown ameliorates motor deficits and extends lifespan in SOD1G93A ALS mice finding
- ★ ChAT-Cre;Mettl14floxed mice represent a novel sporadic-ALS-like mouse model with TDP43/FUS cytoplasmic aggregation, NMJ denervation, and premature death resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| transcriptomic/expression analysis (Answer ALS dataset) | human familial/sporadic ALS iPSC-derived MNs and postmortem cortex | none (disease state) | METTL3/5/14/16 mRNA expression levels | — |
| m6A ELISA and dot blot | familial ALS iPSC-MNs (SOD1+/L144F, C9ORF72exp~800G4C2, TDP43G298S) vs isogenic/healthy controls | CPA (cyclopiazonic acid) ER stress | m6A methylation percentage in poly(A) mRNA over differentiation timecourse | — |
| m6A ELISA and dot blot with pharmacologic METTL3 inhibition | wild-type human iPSC-MNs | STM2457 (METTL3 inhibitor, 20 μM) | m6A levels and neurite degeneration index | — |
| shRNA knockdown and m6A quantification | HEK293T cells | METTL3-shRNA or METTL14-shRNA | m6A-mRNA levels | — |
| lentiviral shRNA knockdown | motor neurons | LV-shMETTL3 / LV-shMETTL14 | neurite degeneration | lentivirus |
| conditional gene knockout (Cre-lox) | mouse spinal motor neurons (Olig2-Cre;Mettl14floxed and ChAT-Cre;Mettl14floxed) | Mettl14 knockout | MN survival, body weight, lifespan, Mettl14 protein loss | — |
| immunostaining/immunohistochemistry | ChAT-Cre;Mettl14floxed mouse lumbar spinal cord | Mettl14 knockout | ChAT+ MN counts, C-bouton density, Iba1+ microglia activation, Tdp43/Fus cytoplasmic localization | — |
| neuromuscular junction confocal imaging | ChAT-Cre;Mettl14floxed mouse gastrocnemius muscle | Mettl14 knockout | NMJ denervation ratio and endplate area (SV2/NF and α-BTX staining) | confocal microscopy |
- ▼ ALS iPSC-MNs show reduced m6A levels at day 4 post-CPA, preceding drastic MN loss seen by day 7
- ▼ STM2457 treatment sharply reduces m6A after 4 days and causes dramatic neurite degeneration by day 6
- – ChAT-Cre;Mettl14floxed mice (n>90) exhibit premature death between P160-P300
- ▼ ChAT+ MN numbers in lumbar spinal cord decline significantly starting after P100, with significant loss by P160
- ▼ C-bouton cholinergic nerve terminals on MNs decrease from as early as P70
- ▲ Microglial (Iba1+) activation is significantly increased in ChAT-Cre;Mettl14floxed spinal cords at P160 but not before P120
- – Tdp43 and Fus shift from nuclear to cytoplasmic aggregation in mutant MNs after P120
- – ChAT-Cre;Mettl14floxed mice show reduced NMJ endplate area and increased muscle denervation
- count n = 3 independent experiments (m6A ELISA/dot blot quantification in iPSC-MN experiments (Fig. 1b-g, i, j))
- count n = 6 independent experiments (degeneration index quantification after STM2457 treatment (Fig. 1l))
- count n > 90 (ChAT-Cre;Mettl14floxed mice (male and female) exhibiting premature death)
- count n = 3-5 mice per timepoint (P30, P70, P100, P160, P250) (ChAT+ MN quantification in lumbar spinal cord over time)
- count n = 5 mice (Iba1+ microglial activation staining/quantification)
- count n = 3 mice (Tdp43 aggregate and NMJ denervation/endplate area quantification)
- other STM2457 dose = 20 μM applied for 6 days (METTL3 inhibitor treatment protocol during iPSC-MN differentiation)
- other lethality onset ~P24-P28 in Olig2-Cre;Mettl14floxed mice (developmental-stage Mettl14 knockout survival)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combines human iPSC-derived motor neuron (MN) experiments and a conditional Mettl14-knockout mouse model to test whether m6A hypomethylation drives ALS-like degeneration. Comparisons between control and ALS/mutant conditions across many outcomes (m6A levels, MN counts, degeneration indices, body weight, microglial activation, NMJ denervation) were assessed with two-tailed t-tests, and premature death was visualized with Kaplan–Meier survival curves. Results are reported as mean ± SD, with significance denoted by P values or 'N.S.' rather than exact values printed in the main text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-tailed Student's t-test | MN degeneration index and m6A levels in iPSC~MN experiments (Fig. 1b–l) | n = 3 or n = 6 independent experiments | not stated |
| two-tailed Student's t-test | ChAT+ MN counts, C-bouton quantification, Iba1+ microglia counts, Tdp43 aggregation, NMJ denervation/endplate area in ChAT-Cre;Mettl14 floxed mice (Figs. 2c–f, 3a–h) | n = 3–5 mice per timepoint/group | not stated |
| Kaplan-Meier survival curve (statistical test for curve comparison not named) | survival of ChAT-Cre;Mettl14 floxed vs littermate control mice (Fig. 2b) | n > 90 mice | not stated |
-
Repeated two-tailed t-tests were used to compare genotype groups at multiple individual timepoints (e.g., MN counts, body weight, C-bouton numbers across P30–P250) within the same figure.↳ Could also: A two-way (genotype × time) ANOVA or a repeated-measures/mixed-effects model, followed by a post-hoc test with multiplicity correction (e.g., Sidak or Tukey) — This would model the genotype-by-time interaction directly, account for correlated repeated measurements from the same animals across timepoints, and control the family-wise error rate across the multiple time-point comparisons in one framework.
-
Survival differences between ChAT-Cre;Mettl14 floxed and control mice are shown as Kaplan–Meier curves without a named statistical test for comparing the curves.↳ Could also: A log-rank (Mantel-Cox) test, optionally paired with a Cox proportional hazards model — This is a standard approach for formally quantifying the difference between two survival curves and can also yield a hazard ratio with a confidence interval, adding a magnitude-of-effect measure alongside the visual comparison.
-
Group comparisons throughout are summarized as mean ± SD with significance denoted by P value thresholds or 'N.S.'↳ Could also: Reporting exact P values alongside an effect size measure (e.g., Cohen's d or mean difference with 95% CI) — Exact P values and effect sizes convey the magnitude and precision of a difference in addition to statistical significance, which can be especially informative when sample sizes are small (n = 3–6).
-
Sample sizes are relatively small (e.g., n = 3 independent experiments or n = 3–5 mice per group) without a stated power analysis.↳ Could also: An a priori power analysis to justify sample size, or nonparametric tests (e.g., Mann-Whitney U) as a complement to the t-test — Power analysis helps confirm adequate sensitivity to detect the effect sizes of interest, and nonparametric alternatives can be useful when normality is difficult to verify with small n.
-
Multiple distinct familial ALS iPSC lines (SOD1, C9ORF72, TDP43) were each compared separately to their isogenic/healthy controls using pairwise t-tests.↳ Could also: A one-way ANOVA across all iPSC lines and controls simultaneously, with a post-hoc test (e.g., Dunnett's, comparing each line to control) — Testing all groups within a single omnibus model can reduce the number of separate pairwise tests performed and provides a unified control of the error rate across the several ALS lines examined together.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40307231
Paper: Yen et al. 2025, Nat Commun 16. "The motor neuron m6A repertoire governs neuronal homeostasis and FTO inhibition mitigates ALS symptom manifestation." DOI 10.1038/s41467-025-59117-2 · PMCID PMC12043976.
Code: https://github.com/jaclab-multiomic/Yen-et-al.-2025-Nat-Comm (MIT, public, 5 files, ~22 KB of R scripts; last push 2025-02-17). Authors' own code (not P16).
Processed data (inputs to the scripts): Figshare project 238025 —
mettl_integrated_slim.rds.gz(3.5 GB, art 28431005, file 52420691, md5 eecbc9be…) — the integrated WNN multiome Seurat object (RNA+ATAC). Primary input.cholinergicneurons_Mettl14.rds.gz,skeletalmns.rds.gz, metadata csvs.
Pipeline-derived results (the scripts are Figures.R, ArchR_peak2geneLinkage.R, stackvlnplot_edited.R)
| Result | Pipeline | In scope? | Why |
|---|---|---|---|
| DEG counts KO-vs-WT per neuron type (Skeletal MNs 652=291↓/361↑; Visceral MNs 500=190↓/310↑; Cholinergic INs 604=165↓/439↑) | Seurat FindMarkers(test.use="bimod"), p_adj<0.05, on the slim object |
YES — primary target | Deterministic test, exact integer counts reported in Results, fully specified in Figures.R, runs from the shipped object alone. |
| Cell types / clusters (Skeletal MNs, Visceral MNs, Cholinergic INs + Excitatory, Inhibitory, Oligodendrocytes, Astrocytes) | stored annotated_cluster Idents |
YES (sanity) | Derivable directly from object metadata. |
| Replicates per genotype (3 KO: r1/r3/r4; 3 WT: r3/r4/r5) | orig.ident metadata |
YES (sanity) | Direct count. |
| Fig 6D/7D cell-type proportions per genotype | prop.table(table(Idents,genotype)) |
YES (report; no exact % in text) | Deterministic table; paper only states "proportions unaffected" qualitatively — report table, no numeric grade. |
| Fig 7E/7F GO/KEGG terms on m6A-tagged DEGs | clusterProfiler enrichGO/enrichKEGG | partial / optional | Downstream of DEGs; version-sensitive (clusterProfiler 4.6.2, org.Mm.eg.db 3.16.0). Last-20% — attempt only if DEGs match. |
| Fig 5d m6A sites: 30,340 sites / 7,921 genes | EpiNano + m6anet nanopore union | OUT | Wet-lab nanopore DRS pipeline upstream of the deposited gene list; raw DRS not the focus; out of multiome scope. |
| Fig S8 skeletal-MN α/γ/γ* label transfer | TransferData vs Blum et al ref |
OUT | Authors note Seurat-version-dependent + non-deterministic; 14 cells dropped manually. Explicitly flagged unstable in code comments. |
| ArchR peak2gene linkage | ArchR | OUT (optional) | Needs full ArchR project (not on Figshare); heavy. |
80/20 plan
Reproduce the DEG counts (652/500/604 with up/down split) — the clearly-specified,
deterministic, integer-valued primary claim — by running the exact FindMarkers bimod
step from Figures.R on the shipped mettl_integrated_slim.rds. Report cell-type list,
replicate count, and proportion table as supporting sanity checks. Do NOT chase GO/KEGG
term lists, label transfer, or the nanopore m6A pipeline (last 20%, version-sensitive /
out of scope).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
All sanity claims reproduce 1:1 (3 WT+3 KO replicates, 7 exact cell-type labels, genotype-stable proportions), but the primary DEG-count claim fails: 3301/1903/768 reproduced vs reported 652/500/604 (up to 5.1× too many). The discrepancy sits on the authors'/data-availability side — the deposited Figures.R FindMarkers() call is truncated (exact thresholds lost), the derived DEG CSV is not on figshare, and Seurat 4→5.1.0 version drift inflates raw calls; a threshold sweep recovered none of the three counts, so nothing was forced. Severity is severe in magnitude yet the qualitative biology (up>down skew, subtype DEG-burden ranking) is preserved and there is no fabrication signal — the blocker is incomplete deposited code/data, not a fabricated value.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.