Gene co-expression network analysis in human spinal cord highlights mechanisms underlying amyotrophic lateral sclerosis susceptibility.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
EXACT 1:1 reproduction of the in-scope pipeline result. The repo dhglab/ALS-Gene-Network-Manuscript @ b0a869c is fully self-contained: it ships the processed GTEx spinal-cord regressed expression matrix (GTEX.expr4.RData, 14577 genes x 62 samples), the covariate metadata, and the authors' output (results/Modules.RData). Re-running the module-defining chain of code/1_WGCNA.R on «our HPC» with WGCNA 1.67 (the authors' exact version), dynamicTreeCut 1.63-1, fastcluster 1.1.25 (seed 88, pickSoftThreshold, adjacency power=10, TOM, hclust average, cutreeHybrid mms=100/ds=2/cutHeight=0.9999/pamStage=F, mergeCloseModules dcor=0.1) regenerates the shipped Modules.RData EXACTLY: 13 non-grey modules (paper reports 13), every one of the 14577 genes receives the identical module color, fraction-identical = 1.000 and adjusted Rand index = 1.000, and the sample/gene counts (62/14577) and soft power (10) match the paper. All 8 in-scope claims grade exact. Described well enough to reproduce 1:1. NOT attempted (out of scope, not in repo): raw SRP064478 realignment, downstream GO/GWAS/cell-type enrichment, SC.Mx module labels, and hub-gene calls.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-18 ⛓ dd895b5f2d0e
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetGene co-expression network analysis of the human spinal cord can provide a systems-level, unbiased framework to identify the biological processes and cell types through which ALS genetic risk factors act, thereby revealing causal mechanisms of ALS pathophysiology.
- ★ WGCNA on control human cervical spinal cord RNA-seq identifies 13 co-expression modules (SC.M1-M13), each representing distinct biological processes or cell types. finding
- ★ ALS common genetic risk (via partitioned SNP heritability) is significantly enriched in module SC.M4 (RNA processing/gene regulation) and SC.M2 (intracellular transport/autophagy, oligodendrocyte-enriched). finding
- ★ Top hub genes of SC.M2 include ALS-implicated genes KPNA3, TMED2, and NCOA4, the latter regulating ferritin autophagy, implicating this process in ALS pathophysiology. mechanism
- ★ Genetic modifiers of C9orf72 and FUS toxicity from yeast screens are significantly enriched in the ribosomal-function module SC.M6. finding
- ★ Differentially expressed genes in postmortem ALS spinal cord are enriched in SC.M12 (upregulated, microglial/immune) and SC.M3 (downregulated, endothelial), but these DEGs are not enriched in causal genetic risk, suggesting they are consequences rather than drivers of disease. finding
- ALS GWAS enrichment in SC.M2/SC.M4 is disease-specific, as IBD risk variants enrich a different module (SC.M12) and ASD, AD, FTD, and PSP risk variants show no module enrichment. finding
- Literature-curated large-effect familial ALS risk genes are nominally (but not FDR-significantly) enriched in SC.M4. finding
- 10 of 13 GTEx-derived modules are preserved in an independent postmortem ALS spinal cord dataset, with SC.M2, SC.M4, SC.M5, and SC.M8 strongly preserved. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq co-expression network construction (WGCNA) | human control cervical spinal cord (GTEx, n=62) | none | gene co-expression modules, module eigengenes | — |
| stratified LD score regression / partitioned heritability | ALS GWAS summary statistics (ALS n=20,806; CTL n=59,804) | none | SNP-heritability enrichment per co-expression module | — |
| protein-protein interaction (PPI) network analysis with permutation testing | top hub genes of SC.M4, SC.M2, SC.M6 (human) | none | network connectivity/significance (1000 permutations) | — |
| genetic screen re-analysis (toxicity modifiers) | yeast (C9orf72/FUS toxicity models) | C9orf72 and FUS toxicity genetic screens | enrichment of modifier genes within co-expression modules | — |
| bulk RNA-seq differential expression | human postmortem control vs ALS cervical spinal cord | ALS disease vs control | DEGs (112 up, 70 down) and module eigengene-disease association | — |
| RNA-seq of laser capture microdissected tissue | human ALS postmortem motor neurons primed for degeneration | ALS disease (non-degenerated MNs) vs control | DEGs (1290 up, 310 down) and module enrichment | — |
| RNA-seq | human iPSC-derived motor neurons from ALS8 patients | ALS8 patient-derived vs control iPSC-MNs | DEGs (364 up, 289 down) and module enrichment | — |
| spatial transcriptomic sequencing | SOD1-G93A mouse spinal cord | SOD1-G93A mutation across disease stages (p30, p70, p100, p120) | stage-specific DEGs and module enrichment | — |
- – 13 co-expression modules identified in control spinal cord, grouped into 5 biological themes (morphogenesis/development, epigenetics/gene regulation, neuronal signaling, immune activation, intracellular transport/metabolism/sensing)
- ▲ SC.M2 significantly enriched for ALS common genetic risk via partitioned heritability FDR=0.047
- ▲ SC.M4 significantly enriched for ALS common genetic risk via partitioned heritability FDR=0.041
- ▲ SC.M2 enriched with mature and myelinating oligodendrocyte markers OR=5.3 (FDR<10e-5); OR=2.9 (FDR<0.05)
- ▲ Literature-curated familial ALS risk genes nominally enriched in SC.M4 OR=2.52, P=0.0124 (FDR=0.174, ns)
- ▲ C9orf72/FUS toxicity modifiers significantly enriched in ribosomal module SC.M6, including RPL/RPS genes, TNPO1, NOB1, HABP4
- – Postmortem ALS spinal cord DEGs enriched: upregulated genes in SC.M12 (microglial), downregulated genes in SC.M3 (endothelial); SC.M1 (astrocyte) eigengene upregulated in ALS FDR=5.92e-39 (SC.M12 up); FDR=0.00152 (SC.M3 down)
- – IBD common risk variants enriched in SC.M12; ASD, AD, FTD, PSP show no module enrichment, unlike ALS FDR=0.047 (IBD/SC.M12)
- other OR=2.52, P=0.0124, FDR=0.174 (SC.M4 enrichment for literature-curated large-effect familial ALS risk genes)
- pvalue FDR=0.047 (SC.M2 enrichment for ALS common risk variant heritability)
- pvalue FDR=0.041 (SC.M4 enrichment for ALS common risk variant heritability)
- fold_change OR=5.3, FDR<10e-5 (SC.M2 enrichment for mature oligodendrocyte markers)
- fold_change OR=2.9, FDR<0.05 (SC.M2 enrichment for myelinating oligodendrocyte markers)
- pvalue p=0.001 (SC.M4 PPI); p=0.018 (SC.M2 PPI) (Significance of PPI networks from top 300 hub genes, derived from 1000 permutations)
- count 112 up-regulated and 70 down-regulated DEGs (Postmortem ALS vs control cervical spinal cord differential expression)
- pvalue FDR=5.92e-39 (SC.M12 up); FDR=0.00152 (SC.M3 down) (Enrichment of postmortem ALS DEGs within co-expression modules)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper constructed a weighted gene co-expression network (WGCNA) from RNA-seq data of 62 control human cervical spinal cord samples (GTEx), identifying 13 modules annotated by GO enrichment. Module relevance to ALS was assessed via Fisher's-exact-style enrichment tests for curated rare large-effect ALS genes, stratified LD score regression to partition common SNP heritability across modules, and odds-ratio-based enrichment of DEGs from four external ALS datasets. Results were reported as odds ratios and FDR-corrected p-values throughout; no traditional group-mean comparisons were the primary analysis.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Weighted Gene Co-Expression Network Analysis (WGCNA) | Construction of gene co-expression modules from GTEx control cervical spinal cord RNA-seq | 62 GTEx control cervical spinal cord samples | not stated |
| Enrichment test reported as odds ratio with FDR-corrected p-value (Fisher's exact implied) | Enrichment of literature-curated large-effect ALS risk genes (~50 genes) within each of 13 co-expression modules | ~50 large-effect ALS risk genes | not stated |
| Stratified LD score regression (partitioned heritability enrichment) | Enrichment of common GWAS risk variants for ALS, IBD, ASD, AD, FTD, and PSP across co-expression modules | ALS GWAS summary statistics: n=20,806 cases, n=59,804 controls | not stated |
| GO term enrichment test (FDR-corrected, threshold 0.05) | Biological annotation of all 13 co-expression modules and cell-type marker enrichment | — | not stated |
| Permutation test (1,000 permutations) | Statistical significance of direct and indirect protein–protein interaction (PPI) networks for SC.M4 (p=0.001), SC.M2 (p=0.018), and SC.M6 hub genes | Top 300 hub genes (SC.M4, SC.M2) or top 100 hub genes (SC.M6) | not stated |
| Module preservation analysis (Z summary statistic) | Assessment of whether GTEx-derived modules were preserved in the postmortem ALS/control cervical spinal cord dataset | — | not stated |
| Odds-ratio-based DEG enrichment test with FDR correction | Enrichment of up- and down-regulated DEGs from four ALS-related datasets (postmortem spinal cord, laser-captured motor neurons, iPSC-derived motor neurons, SOD1-G93A mouse at four timepoints) within each co-expression module | 112 up- and 70 down-regulated DEGs (postmortem spinal cord); 1,290 up and 310 down (laser-captured MNs); 364 up and 289 down (iPSC-derived MNs) | not stated |
| Module eigengene linear model (first principal component of module genes regressed on disease status) | Differential eigengene expression between ALS and control postmortem spinal cord samples across all 13 modules | — | not stated |
-
Co-expression networks were constructed using WGCNA on bulk RNA-seq from 62 samples, with cell-type identity inferred post-hoc via marker gene enrichment.↳ Could also: Single-cell or single-nucleus RNA-seq with cell-type-resolved co-expression or regulon inference (e.g., SCENIC) could also be applied to the same biological question. — Bulk-tissue WGCNA aggregates signals across heterogeneous cell populations; single-cell approaches would resolve cell-type-specific co-expression directly without requiring marker-based annotation as a separate step.
-
Common variant enrichment in modules was assessed using stratified LD score regression (S-LDSC) with GWAS summary statistics.↳ Could also: MAGMA gene-set analysis applied to the same GWAS summary statistics would also test for module-level enrichment of common variant associations. — MAGMA and S-LDSC rely on complementary assumptions; MAGMA operates at the gene level and is less sensitive to annotation of LD structure, providing an orthogonal enrichment estimate that could corroborate or qualify the S-LDSC findings.
-
Enrichment of discrete gene lists (rare ALS risk genes, DEGs defined by a significance threshold) within modules was assessed as odds ratios with FDR correction.↳ Could also: Gene set enrichment analysis (GSEA) using a continuous ranking of all genes (e.g., by fold-change or association statistic) could also quantify module-level enrichment. — GSEA avoids dependence on an arbitrary significance threshold to define the DEG list and can detect coordinated, sub-threshold shifts across an entire module, which a binary overlap test may miss.
-
Confounders were regressed out of expression data using a linear model specifying 11 measured covariates before WGCNA.↳ Could also: Probabilistic estimation of expression residuals (PEER factors) or surrogate variable analysis (SVA) could also be used to capture latent technical and biological variation. — PEER and SVA can account for unmeasured or unknown confounders not captured by the 11 specified covariates, potentially reducing residual technical variation that could otherwise structure co-expression modules.
-
Module dysregulation in ALS was assessed separately in each of four external datasets, with results described individually.↳ Could also: A meta-analytic framework combining module eigengene effect estimates (e.g., random-effects meta-analysis) across the four datasets could also summarize overall dysregulation. — Meta-analysis yields a single variance-weighted effect estimate per module and provides formal tests for heterogeneity across datasets, making it easier to distinguish consistent from context-specific dysregulation.
-
PPI network significance was evaluated using 1,000 permutations comparing observed network connectivity to randomly sampled gene sets of equal size.↳ Could also: A larger number of permutations (e.g., 10,000) or reporting of a confidence interval on the permutation-derived p-value would also characterize uncertainty in the network significance estimate. — With 1,000 permutations the minimum resolvable p-value is 0.001; more permutations or a CI on the p-value would clarify how stable the reported network significance is and whether it approaches that boundary.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 33707641 (ALS spinal cord gene co-expression network, WGCNA)
Wang JC, Ramaswami G, Geschwind DH. Gene co-expression network analysis in human spinal cord highlights mechanisms underlying amyotrophic lateral sclerosis susceptibility. Sci Rep 2021. PMCID PMC7970949. DOI 10.1038/s41598-021-85061-4.
Repo: https://github.com/dhglab/ALS-Gene-Network-Manuscript @ b0a869cba2b453d62079c640363165de9cee530c Data accession in brief: sra:SRP064478. (See note on data provenance below.)
What the repo actually ships (self-contained)
code/1_WGCNA.R + bundled inputs + bundled authors' output:
data/GTEX.expr4.RData— QC'd, regressed expression matrixGTEX.expr.reg(genes x samples)metadata/GTEX.meta4.RData—GTEX.meta.totalcovariate/trait table (seqPCs, RIN, Hardy, Age, center, ethnicity, TRISCHD)results/Modules.RData— authors' OUTPUT:geneTree,moduleColors,MEsresults/*.pdf— soft-threshold selection + signed dendrograms (authors' figures)
The network is built on a GTEx spinal cord (cervical c-1) expression matrix, not on raw SRP064478 reads. The repo provides the processed matrix directly, so the WGCNA step is fully reproducible offline without re-aligning SRA. SRP064478 is the upstream raw-read accession; raw realignment/quantification is NOT in the shipped code and is out of scope here (the paper's processed matrix is the in-scope input).
In scope (pipeline-derived, reproducible from shipped code+data)
- pickSoftThreshold scale-free topology table → soft power = 10 (Results).
- adjacency → TOM → hclust(average) gene dendrogram.
- cutreeHybrid (minClusterSize=100, deepSplit=2, cutHeight=0.9999, pamStage=F)
- mergeCloseModules (cutHeight=0.1) → module assignment
moduleColors.
- mergeCloseModules (cutHeight=0.1) → module assignment
- Number of co-expression modules = 13 (Results) — count of distinct non-grey modules.
- Module eigengenes
MEs. - 1:1 self-consistency check: does re-running
1_WGCNA.R(set.seed(88)) regenerate the shippedresults/Modules.RData(same module count, same gene→module assignment)?
Out of scope (not pipeline-reproducible from this repo)
- Raw SRP064478 read alignment / expression quantification (no code shipped; the QC'd matrix is provided directly).
- GO/pathway enrichment of modules (SC.M2/M4/M6 functional labels), ALS GWAS enrichment (MAGMA/stratified-LDSC), cell-type enrichment, rare-variant / TDP-43 modifier analyses — these are downstream of module assignment and rely on external code/data not in this repo. Mapping of authors' color modules → "SC.M1..M13" labels is not given in the shipped objects.
- Hub genes (KPNA3/TMED2/NCOA4; TARDBP/EWSR1) — derived from module + connectivity + the SC.Mx naming, not reproducible without the (unshipped) labeling step.
Pipeline tool
Standard WGCNA (CRAN) run via the authors' 1_WGCNA.R. Per P16 this is a valid
reproduction regardless of who wrote the driver.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.