Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Gene co-expression network analysis in human spinal cord highlights mechanisms underlying amyotrophic lateral sclerosis susceptibility.

Sci Rep · 2021
L1 100/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
  • No authors-side cause for any deviation
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

EXACT 1:1 reproduction of the in-scope pipeline result. The repo dhglab/ALS-Gene-Network-Manuscript @ b0a869c is fully self-contained: it ships the processed GTEx spinal-cord regressed expression matrix (GTEX.expr4.RData, 14577 genes x 62 samples), the covariate metadata, and the authors' output (results/Modules.RData). Re-running the module-defining chain of code/1_WGCNA.R on «our HPC» with WGCNA 1.67 (the authors' exact version), dynamicTreeCut 1.63-1, fastcluster 1.1.25 (seed 88, pickSoftThreshold, adjacency power=10, TOM, hclust average, cutreeHybrid mms=100/ds=2/cutHeight=0.9999/pamStage=F, mergeCloseModules dcor=0.1) regenerates the shipped Modules.RData EXACTLY: 13 non-grey modules (paper reports 13), every one of the 14577 genes receives the identical module color, fraction-identical = 1.000 and adjusted Rand index = 1.000, and the sample/gene counts (62/14577) and soft power (10) match the paper. All 8 in-scope claims grade exact. Described well enough to reproduce 1:1. NOT attempted (out of scope, not in repo): raw SRP064478 realignment, downstream GO/GWAS/cell-type enrichment, SC.Mx module labels, and hub-gene calls.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ dd895b5f2d0e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Gene co-expression network analysis of the human spinal cord can provide a systems-level, unbiased framework to identify the biological processes and cell types through which ALS genetic risk factors act, thereby revealing causal mechanisms of ALS pathophysiology.

Core claims
  • WGCNA on control human cervical spinal cord RNA-seq identifies 13 co-expression modules (SC.M1-M13), each representing distinct biological processes or cell types. finding
  • ALS common genetic risk (via partitioned SNP heritability) is significantly enriched in module SC.M4 (RNA processing/gene regulation) and SC.M2 (intracellular transport/autophagy, oligodendrocyte-enriched). finding
  • Top hub genes of SC.M2 include ALS-implicated genes KPNA3, TMED2, and NCOA4, the latter regulating ferritin autophagy, implicating this process in ALS pathophysiology. mechanism
  • Genetic modifiers of C9orf72 and FUS toxicity from yeast screens are significantly enriched in the ribosomal-function module SC.M6. finding
  • Differentially expressed genes in postmortem ALS spinal cord are enriched in SC.M12 (upregulated, microglial/immune) and SC.M3 (downregulated, endothelial), but these DEGs are not enriched in causal genetic risk, suggesting they are consequences rather than drivers of disease. finding
  • ALS GWAS enrichment in SC.M2/SC.M4 is disease-specific, as IBD risk variants enrich a different module (SC.M12) and ASD, AD, FTD, and PSP risk variants show no module enrichment. finding
  • Literature-curated large-effect familial ALS risk genes are nominally (but not FDR-significantly) enriched in SC.M4. finding
  • 10 of 13 GTEx-derived modules are preserved in an independent postmortem ALS spinal cord dataset, with SC.M2, SC.M4, SC.M5, and SC.M8 strongly preserved. method
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq co-expression network construction (WGCNA) human control cervical spinal cord (GTEx, n=62) none gene co-expression modules, module eigengenes
stratified LD score regression / partitioned heritability ALS GWAS summary statistics (ALS n=20,806; CTL n=59,804) none SNP-heritability enrichment per co-expression module
protein-protein interaction (PPI) network analysis with permutation testing top hub genes of SC.M4, SC.M2, SC.M6 (human) none network connectivity/significance (1000 permutations)
genetic screen re-analysis (toxicity modifiers) yeast (C9orf72/FUS toxicity models) C9orf72 and FUS toxicity genetic screens enrichment of modifier genes within co-expression modules
bulk RNA-seq differential expression human postmortem control vs ALS cervical spinal cord ALS disease vs control DEGs (112 up, 70 down) and module eigengene-disease association
RNA-seq of laser capture microdissected tissue human ALS postmortem motor neurons primed for degeneration ALS disease (non-degenerated MNs) vs control DEGs (1290 up, 310 down) and module enrichment
RNA-seq human iPSC-derived motor neurons from ALS8 patients ALS8 patient-derived vs control iPSC-MNs DEGs (364 up, 289 down) and module enrichment
spatial transcriptomic sequencing SOD1-G93A mouse spinal cord SOD1-G93A mutation across disease stages (p30, p70, p100, p120) stage-specific DEGs and module enrichment
Key results
  • 13 co-expression modules identified in control spinal cord, grouped into 5 biological themes (morphogenesis/development, epigenetics/gene regulation, neuronal signaling, immune activation, intracellular transport/metabolism/sensing)
  • SC.M2 significantly enriched for ALS common genetic risk via partitioned heritability FDR=0.047
  • SC.M4 significantly enriched for ALS common genetic risk via partitioned heritability FDR=0.041
  • SC.M2 enriched with mature and myelinating oligodendrocyte markers OR=5.3 (FDR<10e-5); OR=2.9 (FDR<0.05)
  • Literature-curated familial ALS risk genes nominally enriched in SC.M4 OR=2.52, P=0.0124 (FDR=0.174, ns)
  • C9orf72/FUS toxicity modifiers significantly enriched in ribosomal module SC.M6, including RPL/RPS genes, TNPO1, NOB1, HABP4
  • Postmortem ALS spinal cord DEGs enriched: upregulated genes in SC.M12 (microglial), downregulated genes in SC.M3 (endothelial); SC.M1 (astrocyte) eigengene upregulated in ALS FDR=5.92e-39 (SC.M12 up); FDR=0.00152 (SC.M3 down)
  • IBD common risk variants enriched in SC.M12; ASD, AD, FTD, PSP show no module enrichment, unlike ALS FDR=0.047 (IBD/SC.M12)
Key statistics
  • other OR=2.52, P=0.0124, FDR=0.174 (SC.M4 enrichment for literature-curated large-effect familial ALS risk genes)
  • pvalue FDR=0.047 (SC.M2 enrichment for ALS common risk variant heritability)
  • pvalue FDR=0.041 (SC.M4 enrichment for ALS common risk variant heritability)
  • fold_change OR=5.3, FDR<10e-5 (SC.M2 enrichment for mature oligodendrocyte markers)
  • fold_change OR=2.9, FDR<0.05 (SC.M2 enrichment for myelinating oligodendrocyte markers)
  • pvalue p=0.001 (SC.M4 PPI); p=0.018 (SC.M2 PPI) (Significance of PPI networks from top 300 hub genes, derived from 1000 permutations)
  • count 112 up-regulated and 70 down-regulated DEGs (Postmortem ALS vs control cervical spinal cord differential expression)
  • pvalue FDR=5.92e-39 (SC.M12 up); FDR=0.00152 (SC.M3 down) (Enrichment of postmortem ALS DEGs within co-expression modules)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper constructed a weighted gene co-expression network (WGCNA) from RNA-seq data of 62 control human cervical spinal cord samples (GTEx), identifying 13 modules annotated by GO enrichment. Module relevance to ALS was assessed via Fisher's-exact-style enrichment tests for curated rare large-effect ALS genes, stratified LD score regression to partition common SNP heritability across modules, and odds-ratio-based enrichment of DEGs from four external ALS datasets. Results were reported as odds ratios and FDR-corrected p-values throughout; no traditional group-mean comparisons were the primary analysis.

Replicationbiological Sample sizen=62 GTEx control cervical spinal cord samples for network construction; external RNA-seq datasets from previously published studies used as provided; GWAS summary statistics from published ALS GWAS (n=20,806 cases + n=59,804 controls) GroupsALS patients vs. controls in external validation datasets; co-expression modules vs. background gene universe for enrichment analyses Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionFDR correction (specific algorithm not named in the text; Benjamini-Hochberg is standard for this context)
Statistical tests used
Test Applied to n Assumptions
Weighted Gene Co-Expression Network Analysis (WGCNA) Construction of gene co-expression modules from GTEx control cervical spinal cord RNA-seq 62 GTEx control cervical spinal cord samples not stated
Enrichment test reported as odds ratio with FDR-corrected p-value (Fisher's exact implied) Enrichment of literature-curated large-effect ALS risk genes (~50 genes) within each of 13 co-expression modules ~50 large-effect ALS risk genes not stated
Stratified LD score regression (partitioned heritability enrichment) Enrichment of common GWAS risk variants for ALS, IBD, ASD, AD, FTD, and PSP across co-expression modules ALS GWAS summary statistics: n=20,806 cases, n=59,804 controls not stated
GO term enrichment test (FDR-corrected, threshold 0.05) Biological annotation of all 13 co-expression modules and cell-type marker enrichment not stated
Permutation test (1,000 permutations) Statistical significance of direct and indirect protein–protein interaction (PPI) networks for SC.M4 (p=0.001), SC.M2 (p=0.018), and SC.M6 hub genes Top 300 hub genes (SC.M4, SC.M2) or top 100 hub genes (SC.M6) not stated
Module preservation analysis (Z summary statistic) Assessment of whether GTEx-derived modules were preserved in the postmortem ALS/control cervical spinal cord dataset not stated
Odds-ratio-based DEG enrichment test with FDR correction Enrichment of up- and down-regulated DEGs from four ALS-related datasets (postmortem spinal cord, laser-captured motor neurons, iPSC-derived motor neurons, SOD1-G93A mouse at four timepoints) within each co-expression module 112 up- and 70 down-regulated DEGs (postmortem spinal cord); 1,290 up and 310 down (laser-captured MNs); 364 up and 289 down (iPSC-derived MNs) not stated
Module eigengene linear model (first principal component of module genes regressed on disease status) Differential eigengene expression between ALS and control postmortem spinal cord samples across all 13 modules not stated
Approaches that could also have been used
  • Co-expression networks were constructed using WGCNA on bulk RNA-seq from 62 samples, with cell-type identity inferred post-hoc via marker gene enrichment.
    Could also: Single-cell or single-nucleus RNA-seq with cell-type-resolved co-expression or regulon inference (e.g., SCENIC) could also be applied to the same biological question. — Bulk-tissue WGCNA aggregates signals across heterogeneous cell populations; single-cell approaches would resolve cell-type-specific co-expression directly without requiring marker-based annotation as a separate step.
  • Common variant enrichment in modules was assessed using stratified LD score regression (S-LDSC) with GWAS summary statistics.
    Could also: MAGMA gene-set analysis applied to the same GWAS summary statistics would also test for module-level enrichment of common variant associations. — MAGMA and S-LDSC rely on complementary assumptions; MAGMA operates at the gene level and is less sensitive to annotation of LD structure, providing an orthogonal enrichment estimate that could corroborate or qualify the S-LDSC findings.
  • Enrichment of discrete gene lists (rare ALS risk genes, DEGs defined by a significance threshold) within modules was assessed as odds ratios with FDR correction.
    Could also: Gene set enrichment analysis (GSEA) using a continuous ranking of all genes (e.g., by fold-change or association statistic) could also quantify module-level enrichment. — GSEA avoids dependence on an arbitrary significance threshold to define the DEG list and can detect coordinated, sub-threshold shifts across an entire module, which a binary overlap test may miss.
  • Confounders were regressed out of expression data using a linear model specifying 11 measured covariates before WGCNA.
    Could also: Probabilistic estimation of expression residuals (PEER factors) or surrogate variable analysis (SVA) could also be used to capture latent technical and biological variation. — PEER and SVA can account for unmeasured or unknown confounders not captured by the 11 specified covariates, potentially reducing residual technical variation that could otherwise structure co-expression modules.
  • Module dysregulation in ALS was assessed separately in each of four external datasets, with results described individually.
    Could also: A meta-analytic framework combining module eigengene effect estimates (e.g., random-effects meta-analysis) across the four datasets could also summarize overall dysregulation. — Meta-analysis yields a single variance-weighted effect estimate per module and provides formal tests for heterogeneity across datasets, making it easier to distinguish consistent from context-specific dysregulation.
  • PPI network significance was evaluated using 1,000 permutations comparing observed network connectivity to randomly sampled gene sets of equal size.
    Could also: A larger number of permutations (e.g., 10,000) or reporting of a confidence interval on the permutation-derived p-value would also characterize uncertainty in the network significance estimate. — With 1,000 permutations the minimum resolvable p-value is 0.001; more permutations or a CI on the p-value would clarify how stable the reported network significance is and whether it approaches that boundary.
Software: WGCNA (R package) · Stratified LD score regression (ldsc)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 33707641 (ALS spinal cord gene co-expression network, WGCNA)

Wang JC, Ramaswami G, Geschwind DH. Gene co-expression network analysis in human spinal cord highlights mechanisms underlying amyotrophic lateral sclerosis susceptibility. Sci Rep 2021. PMCID PMC7970949. DOI 10.1038/s41598-021-85061-4.

Repo: https://github.com/dhglab/ALS-Gene-Network-Manuscript @ b0a869cba2b453d62079c640363165de9cee530c Data accession in brief: sra:SRP064478. (See note on data provenance below.)

What the repo actually ships (self-contained)

code/1_WGCNA.R + bundled inputs + bundled authors' output:

  • data/GTEX.expr4.RData — QC'd, regressed expression matrix GTEX.expr.reg (genes x samples)
  • metadata/GTEX.meta4.RDataGTEX.meta.total covariate/trait table (seqPCs, RIN, Hardy, Age, center, ethnicity, TRISCHD)
  • results/Modules.RData — authors' OUTPUT: geneTree, moduleColors, MEs
  • results/*.pdf — soft-threshold selection + signed dendrograms (authors' figures)

The network is built on a GTEx spinal cord (cervical c-1) expression matrix, not on raw SRP064478 reads. The repo provides the processed matrix directly, so the WGCNA step is fully reproducible offline without re-aligning SRA. SRP064478 is the upstream raw-read accession; raw realignment/quantification is NOT in the shipped code and is out of scope here (the paper's processed matrix is the in-scope input).

In scope (pipeline-derived, reproducible from shipped code+data)

  1. pickSoftThreshold scale-free topology table → soft power = 10 (Results).
  2. adjacency → TOM → hclust(average) gene dendrogram.
  3. cutreeHybrid (minClusterSize=100, deepSplit=2, cutHeight=0.9999, pamStage=F)
    • mergeCloseModules (cutHeight=0.1) → module assignment moduleColors.
  4. Number of co-expression modules = 13 (Results) — count of distinct non-grey modules.
  5. Module eigengenes MEs.
  6. 1:1 self-consistency check: does re-running 1_WGCNA.R (set.seed(88)) regenerate the shipped results/Modules.RData (same module count, same gene→module assignment)?

Out of scope (not pipeline-reproducible from this repo)

  • Raw SRP064478 read alignment / expression quantification (no code shipped; the QC'd matrix is provided directly).
  • GO/pathway enrichment of modules (SC.M2/M4/M6 functional labels), ALS GWAS enrichment (MAGMA/stratified-LDSC), cell-type enrichment, rare-variant / TDP-43 modifier analyses — these are downstream of module assignment and rely on external code/data not in this repo. Mapping of authors' color modules → "SC.M1..M13" labels is not given in the shipped objects.
  • Hub genes (KPNA3/TMED2/NCOA4; TARDBP/EWSR1) — derived from module + connectivity + the SC.Mx naming, not reproducible without the (unshipped) labeling step.

Pipeline tool

Standard WGCNA (CRAN) run via the authors' 1_WGCNA.R. Per P16 this is a valid reproduction regardless of who wrote the driver.

C1
Reported
62 GTEx spinal-cord samples
Reproduced
62 samples
exact
C2
Reported
14577 genes retained
Reproduced
14577 genes
exact
C3
Reported
soft power = 10
Reproduced
10 (signed R^2 = 0.66 at power 10, first power above the 0.65 cutoff)
exact
C4
Reported
cutreeHybrid minClusterSize = 100
Reproduced
100
exact
C5
Reported
cutreeHybrid deepSplit = 2
Reproduced
2
exact
C6
Reported
mergeCloseModules cutHeight = 0.1
Reproduced
0.1
exact
C7
Reported
13 co-expression modules
Reproduced
13 modules (non-grey)
exact
C8
Reported
shipped Modules.RData gene->module assignment
Reproduced
fraction identical = 1.000, adjusted Rand index = 1.000 over all 14577 genes; every shipped module Jaccard = 1.0
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

165.1 k
tokens (I/O) · 12.9 M incl. cache
38 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.