CASK loss of function differentially regulates neuronal maturation and synaptic function in human induced cortical excitatory neurons.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce, near-1:1 in magnitude with one important caveat. GEO GSE199910 ships processed abundances (kallisto d28, RSEM d7), so DEG counts were reproduced without re-alignment via tximport+DESeq2 (d28) and t-test (d7) on «our HPC». The repo (UMMS-Biocore/dolphinnext) is a generic platform, so this is a P16 standard-tool reproduction on the paper's own data. Day28: strict 'DEG in both KO lines' intersection gives 771 (down-skewed 280/491), NOT the reported 1742; but a UNION criterion ('DEG in either line') gives 1437 with a balanced 702/735 split that matches the reported 906/838 balance and magnitude (~82%). So the reported number is recoverable from the deposited data, but the Methods wording 'commonly dysregulated in both' does not literally match the intersection it describes -- an analysis-specification ambiguity, not fabrication (no fabrication signal; gene-level directions for RELN/STX1B/RAB3A reproduce). Day7 undershoots under strict intersection (217 vs 876); same loosening direction applies. Residual gaps attributable to the underspecified 'min 5 TPM' rule, the unstated DESeq2 design formula, and intersection-vs-union choice. NOT attempted (out of scope): electrophysiology, morphology, immunostaining, synapse counts (wet-lab), and the GO/network enrichment narrative (downstream of the DEG list).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 62assessed: 2026-06-15 ⛓ 48bae82f2897
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates the cell-autonomous, developmental-stage-specific roles of CASK in human cortical excitatory neurons, testing how CASK loss-of-function affects neuronal maturation (outgrowth/morphology) versus later synaptic transmission and network activity.
- ★ Immature (day 7) CASK KO induced neurons show increased dendritic complexity/neurite overgrowth and upregulated gene networks for cell adhesion, neurite outgrowth, and cytoskeletal organization. finding
- ★ Mature (day 28) CASK KO induced neurons show normal neuronal morphology and unchanged synapse (SYP/PSD95) density and volume. finding
- ★ Mature CASK KO induced neurons exhibit defects in neuronal spiking, synchronized network activity (MEA), and decreased sEPSC frequency, indicating a presynaptic functional defect. finding
- ★ CASK mutations regulate a core, shared set of DEGs across independent genetic backgrounds, based on overlap between this dataset and the Becker et al. (2020) patient iPSC-neuron dataset. finding
- ★ TNIK is a major direct CASK protein-protein interaction hub among day 7 DEGs and links CASK LOF to WNT signaling pathway dysregulation. mechanism
- ★ Generation of two independent isogenic CASK KO human embryonic stem cell lines (KO1, KO2) via CRISPR/Cas9 for cortical excitatory neuron (Ngn2-iN) differentiation. resource
- Bulk RNA-seq of day 7 iNs identified 876 shared DEGs (420 up, 456 down) between CASK KO and WT. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq | day 7 human ESC-derived cortical excitatory induced neurons (Ngn2-iN) | CASK KO (CRISPR) | differentially expressed genes (DESeq2) | — |
| qRT-PCR | day 7 Ngn2-iN, total RNA | CASK KO | validation of up/down-regulated DEGs relative to GAPDH | — |
| gene set enrichment / PPI network analysis (GSEA, ToppGene/ToppCluster) | day 7 iN DEG list | CASK KO | GO term enrichment and CASK protein-protein interaction network | ToppGene/ToppCluster |
| neurite outgrowth morphometrics (confocal imaging with SYN-EGFP sparse labeling) | day 7 Ngn2-iN | CASK KO | soma size, total dendritic length, branch points, primary processes | Imaris (Bitplane) |
| neuronal morphology and synaptic puncta immunostaining (SYP, PSD95, MAP2) | day 28 Ngn2-iN co-cultured with mouse glia | CASK KO | dendritic morphology parameters and synaptic puncta density/volume | confocal imaging + Imaris |
| high-density CMOS microelectrode array (MEA) recording | day 21 and day 28 Ngn2-iN co-cultured with mouse glia | CASK KO | spike firing rate, spike amplitude, network burst synchrony/duration | Maxwell Biosystems MEA |
| whole-cell patch-clamp electrophysiology | mature Ngn2-iN | CASK KO | spontaneous excitatory postsynaptic current (sEPSC) frequency | — |
| immunoblot (Western blot) | day 4 Ngn2-iN cells | CASK KO | CASK protein expression (TUJ1 loading control) | anti-CASK antibody |
- – 876 shared DEGs identified in day 7 CASK KO1/KO2 iNs vs WT (420 up-regulated, 456 down-regulated) 876 DEGs (420 up / 456 down)
- ▲ Up-regulated day 7 DEGs enriched for cell projection morphogenesis, neurodevelopment, cell adhesion, and synaptic membrane/structure GO terms
- ▲ Day 7 CASK KO iNs show increased total dendritic length and number of branch points vs WT, with no change in soma size or primary processes
- – Day 28 CASK KO iNs show no significant change in dendritic complexity (except slight soma size decrease in KO1) and no change in SYP/PSD95 puncta density or volume
- ▼ CASK KO iNs show decreased neuronal firing rate and spike amplitude on MEA at days 21 and 28
- ▲ WT iN spike amplitude increased from day 21 to day 28 60 to 90 μV
- ▲ WT iN network burst duration increased from day 21 to day 28 0.8 s to 1.3 s
- – PPI network analysis of day 7 DEGs identified TNIK as a direct CASK interactor and major hub linked to WNT signaling genes 4 direct / 32 indirect interactors
- count 876 DEGs (420 up-regulated, 456 down-regulated) (day 7 CASK KO vs WT bulk RNA-seq)
- fold_change FC ≥ 1.2 or ≤ 0.8, p < 0.05, minimum 5 TPM per gene (DEG cutoff criteria for RNA-seq analysis)
- pvalue 9.60E-07 (cell projection morphogenesis GO term enrichment (up-regulated DEGs))
- pvalue 1.26E-05 (neurodevelopment GO term enrichment (up-regulated DEGs))
- pvalue 5.85E-06 (cell adhesion GO term enrichment (up-regulated DEGs))
- pvalue synapse 3.09E-06; postsynapse 2.01E-08 (GO Cellular Component enrichment (up-regulated DEGs))
- count 4 direct interactors, 32 indirect interactors (CASK PPI network among day 7 DEGs)
- count 4 replicates WT & KO1, 3 replicates KO2 (bulk RNA-seq culture replicates per genotype)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used isogenic CASK knockout hESC lines (two independent KO clones vs. one WT parental line) differentiated into cortical excitatory induced neurons, assayed at day 7 (immature) and day 28 (mature) time points. Differential gene expression was assessed by bulk RNA-seq analyzed with DESeq2, validated by qRT-PCR; morphometric and electrophysiological outcomes were compared between genotypes using Student's t-tests. Results were reported with means ± SEM and asterisk-based significance thresholds.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 Wald test (negative binomial model) | Differential gene expression in day 7 and day 28 bulk RNA-seq (Figures 2A, and implied for day 28); cutoff FC ≥1.2 or ≤0.8, p ≤0.05, minimum 5 TPM | 4 replicates WT, 4 replicates KO1, 3 replicates KO2 (day 7 RNA-seq) | not stated |
| GSEA / ToppGene enrichment test with Bonferroni correction | Gene ontology enrichment of up- and down-regulated DEGs (Figure 2B, Figures S1–S2); threshold p < 0.05 after Bonferroni | 876 shared DEGs (420 up, 456 down) | not stated |
| Student's t-test (two-sample, unpaired implied) | qRT-PCR validation of 25 up-regulated and 12 down-regulated DEGs (Figure 2C, Figure S1B); WT vs. KO1 and WT vs. KO2 separately | 4–5 independent culture samples per genotype; each reaction in triplicate | not stated |
| Student's t-test (two-sample, unpaired implied) | Neurite outgrowth morphometrics at day 7: soma size, total dendritic length, branch points, primary processes (Figure 3B); WT vs. KO1 and WT vs. KO2 separately | Number of cells per independent culture replicate shown in figure bars; replicates not stated numerically in excerpt | not stated |
| Student's t-test (two-sample, unpaired implied) | Day 28 neurite morphometrics and synaptic puncta density/volume (Figures 4B, 4D); WT vs. KO1 and WT vs. KO2 separately; 4 independent culture batches | 4 independent culture batches; cell counts per bar shown in figures | not stated |
-
Each CASK KO line was compared to WT with separate t-tests across multiple morphometric parameters (soma size, total dendritic length, branch points, primary processes) and across ~37 qRT-PCR targets, without a stated correction for the resulting family of comparisons↳ Could also: A linear mixed-effects model or one-way ANOVA with a post-hoc correction (e.g., Tukey HSD or Benjamini-Hochberg FDR) applied across the parameter/gene family would also control the expected rate of false positives within each comparison family — When multiple outcome measures are tested in the same experiment, a correction approach makes the false-discovery rate explicit and allows readers to interpret the ensemble of p-values in context
-
Dispersion was reported as SEM throughout↳ Could also: SD or a 95% confidence interval would also describe the spread of the data — With small n (3–5 biological replicates), SD conveys the actual variability of observations more directly, while 95% CIs communicate estimation uncertainty and are often recommended by reporting guidelines for small samples
-
The two KO lines were each compared to WT independently, and shared DEGs were identified by intersection↳ Could also: A single DESeq2 model with genotype as a multi-level factor (or an interaction model) followed by a contrast testing the common KO effect could also identify genes consistently altered in both KO lines in a single statistical framework — A joint model uses all replicate information simultaneously, can provide a single adjusted p-value for the shared effect, and avoids the implicit multiple-testing involved in taking the intersection of two separate test results
-
The DESeq2 cutoff is stated as 'p ≤ 0.05' without explicitly naming whether the raw or BH-adjusted p-value (padj) was used↳ Could also: Explicitly filtering on padj ≤ 0.05 (or a stated FDR threshold such as 0.1) and reporting it as such would also make the multiple-testing correction strategy unambiguous — DESeq2 reports both raw p and padj; clearly naming the filter allows readers to assess the expected false-discovery rate across the thousands of genes tested
-
The fold-change thresholds (FC ≥1.2 or ≤0.8) were used as a DEG filter alongside the p-value cutoff↳ Could also: Reporting log2 fold change with its standard error (or shrinkage estimate as provided by DESeq2's lfcShrink) alongside the significance threshold would also quantify the magnitude of each gene's differential expression — Shrinkage-based log2 fold change estimates from DESeq2 stabilize estimates for low-count genes and provide a standardized effect size that facilitates comparison across genes and datasets
-
Neuronal morphometric data at both day 7 and day 28 were collected from multiple cells nested within independent culture batches, analyzed with standard t-tests↳ Could also: A mixed-effects model or hierarchical linear model treating culture replicate as a random effect and genotype as a fixed effect would also account for the non-independence of cells measured within the same batch — When multiple cells are measured per culture replicate, cells within the same replicate are not independent; a hierarchical model partitions variance appropriately between the cell and replicate levels, which can affect both the estimated standard error and the resulting p-value
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
1 downstream papers · 2 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Presynaptic dysfunction in CASK-related neurodevelop... 2020 · 33 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36262316 (CASK LoF in human cortical excitatory neurons, iScience 2022)
Paper
McSweeney et al. 2022, iScience 25:105187. PMID 36262316 / PMC9574418. Data: GEO GSE199910 (bulk RNA-seq). "Code": github.com/UMMS-Biocore/dolphinnext (a generic Nextflow RNA-seq platform, NOT authors' bespoke analysis script — per BRIEF rule P16, running the same standard tools on the paper's data is an equally valid reproduction).
Bulk RNA-seq design (from GEO + Methods)
24 samples. Two timepoints x three genotypes x 4 replicates:
- Day 7 cortical neurons: WT, CASK KO#1, CASK KO#2 (4 reps each = 12)
- Day 28 cortical neurons (+ mouse glia): WT, KO#1, KO#2 (4 reps each = 12) GEO supplementary = transcript/gene abundance tarballs: GSE199910_d7_abundance.tar.gz, GSE199910_d28_abundance.tar.gz
IN SCOPE — pipeline-derived DE results
- C1 (Day 28, primary, well-specified): kallisto v0.46.0 (GRCh38 v96 + GRCm38 v96 concatenated) -> DESeq2 v1.28.1. Reported: 1742 DEGs (906 up, 838 down) common to KO1 & KO2 vs WT. Threshold FC>=1.2 or <=0.8, p<=0.05, min 5 TPM. REPRODUCE: deposited d28 kallisto abundances -> tximport -> DESeq2 -> per-KO contrasts vs WT -> intersect same-direction at the stated cutoffs -> count.
- C2 (Day 7, secondary): RSEM v1.2.28 (UCSC hg19 refGene) -> Student's t-test, |log2FC|>1? (text says FC>=1.2/<=0.8, p<0.05, 5 TPM). Reported: 876 DEGs (420 up, 456 down) common to both KO lines. REPRODUCE from deposited d7 abundances (TPM) -> per-gene t-test KO vs WT -> intersect.
OUT OF SCOPE (not pipeline / not attempted)
- Electrophysiology (MEA spiking, synaptic transmission, burst firing) — wet-lab.
- Morphology / immunostaining / synapse counts — wet-lab imaging.
- GO/network enrichment narrative — downstream of the DEG list, not re-derived.
Reproduction approach
Start from deposited abundances (no re-alignment needed — 80/20). Day 28 DESeq2 is the cleanest claim and the primary target. Day 7 t-test is secondary. All compute on «our HPC»/«infra».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Reproduced from the authors' own GEO-deposited abundances (kallisto d28 / RSEM d7), so data identity and endpoint comparability are strong (q1/q2 green). The deviation — day28 1742 vs 771 strict / 1437 union, day7 876 vs 217 — sits not in the DE statistics (gene directions reproduce) but in an authors-side specification problem: the Methods say 'in both KO lines' (intersection) while the reported balanced count only matches a union criterion, plus an undefined '5 TPM' rule and unstated design formula. Values are therefore partly derivable (~82% recovered, no fabrication) and the core biological conclusion holds qualitatively but with limited numeric confirmation. Overall a solid yellow: explainable, data-consistent deviations driven by underspecified/contradictory methods, not irreproducibility or fabrication.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.