Prediction of Alzheimer's disease-specific phospholipase c gamma-1 SNV by deep learning-based approach for high-throughput screening.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for a PARTIAL, high-confidence reproduction. The paper's deep-learning core is SpliceAI (Illumina pretrained CNN, third-party tool -> P16-valid, fully deterministic). Running SpliceAI 1.3.1 (GRCh38, bundled GENCODE-V24 annotation, default params) on the human PLCG1 exon-26 variant reproduces the central claim 1:1 qualitatively: DS_AL=0.99 (maximal acceptor-loss, >0.8 high-precision threshold) at the true acceptor-G chr20:41,172,420, biologically coherent (the G is the canonical AG acceptor of exon 26). AUDITABILITY FLAG: the paper's stated coordinates are internally inconsistent and wrong against GRCh38 (41,172,421 AND 41,172,423 are both ref 'A', not 'G'; one SNV cannot be at two positions; no rsID/HGVS) -> resolved to nearest true G. NOT attempted: exact 14 Fig.5 delta values (figure-only, last-20%); the 5xFAD mouse RNA-seq DEGs (GSE151270/GSE147792), H3K27ac ChIP context (GSE52386), and blood/cortex SNV calling (wet-lab primary data, not the DL claim, out of scope).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 48assessed: 2026-06-15 ⛓ 50c02404c521
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests whether combining genome-wide association study (GWAS)/high-throughput RNA-seq data with a deep learning-based exon splicing prediction tool can identify Alzheimer's disease (AD)-specific single-nucleotide variants (SNVs) and abnormal exon splicing in the PLCγ1 gene that may serve as predictive/diagnostic biomarkers for AD.
- ★ An AD-specific frameshift single-nucleotide insertion in exon 27 of PLCγ1 substitutes isoleucine 970 to asparagine in the 5xFAD AD mouse model. finding
- ★ The SNV in exon 27 of PLCγ1 is associated with abnormal exon splicing (exon skipping) during mRNA maturation in AD. finding
- ★ Deep learning (SpliceAI) trained on human genome predicted 14 splicing sites in human PLCγ1, one of which (exon 26) completely matched the SNV position in mouse exon 27. finding
- ★ AD-associated SNVs are concentrated in the H3K27ac-enriched intron regions of PLCγ1, accumulating during forebrain development from E11.5 to P56. finding
- ★ Combining in silico GWAS/RNA-seq analysis with deep learning-based splicing prediction can identify clinically relevant SNVs for AD prediction. method
- The isoleucine site at exon 27 of mouse PLCγ1 is evolutionarily conserved across species including human (exon 26), with full-length amino acid sequence matching completely. finding
- A pipeline for RNA-seq-based SNV and InDel detection (SAMtools, vcfutils, ANNOVAR) profiles genetic variation in PLC subfamily genes. method
- PLC subfamily transcript changes do not produce significant protein-level differences between WT and 5xFAD by Western blot. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| total RNA-seq | cortex of 6-mo-old WT and 5xFAD mice | 5xFAD transgenic (AD model) vs WT | differential gene expression and transcript levels | NGS, paired-end 100, DNA NanoBall (DNB) |
| RNA-seq-based SNV/InDel calling | cortex of WT and 5xFAD mice | 5xFAD vs WT | SNVs and insertions/deletions in PLCγ1 and PLCβ genes | SAMtools mpileup, UCSC mm10 reference, vcfutils.pl varFilter, ANNOVAR |
| Western blot | WT and 5xFAD mouse cortex | 5xFAD vs WT | PLCβ1/β2/β3/β4 and PLCγ1 protein levels | — |
| RT-PCR | WT and 5xFAD mouse cortex (pre-mRNA and mature mRNA) | 5xFAD vs WT | exon inclusion/skipping near exon 27 of PLCγ1 | — |
| ChIP-seq (H3K27ac) | mouse forebrain across development (E11.5 to P56) | developmental stage / none | histone acetylation enrichment profile at PLCγ1 gene body | — |
| deep learning splice prediction (SpliceAI) | human PLCγ1 genomic sequence | in silico single-nucleotide substitution | predicted splice sites with delta scores and positions | SpliceAI deep neural network |
| cross-species sequence alignment | PLCγ1 of various species including human and mouse | none | conservation of isoleucine site / amino acid sequence | UCSC genome browser |
- – Frameshift 'A' insertion at chr2:160,759,682 in exon 27 of PLCγ1 changes Ile970 to Asn in 5xFAD
- – Exons 26-30 of PLCγ1 deleted in mature mRNA of 5xFAD cortex while exons 27-32 conserved in pre-mRNA of both WT and 5xFAD
- – SpliceAI predicted 14 splicing sites in human PLCγ1; exon 26 SNV at chr20:41,172,421 (G>A/C/T) correlated with mouse exon 27
- – 163 total variations (132 SNVs, 31 InDels) across 5 genes; PLCγ1 had 13 variants; most located in introns (152 variants)
- – 1,472 genes up-regulated and 653 genes down-regulated in 5xFAD cortex
- ▲ PLCβ2 expression up-regulated ~2.4-fold in 5xFAD cortex, though its endogenous level remained low and protein undetected ~2.4-fold
- ▲ Histone acetylation (H3K27ac) gradually accumulated at PLCγ1 introns through forebrain development, highest at P56, overlapping AD SNVs
- – Stop-gain mutation in PLCβ3 (chr2:6,963,413 G>A) creates terminal codon at Gln352 in 5xFAD cortex
- count 17,109 genes (DEGs between WT and 5xFAD) (genes profiled in heatmap/clustering of total RNA-seq)
- count 1,472 up-regulated; 653 down-regulated (DEGs in 5xFAD mouse cortex)
- fold_change ~2.4 times (PLCβ2 up-regulation in 5xFAD cortex)
- count 163 variations (132 SNVs, 31 InDels) (total variants identified across five PLC genes)
- count 152 intronic, 5 exonic, 6 splicing-region variants (location distribution of identified variations)
- count PLCγ1=13, PLCβ1=66, PLCβ2=6, PLCβ3=3, PLCβ4=75 variants (variant counts per gene in 5xFAD cortex)
- count 14 splicing sites (SpliceAI-predicted splicing sites in human PLCγ1)
- other human >95-100% vs mouse ~63% alternative pre-mRNA splicing rate (species difference in average AS rate)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combined high-throughput total RNA-seq (WT vs 5xFAD mouse cortex) with a bioinformatics variant-calling pipeline (SAMtools/vcfutils/ANNOVAR) to identify AD-specific SNVs and InDels in PLCγ1 and related PLC genes, and ChIP-seq to profile H3K27ac histone acetylation during forebrain development. Differential gene expression was reported as DEG counts and fold-changes without naming the underlying statistical test or correction method. Wet-lab validation (Western blot, RT-PCR) was described qualitatively, and splice-site prediction used the SpliceAI deep learning tool with delta scores as the decision criterion.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential gene expression analysis (specific test not named) | Genome-wide comparison of WT vs 5xFAD cortex yielding 17,109 DEGs, 1,472 up-regulated, 653 down-regulated (Fig. 1B–C) | — | not stated |
| SNV/InDel variant calling pipeline (SAMtools mpileup + vcfutils.pl varFilter + ANNOVAR annotation) | Identification of 163 variants across PLCγ1, PLCβ1–4 in 5xFAD vs WT cortex (Fig. 2, Fig. 3, SI Appendix Tables S1–S5) | — | not stated |
| SpliceAI deep learning prediction (delta score criterion) | Prediction of 14 splicing sites in human PLCγ1 gene and correlation with mouse AD-model SNV (Fig. 5) | — | na |
| Western blot comparison (specific statistical test not named) | PLC subfamily protein levels in WT vs 5xFAD cortex, stated as 'no significant differences' (SI Appendix Fig. S3) | — | not stated |
| ChIP-seq enrichment profiling (specific statistical test not named) | H3K27ac accumulation across forebrain developmental stages E11.5 to P56 at PLCγ1 locus (Fig. 4) | — | not stated |
-
The differential gene expression method is not named; DEG counts and fold-changes are reported without specifying the statistical model or normalization↳ Could also: Explicitly named tools such as DESeq2 (negative binomial Wald test), edgeR (exact test or GLM), or limma-voom could also be used and reported by name — Naming the method, its distributional assumptions, and normalization strategy (e.g., TMM, RLE) allows readers to assess model fit and reproduce the DEG list
-
No multiplicity correction is described for the genome-wide DEG analysis across ~17,000+ genes↳ Could also: Benjamini-Hochberg FDR correction applied to the full set of tested genes, with an explicit FDR threshold (e.g., FDR < 0.05), is a standard approach in RNA-seq workflows — Stating the correction method and significance threshold contextualizes the DEG count and allows readers to interpret the false-discovery rate among reported hits
-
Sample size (number of biological replicates per group) is not reported↳ Could also: An explicit statement of n per group and, where feasible, a power calculation or effect-size estimate could also be included — Reporting biological n informs the reader about the study's capacity to detect the observed expression differences and supports reproducibility
-
Western blot group comparisons are stated as significant or non-significant without a named test or reported statistic↳ Could also: A Student's t-test or Mann-Whitney U with reported test statistic and exact p-value could also be used for pairwise comparisons of protein band intensities — Exact p-values and a measure of effect magnitude allow quantitative interpretation of the difference and its uncertainty
-
No measure of spread is reported for any quantitative outcome (no SD, SEM, or CI)↳ Could also: Mean ± SD or 95% CI could also be used to communicate variability around central estimates in Western blot or expression data — Dispersion measures convey whether observed differences are consistent across replicates and are especially informative when n is small
-
SpliceAI splice-site predictions are reported without benchmarking against an independent prediction tool↳ Could also: Cross-validation with an orthogonal splice predictor (e.g., MaxEntScan, SPIDEX, or ESEfinder) could also be used to assess concordance — Agreement across independent algorithms with different architectures increases confidence that reported splice sites are not specific to one tool's scoring scheme
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
H3K27ac enrichment at PLCG1 introns increases progressively during mouse forebrain development (E11.5–P56), peaking at P56 and overlapping AD-associated SNV positions.ChIP-seq mouse forebrain up 2021×1papers★ This paper is the founder (earliest)
-
SpliceAI predicts 14 splice sites in human PLCG1; exon 26 SNV at chr20:41172421 (G>A/C/T) is the human equivalent of the AD-associated mouse exon 27 frameshift locus.other human 2021×1papers★ This paper is the founder (earliest)
-
PLCG1 mature mRNA lacks exons 26-30 in 5xFAD mouse cortex while pre-mRNA exons 27-32 are retained in both WT and 5xFAD, indicating aberrant splicing downstream of the frameshift.qPCR mouse cortex 2021×1papers★ This paper is the founder (earliest)
-
PLCB2 mRNA is upregulated ~2.4-fold in 5xFAD mouse cortex relative to WT, though protein remains undetectable by western blot.RNA-seq mouse cortex up 2021×1papers★ This paper is the founder (earliest)
-
PLCB3 harbors a stop-gain SNV (Gln352*) in 5xFAD mouse cortex identified by RNA-seq-based variant calling.RNA-seq mouse cortex 2021×1papers★ This paper is the founder (earliest)
-
PLCG1 harbors a frameshift 'A' insertion in exon 27 (Ile970Asn) in 5xFAD mouse cortex detected by RNA-seq-based variant calling.RNA-seq mouse cortex 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
25 downstream papers · 3 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Tau promotes neurodegeneration through global chroma... 2014 · 400 cites
- Meta-Analysis of Gene Expression Changes in the Bloo... 2019 · 35 cites
- Exploring the Key Genes and Identification of Potent... 2021 · 33 cites
- Differential Expression of mRNAs in the Brain Tissue... 2019 · 30 cites
- Revelation of Pivotal Genes Pertinent to Alzheimer's... 2022 · 22 cites
- Detecting Diagnostic Biomarkers of Alzheimer's Disea... 2019 · 20 cites
- Rapid and pervasive changes in genome-wide enhancer... 2013 · 329 cites
- Genome-wide identification and characterization of f... 2014 · 250 cites
- Actin cytoskeletal remodeling with protrusion format... 2015 · 188 cites
- Hippo Signaling Plays an Essential Role in Cell Stat... 2018 · 163 cites
- A discrete transition zone organizes the topological... 2015 · 53 cites
- Reprogramming of DNA methylation at NEUROD2-bound se... 2019 · 40 cites
- Identifying New COVID-19 Receptor Neuropilin-1 in Se... 2021 · 28 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33397809
Paper: Kim SH, Yang S, Lim KH, Ko E, Jang HJ, Kang M, Suh PG, Joo JY. "Prediction of Alzheimer's disease-specific phospholipase C-γ1 SNV by deep learning-based approach for high-throughput screening." PNAS 2021;118(3):e2011250118. PMID 33397809 · PMCID PMC7826347 · DOI 10.1073/pnas.2011250118.
What the paper does (overview)
The study identifies an Alzheimer's-disease (AD)-specific single-nucleotide variant (SNV) in PLCG1 (phospholipase C-γ1). Mouse (5xFAD model) and human analyses are combined. The computational / deep-learning component is the use of SpliceAI (Illumina; a residual convolutional neural network trained on GENCODE V24) to predict whether the PLCG1 nucleotide alteration creates/affects an RNA splice site. Figure 5 reports "the total 14 accurate prediction splicing sites in the human PLCG1 gene, with the delta scores and position information."
In scope (pipeline-derived → attempted)
| result | pipeline / tool | reproducible? |
|---|---|---|
| R1 SpliceAI splice-altering delta scores for the human PLCG1 AD-SNV (Fig 5) | SpliceAI 1.x (pretrained, GRCh38/GENCODE V24) — third-party tool, run on the paper's variant | YES — deterministic given variant + reference + bundled annotation. This is the primary reproduction target. |
SpliceAI is a third-party GitHub tool (github.com/Illumina/SpliceAI). Per BRIEF P16, applying it to the paper's own variant is an equally valid reproduction. SpliceAI output is fully deterministic (no training, pretrained model), so the delta scores are exactly recomputable.
Out of scope (NOT attempted, with reason)
| result | why out of scope |
|---|---|
| 5xFAD mouse cortex RNA-seq DEGs ("17,109 DEGs WT vs 5xFAD"; GSE151270 / GSE147792) | wet-lab-generated primary data + standard DEG pipeline; not the paper's deep-learning claim; large, peripheral to the SpliceAI result. 80/20: skipped. |
| H3K27ac ChIP-seq during brain development (re-using GEO GSE52386) | GSE52386 is adapted external ChIP-seq context data, not a deep-learning result; descriptive. The auto-mined data:GSE52386 accession belongs to this peripheral analysis, not to the SpliceAI result. Not attempted. |
| Mouse Plcg1 exon-27 single-"A" insertion frameshift (Ile970→Asn) | derived from wet-lab Sanger/variant calling on 5xFAD tissue, described manually; not a runnable pipeline output. Out of scope. |
| Blood vs cortex AD-specific PLCβ/PLCγ SNVs (Suppl. Table S6) | variant calling on the authors' private sequencing; data not the focus and not the DL claim. |
| Wet-lab validations (Western blot, IHC, electrophysiology, behaviour) | non-computational. |
Key auditability flag (HARD RULE 5)
The paper's stated human variant coordinates are internally inconsistent:
"The single nucleotide 'G' at position 41,172,421 in chromosome 20 was substituted with 'A,' 'C,' or 'T' at position 41,172,423."
- A single G→{A,C,T} substitution cannot be at two different positions (41,172,421 and 41,172,423).
- In GRCh38, chr20:41,172,421 = A and chr20:41,172,423 = A — neither is
G. The nearest reference G bases are chr20:41,172,420 and chr20:41,172,422
(verified via UCSC hg38: chr20:41,172,415–41,172,430 =
GCCCAGAGATTGGCAC). - The paper gives no rsID / HGVS, so the exact intended variant is not unambiguously recoverable from the text.
Mitigation: we run SpliceAI on the two true flanking G positions (chr20:41,172,420 and 41,172,422), each as G→A, G→C, G→T (6 variants), GRCh38, bundled GENCODE-V24 annotation. We report the delta scores and compare the qualitative claim ("this SNV creates a novel splice site in exon 26 of PLCG1"). The exact 14 numeric values of Fig 5 are not given in the text (figure-only), so a strict 14-value 1:1 match is part of the optional last-20% and is not forced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's computational core is the deterministic, pretrained SpliceAI CNN, and the central qualitative claim — the AD-specific PLCG1 SNV disrupts the exon-26 splice acceptor — is reproduced 1:1 with a maximal DS_AL=0.99 at the true acceptor-G (chr20:41,172,420), biologically coherent and not fabrication-suspect. The deviations are (i) authors'-side coordinate errors (chr20:41,172,421/41,172,423 are both reference 'A', not 'G', and inconsistent for a single SNV), forcing us to self-define the input, and (ii) Fig.5's 14 delta scores are figure-only with unstated enumeration parameters. Net: the core conclusion holds (q7 green), but sloppy/contradictory input reporting on the authors' side (q4 red) plus the unreproducible figure-only values keep the overall judgement at yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.