Cross-species comparative hippocampal transcriptomics in Alzheimer's disease.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the core pipeline leg. The paper is a public-data re-analysis with NO authors' analysis repo (harvested 'code' = sra-tools, a download tool only); per BRIEF P16 I reproduced by re-implementing the paper's described DESeq2 recipe (mean CPM>=2 filter, NB Wald + lfcShrink, DEGs at unadjusted p<0.05) and running it on the paper's own GEO data for the APP/PS1 mouse model. Result: 1:1-in-spirit, partial on the exact number. The reported 1768 APP/PS1 DEGs is BRACKETED by my reproductions (632 < 1768 < 2263) and the simplest single-dataset run is within +4.7% (1852). The +/- spread is fully explained by two choices the Methods leave unstated: (a) whether the two combined datasets (GSE149661 12mo + GSE145907 8mo) were batch-corrected, (b) gene-ID harmonization (ENSMUSG vs symbol). NOT a fabrication concern — 1768 is reproducible to within method-choice variance from the shipped public data. Strong positive control: the APP/PS1 transgene (App) is the single strongest DEG, confirming the DE direction and pipeline are correct, and the microglial/immune signature matches the paper. NOT attempted (out of 80/20 scope): the 5xFAD & habeta-KI models (syn16798173/syn18634479 — AMP-AD/Synapse credentialed access, data_restricted), the human microarray leg (5 GSE studies + sva + limma), and all downstream steps (STRING PPI, clusterProfiler GO/KEGG, GOSemSim, ARACNe master regulators, GSEA — the hard ~20%, multi-stage with many unpinned params). DESeq2 version used was 1.50.2 (paper: 1.28.1; pinned conda solve failed in budget; NB Wald test stable across versions — documented in agreement.json).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 75assessed: 2026-06-15 ⛓ 800212ee818d
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests to what extent transgenic overexpression (5xFAD, APP/PS1) and humanized knock-in (hAβ-KI) mouse models recapitulate the hippocampal transcriptomic alterations of human early-onset (EOAD) and late-onset (LOAD) Alzheimer's disease.
- ★ All three mouse models (hAβ-KI, 5xFAD, APP/PS1) share more differentially expressed genes, GO biological processes, and KEGG pathways with LOAD than with EOAD patients. finding
- ★ The hAβ-KI knock-in model is the most specific to LOAD, with ~92% of its enriched GOBP terms overlapping LOAD versus ~32% with EOAD. finding
- ★ More biological processes than individual genes are altered/conserved across human AD and mouse models, indicating functional rather than gene-level convergence. finding
- ★ Innate immune response genes (C1QB, CD33, CD14, S100A6, SLC11A1) are consistently upregulated across mouse models and human AD. finding
- ★ 17 transcription factors act as candidate master regulators of AD, with PARK2 and SOX9 enriched in all three models and human disease. finding
- A regulatory network/master-regulator inference approach (with two-tail GSEA) can identify transcription factors driving AD transcriptional changes across species. method
- Mouse models show greater DEG/GOBP/KEGG overlap with AD than with multiple sclerosis, supporting AD-specificity. finding
- Cross-species comparison used unadjusted p<0.05 DEGs for exploratory analysis because adjusted (BH<0.05) DEGs yielded no model-human overlap. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk hippocampal transcriptomics (publicly available datasets, differential expression analysis) | hAβ-KI, 5xFAD, APP/PS1 mice vs WT; EOAD and LOAD vs cognitively unimpaired humans | transgenic overexpression / humanized knock-in / disease vs control | differentially expressed genes (DEGs) | — |
| Functional enrichment analysis (Gene Ontology biological processes) | mouse models and human AD subtypes (hippocampus) | none (in silico) | enriched GOBP terms and semantic-similarity clusters | — |
| Functional enrichment analysis (KEGG canonical pathways) | mouse models and human AD subtypes (hippocampus) | none (in silico) | enriched KEGG pathways and overlap | — |
| Master regulator analysis with two-tail GSEA | mouse models and human AD DEG sets | none (in silico) | enriched transcription factors / activation state | — |
| Protein-protein interaction network analysis | DEGs shared between hAβ-KI and LOAD or EOAD | none (in silico) | network hub genes and clusters | — |
| qRT-PCR validation | APP/PS1 (n=14) and WT (n=14) mouse hippocampus from two laboratories | transgenic overexpression vs WT | mRNA levels of 7 overlapping DEGs | — |
| qRT-PCR validation | postmortem human hippocampus: CU (n=9), EOAD (n=7), LOAD (n=8), Douglas-Bell Canada Brain Bank | disease vs control | mRNA levels of 7 overlapping DEGs | — |
- – DEGs identified in hippocampus of hAβ-KI, 5xFAD, and APP/PS1 mice vs WT 1537, 3231, and 1768 DEGs respectively
- – hAβ-KI mice share more DEGs with LOAD than with EOAD 381 (24.8%) vs 164 (10.7%)
- – ~92% of hAβ-KI enriched GOBP terms overlap LOAD vs ~32% with EOAD (model-disease) 92% vs 32%
- – Six of ten KEGG pathways enriched in hAβ-KI also enriched in LOAD, only one in EOAD 6/10 vs 1/10
- ▲ S100A6, C1QB, CD33, CD14, SLC11A1 mRNA increased in APP/PS1 mice vs WT
- – S100A6 increased in both LOAD and EOAD; SLC11A1 increased and KCNK decreased in LOAD humans
- – Only PARK2 and SOX9 master regulators enriched in all three models and human disease (95 MR total, 17 in ≥4 groups) 2 of 17
- – hAβ-KI DEG overlap with EOAD (10.7%) not different from MS (9.6%), but LOAD overlap (24.8%) significantly higher 24.8% vs 10.7% vs 9.6%
- count 1537, 3231, 1768 DEGs (DEGs in hAβ-KI, 5xFAD, APP/PS1 vs WT (unadjusted p<0.05))
- pvalue Chi-square adjusted p<0.001 (hAβ-KI shares more DEGs with 5xFAD (25.3%) than APP/PS1 (15.3%))
- pvalue adjusted p=0.0054 (5xFAD-EOAD overlap (14%) vs hAβ-KI-EOAD overlap (10.7%))
- pvalue adjusted p=0.063 (APP/PS1-LOAD (25.3%) vs hAβ-KI-LOAD (24.8%) overlap, not significant)
- pvalue adjusted p=0.017 (glutamatergic/GABAergic synapse KEGG enriched in LOAD)
- count 212 vs 98 DEGs (hAβ-KI DEGs exclusively shared with LOAD vs EOAD)
- count 7 DEGs (DEGs common to all mouse models and human AD)
- count n=14 APP/PS1, n=14 WT; n=9 CU, n=7 EOAD, n=8 LOAD (qRT-PCR validation sample sizes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This cross-species comparative study used publicly available hippocampal RNA-seq transcriptomic datasets from three AD mouse models (5xFAD, APP/PS1, hAβ-KI) and human EOAD and LOAD patients to identify overlapping differentially expressed genes (DEGs), enriched Gene Ontology biological processes (GOBPs), and KEGG pathways. DEGs were defined at an unadjusted p < 0.05 threshold (with sensitivity analyses at BH-FDR < 0.1 and < 0.05), and pairwise overlap proportions were compared using Pearson's chi-squared tests with Yates' continuity correction. Experimental validation of candidate genes was performed by qRT-PCR, with group comparisons on Z-standardized expression values using the Wilcoxon test. Master regulator inference and two-tailed GSEA were additionally applied to characterize transcriptional drivers.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson's chi-squared test with Yates' continuity correction | Mosaic plot comparisons of DEG, GOBP, and KEGG overlap proportions among mouse models, EOAD, LOAD, and MS | — | not stated |
| Wilcoxon rank-sum test | qRT-PCR validation: Z-score comparisons of APP/PS1 vs WT, EOAD vs CU, LOAD vs CU | APP/PS1 n=14, WT n=14, CU n=9, EOAD n=7, LOAD n=8 | not stated |
| Functional enrichment analysis (GO biological processes and KEGG pathways, specific test not stated — hypergeometric assumed) | Enrichment of DEGs in GOBP terms and KEGG canonical pathways for each model and AD subtype | — | not stated |
| Two-tailed gene set enrichment analysis (GSEA) | Inference of activation state of master regulator transcription factors | — | not stated |
| Master regulator analysis (specific algorithm not stated in excerpt — VIPER or equivalent assumed) | Identification of transcription factors potentially driving transcriptional alterations in AD and mouse models | — | not stated |
-
DEGs were defined using unadjusted p < 0.05 for the primary cross-species overlap analysis because no overlaps survived BH-FDR < 0.05↳ Could also: A ranked, threshold-free approach such as GSEA on continuous DE statistics (e.g., log-fold-change or Wald statistic), or a permissive but explicit FDR threshold (e.g., BH < 0.20) with clear labeling as exploratory, could also be used — Threshold-free ranking methods capture the full expression gradient without requiring a binary cutoff and can identify coherent pathway signals even when individual genes do not survive stringent correction, making the exploratory intent more formally grounded
-
Pairwise overlap proportions between groups were compared with Pearson's chi-squared tests with Yates' continuity correction↳ Could also: Fisher's exact test could also be applied for overlap proportion comparisons, particularly in cells with small expected counts — Fisher's exact test does not rely on asymptotic approximations and is preferred when any expected cell count is below ~5, a situation that can arise when comparing small overlap sets
-
qRT-PCR validation comparisons were performed on Z-standardized expression values using the Wilcoxon rank-sum test↳ Could also: A one-way Kruskal-Wallis test followed by Dunn's post-hoc test (or a one-way ANOVA with Tukey HSD if normality holds) could also compare all three human groups (CU, EOAD, LOAD) simultaneously before pairwise contrasts — A single omnibus test before pairwise comparisons controls the family-wise error rate across the three-group comparison and makes the multiplicity structure explicit, whereas running separate two-group Wilcoxon tests increases the chance of a type I error across the family
-
Similarity between mouse models and human AD subtypes was quantified as the proportion of overlapping DEGs (count-based intersection/union fractions)↳ Could also: The Jaccard index, Szymkiewicz-Simpson overlap coefficient, or rank-based correlation of effect sizes (e.g., Spearman's rho on log-fold-changes genome-wide) could also quantify cross-species concordance — Effect-size-based correlations use the full continuous spectrum of differential expression rather than a binary DEG membership, and the overlap coefficient explicitly accounts for asymmetric list sizes, each providing a complementary perspective on model-disease similarity
-
GO biological process and KEGG pathway enrichment was performed on binary DEG lists (standard over-representation analysis)↳ Could also: Gene set enrichment analysis (GSEA) on pre-ranked gene lists (by fold-change or test statistic) or fgsea could also be applied to avoid dependence on a DEG significance threshold — Ranked-list enrichment methods make use of the full expression profile and do not require a binary cutoff decision, which is particularly relevant here given that threshold choice had a large effect on the number of overlapping DEGs
-
Semantic similarity of GO terms was computed and clusters were named manually based on biological role↳ Could also: Automated GO term summarization tools such as REVIGO or rrvgo could also be used to collapse redundant terms and generate cluster labels algorithmically — Automated reduction tools provide a reproducible, parameter-documented alternative to manual labeling, and their distance-based outputs can be directly visualized as treemaps or scatter plots, facilitating comparison across groups
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38292167
Paper: De Bastiani MA, Bellaver B, Carello-Collar G, et al. Cross-species comparative hippocampal transcriptomics in Alzheimer's disease. iScience 2023; 27(1):108671. PMID 38292167 · PMCID PMC10824791 · DOI 10.1016/j.isci.2023.108671.
What kind of paper this is
A secondary / re-analysis study. The authors did not generate new sequencing data; they curated public mouse-model RNA-seq and human-AD microarray datasets and ran a standard bioinformatics pipeline (differential expression → overlap → enrichment → network/master-regulator inference) to compare AD signatures across species and models.
The "Code" field harvested into this RU (github.com/ncbi/sra-tools) is only the
NCBI SRA download utility — there is no authors' analysis repo. Per BRIEF P16
this is fine: we reproduce by running the paper's described pipeline (DESeq2) on
the paper's own input data, which is an equally valid reproduction.
Datasets used by the paper (from Methods + Table references)
| Accession | Species / model | Assay | Role |
|---|---|---|---|
| GSE149661 | mouse APP/PS1, 12 mo, hippocampus | RNA-seq (HTSeq counts shipped) | assigned to this RU |
| GSE145907 | mouse APP/PS1, 8 mo, hippocampus | RNA-seq | combined w/ GSE149661 for APP/PS1 |
| syn16798173 | mouse 5xFAD | RNA-seq (AMP-AD) | other model |
| syn18634479 | mouse hAβ-KI | RNA-seq (AMP-AD) | other model |
| GSE28146/29378/36980/48350/84422 | human AD/CU | microarray | human side |
| GSE123496 | human MS | RNA-seq | specificity control |
In scope (this RU) — pipeline-derived, low ambiguity
Pipeline: GEO processed counts → DESeq2 v1.28.1 (Negative-Binomial,
lfcShrink) → filter genes with mean CPM < 2 → DEGs = unadjusted p < 0.05
(exact recipe quoted from Methods).
- R1 (core, fully specified): Apply that exact DESeq2 recipe to GSE149661 (the assigned accession) — contrast APP-PS1+PBS (n=5) vs WT+PBS (n=5) (the anti-CD8 treatment arm is excluded; it is the original GSE149661 study's intervention, not a genotype contrast). Report the DEG count and top genes.
- R2 (faithful 1:1 attempt, if GSE145907 ships raw counts): Combine GSE149661 (12 mo) + GSE145907 (8 mo) to the paper's n=8 APP/PS1; 8 WT and rerun, to compare directly against the reported 1768 DEGs for the APP/PS1 model.
Sample selection (reconciled with paper's "8 and 12 months-old, n=8 APP/PS1; 8 WT")
- APP/PS1 (8): GSE149661 PBS [GSM4508482,85,87,89,91] (12mo) + GSE145907 [GSM4339185-7] (8mo)
- WT (8): GSE149661 PBS [GSM4508483,84,92,93,95] (12mo) + GSE145907 [GSM4339182-4] (8mo)
- Excluded: GSE149661 anti-CD8 arm [GSM4508481,86,88,90,94].
Out of scope (not attempted, with reason)
- Other models (5xFAD, hAβ-KI): AMP-AD Synapse accessions (syn…) require a
registered/credentialed account →
data_restrictedfor this RU's effort budget. - Human microarray side, sva batch correction, 5-study merge — separate pipeline + many ambiguous curation choices; the 80/20 core is the mouse DESeq2 leg.
- Downstream PPI/STRING, clusterProfiler GO/KEGG, ARACNe master regulators, GOSemSim — multi-stage, many unpinned parameters → the "hard 20%", deferred.
- qRT-PCR validation, wet-lab — non-pipeline.
Expected reported value to compare against
- APP/PS1 model: 1768 DEGs (unadjusted p<0.05), Methods/Table S1.
- Honest caveat baked in: 1768 is the combined GSE149661+GSE145907 number. R1 (GSE149661 alone, n=5/5) is a component reproduction → expect "partial". R2 (n=8/8) is the direct 1:1 — feasibility depends on GSE145907 raw-count format.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a public-data re-analysis with no authors' analysis repo, reproduced by re-running the paper's described DESeq2 recipe on its own GEO data. The headline 1768 APP/PS1 DEGs is bracketed by the reproductions (632 < 1768 < 2263, single-dataset within +4.7% at 1852), and the qualitative core claim is fully confirmed — App is the #1 DEG and the microglial/immune (DAM) signature matches. The deviation sits in input/preprocessing (unstated batch correction across GSE149661+GSE145907 and ENSMUSG-vs-symbol harmonization), i.e. a mix of authors' Methods underspecification and our self-chosen fill-in, not a computational error. No fabrication concern; moderate, explainable count variance with the central conclusion intact → solid partial.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.