Electroacupuncture reshapes the microbial co-occurrence networks related to the behavioral and psychological symptoms of dementia in Alzheimer's disease.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce 1:1 from shipped data. PRIMARY (Fig 2I/J differential metabolic pathways): re-ran the authors' own script F2IJ_Differential Pathway.R (metagenomeSeq fitFeatureModel, cumNorm) on the authors' own shipped PICRUSt2/EC_path_abun table (419 MetaCyc pathways x 144 samples). All four named EA-upregulated pathways reproduce in the EA-vs-Normal post-intervention contrast: DHGLUCONATE-PYR-CAT-PWY (adjP=0.019, AND the single significant pathway of 419 in the 6-month contrast, matching Fig 2I exactly), PWY-7295 (adjP=0.0098) and GALLATE-DEGRADATION-II-PWY (adjP=0.0105) significant & up in EA at 9 months; CHLOROPHYLL-SYN is up in EA but metagenomeSeq returns adjP=NA so significance is unconfirmable (partial). SECONDARY (Fig S7 Pearson correlations): 3 of 4 reported r-values reproduce within ~0.03 once the figure's stated log2-z transform is applied (e.g. -0.464 vs reported -0.46), but the WT/APP group labels are not keyed in the shipped N/S/EA tables so the exact subsets are inferred, and WT6 r=-0.64 is not regenerable (WT cohort not in the shipped table). NOT attempted: raw 16S->OTU/PICRUSt2 (upstream wet-pipeline; processed tables shipped, so reproduced downstream of them); Fig 1 LEfSe LDA/cladogram (un-shipped intermediates); Fig 2K ZiPi (sources a private vendor script); volcano/heatmap/boxplot/OPLS-DA/network-rendering Rmd (read un-shipped Desktop spreadsheets, plotting-only); Fig S5/S6 OFT (wet-lab behaviour). No reproduced value contradicts the paper; no fabrication indicated.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 77assessed: 2026-06-16 ⛓ 47f0a104be6b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusElectroacupuncture ameliorates the behavioral and psychological symptoms of dementia (BPSD) in Alzheimer's disease by reshaping the gut microbial co-occurrence networks, driving keystone species, and regulating core microbiota composition and predictive metabolic pathways.
- ★ Electroacupuncture reshapes microbial co-occurrence network topology and drives keystone species in AD-related BPSD, with R. gnavus emerging as a likely keystone species post-intervention. finding
- ★ Electroacupuncture ameliorates AD-related BPSD-like behavioral phenotypes (e.g., hyperactivity, anxiety) measured by the open field test. finding
- ★ Gut microbial keystone species and composition vary in a stage-specific (age-stratified) manner across the AD continuum, yielding distinct disease-discriminatory microbial markers for early-to-middle versus middle-to-late stage AD. finding
- ★ Electroacupuncture regulates predictive functional metabolic pathways, including glucose, L-arabinose, and gallate degradation and chlorophyllide a biosynthesis, in AD-related BPSD. mechanism
- A Random Forest machine learning model identifies optimal microbial markers discriminating AD-related BPSD by age stage. method
- 16S rRNA sequencing data deposited at NCBI BioProject PRJNA1201107 serves as a resource for the study. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| 16S rRNA gene sequencing | APPswe/PS1δE9 (APP/PS1) and wild-type mice, 6- and 9-month-old (fecal/gut microbiota) | none (disease model vs WT) | gut microbiota composition / relative abundance of bacterial taxa | — |
| 16S rRNA gene sequencing | APP/PS1 mice, 6- and 9-month-old | electroacupuncture (vs Sham and Normal control) | post-intervention microbial composition, co-occurrence network topology, keystone species | HANS-JS502A electroacupuncture device |
| Open field test (OFT, behavioral assay) | APP/PS1 and WT mice, 6- and 9-month-old | electroacupuncture vs Sham vs Normal control | total distance traveled, central zone crossing frequency, distance/duration in central zone | — |
| PICRUSt2 functional prediction | gut microbiota of 6- and 9-month-old APP/PS1 mice | electroacupuncture | predicted metabolic pathway abundance (log2FC) | PICRUSt2 |
| Random Forest machine learning classification | gut microbiota of APP/PS1 vs WT mice | none | importance scores of microbial markers for AD-related BPSD | — |
| LEfSe / Linear discriminant analysis (LDA) | gut microbiota of 6- and 9-month-old APP/PS1 vs WT mice | none | differentially abundant taxa (LDA score) | — |
| Co-occurrence network analysis (Pearson correlation, ZiPi) | gut microbiota of APP/PS1 and WT mice, 6- and 9-month-old | none / electroacupuncture | centrality measures (degree, closeness, betweenness, expected influence), keystone species, module connectivity | — |
- ▲ Bifidobacterium pseudolongum, Lactobacillus hamsteri, and Parabacteroides distasonis were significantly enriched in both 6- and 9-month-old APP/PS1 mice p_adj < 0.05
- ▼ Akkermansia muciniphila and Clostridium cocleatum were significantly depleted in both 6- and 9-month-old APP/PS1 mice p_adj < 0.05
- ▼ In 6-month-old APP/PS1 mice, Mucispirillum schaedleri and C. perfringens were negatively correlated and identified as keystone species r = -0.48, p_adj < 0.05
- ▲ In 9-month-old APP/PS1 mice, B. pseudolongum and L. hamsteri were positively correlated keystone species r = 0.66, p_adj < 0.05
- ▼ In 9-month-old APP/PS1 mice, B. pullicaecorum and C. celatum exhibited significant negative correlation as keystone species r = -0.46, p_adj < 0.05
- ▼ OFT total distance traveled and central zone crossing frequency were significantly greater in APP/PS1 mice than age-matched WT mice; all such indices decreased significantly in EA groups vs Sham post-intervention p_adj < 0.05
- ▲ Electroacupuncture significantly upregulated the DHGLUCONATE-PYR-CAT-PWY glucose degradation pathway in 6-month-old APP/PS1 mice
- – Electroacupuncture altered PWY-7295 (L-arabinose degradation IV), GALLATE-DEGRADATION-II-PWY, and CHLOROPHYLL-SYN (chlorophyllide a biosynthesis I) pathways in 9-month-old APP/PS1 mice
- correlation r = -0.48 (M. schaedleri vs C. perfringens, 6-month-old APP/PS1 network)
- correlation r = -0.64 (M. schaedleri vs C. perfringens, 6-month-old WT network)
- correlation r = 0.66 (B. pseudolongum vs L. hamsteri, 9-month-old APP/PS1 network)
- correlation r = -0.46 (B. pullicaecorum vs C. celatum, 9-month-old APP/PS1 network)
- pvalue p_adj < 0.05 (significance threshold for differential taxa, behavioral indices, and correlations)
- other LDA threshold = 2 (Log10 scale) (LDA score threshold for differentially abundant taxa)
- other Pi = 2.5; Zi = 0.62 (ZiPi plot thresholds dividing microbes into network hubs, module hubs, connectors, peripherals)
- count top 20 species; top 5 species (20 most abundant species in co-occurrence network; top 5 in ZiPi plot)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This animal study compared gut microbiota profiles between APPswe/PS1δE9 (APP/PS1) transgenic and wild-type (WT) mice at two age points (6 and 9 months) and then evaluated electroacupuncture (EA) effects across Normal, Sham, and EA groups at each age. Differential taxa were identified with LEfSe and visualized via volcano plots and heatmaps; co-occurrence networks were built from Pearson correlation coefficients; Random Forest ranked microbial marker importance; and PICRUSt2 predicted functional metabolic pathways. All significance thresholds were expressed as p_adj < 0.05, with LDA scores, Pearson r, and log2 fold-change serving as the primary effect measures.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| LEfSe (Linear discriminant analysis Effect Size) with LDA threshold = 2 (log10 scale) | Differential abundance of taxa between APP/PS1 and WT mice at 6 and 9 months (Figures 1A–B; Tables S1–S4) | — | not stated |
| Pearson correlation coefficient on Z-score-standardized log2-transformed relative abundances | Microbial co-occurrence network construction for all groups and age points (Figures 2E–H, S7) | — | not stated |
| Random Forest (ensemble decision-tree importance scoring) | Ranking of candidate microbial markers for BPSD in early-to-middle and middle-to-late AD stages (Figures S2C–D) | — | na |
| OPLS-DA (Orthogonal partial least squares-discriminant analysis) | Multivariate clustering of microbial composition across all groups post-intervention (Figure S4) | — | not stated |
| PICRUSt2 functional pathway prediction (log2FC reported) | Metabolic pathway differences between EA and control groups in 6- and 9-month APP/PS1 mice (Figures 2I–J) | — | na |
| Unspecified inferential test (p_adj < 0.05 reported; specific test in supplementary Table S5) | Open field test (OFT) behavioral comparisons between APP/PS1 and WT mice at baseline, and between EA, Sham, and Normal groups post-intervention (Figures S5–S6) | — | not stated |
-
Co-occurrence networks were constructed using Pearson correlation coefficients applied to Z-score-standardized, log2-transformed relative abundances↳ Could also: Spearman rank correlation or SparCC (Sparse Correlations for Compositional data) could also be used to build co-occurrence networks from 16S relative-abundance tables — Spearman is more robust to non-normality and outliers common in microbiome data; SparCC was specifically designed to address the compositionality constraint of relative abundances, where proportions are not independent across taxa and can inflate or deflate Pearson coefficients in ways that are difficult to anticipate
-
Differential taxa were identified using LEfSe with a fixed LDA score threshold↳ Could also: DESeq2, edgeR, or ANCOM-BC could also be applied to 16S count data to identify differentially abundant taxa — These methods model integer count data directly with negative-binomial or related distributions and provide taxon-level FDR control with explicit statistical testing; they are widely used when the priority is formal inference alongside discovery
-
Microbial marker importance was evaluated with a Random Forest importance score↳ Could also: LASSO or elastic-net penalized logistic regression could also be used for microbial biomarker selection — Penalized regression yields explicit feature selection with direct probabilistic interpretability, and stability-selection wrappers provide confidence estimates for included features, which can be useful in small-n animal study contexts where Random Forest importance scores may vary across runs
-
Functional metabolic pathway abundances were inferred from 16S amplicon data using PICRUSt2↳ Could also: Shotgun metagenomic sequencing analyzed with HUMAnN3 or similar pipelines could also provide functional pathway profiles — Direct sequencing of the functional gene content bypasses the reference-database and phylogenetic-placement assumptions inherent in PICRUSt2, whose accuracy depends on how well the community is represented by sequenced reference genomes
-
Open field test behavioral outcomes were compared between groups with adjusted p-values; the specific inferential test is not named in the main text↳ Could also: A two-way ANOVA (factors: genotype/treatment × age) followed by Tukey HSD or Dunnett post-hoc correction could also be used to test OFT indices across all groups simultaneously — A factorial ANOVA tests main effects and their interaction within a single model and controls the family-wise error rate across all pairwise comparisons in one step, making explicit whether age and treatment effects are additive or interactive
-
Results throughout are expressed as threshold-based adjusted p-values (p_adj < 0.05) without exact values, and without confidence intervals for behavioral or abundance comparisons↳ Could also: Reporting exact adjusted p-values alongside effect-size estimates with 95% confidence intervals (e.g., fold-change ± CI for abundance, Cohen's d for behavioral endpoints) would also be standard practice — Exact values allow readers to apply their own evidentiary thresholds and to power future studies; confidence intervals convey both the magnitude and the precision of group differences, which is particularly informative in small-n animal experiments where uncertainty can be high
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41676443
Paper: Su et al. 2025, iMetaOmics. "Electroacupuncture reshapes the microbial
co-occurrence networks related to the behavioral and psychological symptoms of
dementia in Alzheimer's disease." DOI 10.1002/imo2.70035 · PMCID PMC12806058.
Repo: https://github.com/sufuyou2021/IMO-2025-0041 @ commit 7b72fdf0354dd00999eef9c50fe6b70e142480ee
Raw data: SRA PRJNA1201107 (16S rRNA amplicon).
Study: 16S microbiome of APP/PS1 (AD model) vs WT mice at 6- and 9-months; an
electroacupuncture (EA) intervention with Normal (N), Sham (S), EA groups,
baseline (_6/_9) and 3-weeks-post (_6W3/_9W3), 12 mice/group.
What the repo ships (processed, human-auditable)
Data/Gut Microbes 6.xlsx,Gut Microbes 9.xlsx— top-20 species relative abundance, 20 species × 72 samples (N/S/EA × 12 × {baseline, W3}).PICRUSt2/EC_path_abun_unstrat_norm_descrip.xls— 419 MetaCyc pathways × 144 samples (the.xlsfiles are actually TSV). Also KO/EC/COG predictions.- R/Rmd/py scripts that render each figure, plus the authors' rendered
.html.
In scope (pipeline-derived, reproducible from shipped data)
| Result | Pipeline | Shipped input | Status |
|---|---|---|---|
| Fig 2I/J Differential metabolic pathways altered by EA | metagenomeSeq fitFeatureModel (cumNorm) — F2IJ_Differential Pathway.R |
EC_path_abun_unstrat_norm_descrip.xls |
PRIMARY — authors' own code + data |
| Fig S7 Pearson correlations between keystone species | FS7_Correlation.py (pandas.corr) |
Gut Microbes 6/9.xlsx |
SECONDARY — group labels (WT/APP) not in shipped data |
Out of scope / not attempted (and why)
- Raw 16S → OTU/ASV table (SRA PRJNA1201107 → the 20-species & PICRUSt2 tables): upstream wet-pipeline; processed tables are shipped so we reproduce downstream of them. (The classic hard 20%.)
- Fig 1 LDA / cladogram (LEfSe) —
F1AB_LDA.Rneeds LEfSe intermediates (lefse_table.txt,lefse_results2.txt,bmap) not shipped. - Fig 2K ZiPi —
F2K_ZIPI.Rsources a private vendor script/PERSONALBIO/work/microbio/m13/bin/script/functions.community_similarity.R(not shipped) →env_unresolvablefor that result. - Fig 1E–H / S2 / S4 volcano/boxplot/OPLS-DA, Fig 2A–H centrality/network
rendering — the
.Rmdread un-shipped per-figure spreadsheets (C:«path» N.xlsx). The underlying abundances overlap withGut Microbes 6/9, but exact per-figure sheets/subsets are not shipped; these are plotting-only (no new computed statistic beyond what we test in S7). - Fig S5/S6 OFT behavior — wet-lab behavioral data, not a bioinformatic pipeline.
Reported values targeted (see original/claims.tsv)
- Fig 2I: EA upregulated DHGLUCONATE-PYR-CAT-PWY (glucose degradation) in 6-mo.
- Fig 2J: EA upregulated PWY-7295 (L-arabinose deg IV), GALLATE-DEGRADATION-II-PWY (gallate deg I), CHLOROPHYLL-SYN (chlorophyllide a biosynth I) in 9-mo.
- Fig S7 r-values: −0.48 (M.schaedleri/C.perfringens APP6), −0.64 (WT6), +0.66 (B.pseudolongum/L.hamsteri APP9), −0.46 (B.pullicaecorum/C.celatum APP9).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The primary functional claim reproduces strongly: re-running the authors' own F2IJ_Differential Pathway.R (metagenomeSeq fitFeatureModel) on their own shipped PICRUSt2 table yields DHGLUCONATE-PYR-CAT-PWY (adjP=0.0189) as the sole significant pathway at 6-mo (matching Fig 2I) plus PWY-7295 (0.0098) and GALLATE-DEGRADATION-II-PWY (0.0105) up in EA at 9-mo. Secondary Fig S7 correlations reproduce 3/4 within ~0.03 once the caption's log2-z transform and inferred N/S/EA groupings are applied. The deviations are on our/data-availability side — inferred cohort subsets, a caption-only transform, one unconfirmable significance (CHLOROPHYLL-SYN adjP=NA), and the WT6 cohort being absent from the shipped table — not contradictions of the paper. Overall a solid, no-fabrication reproduction with explainable input-side caveats → yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.