Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Electroacupuncture reshapes the microbial co-occurrence networks related to the behavioral and psychological symptoms of dementia in Alzheimer's disease.

IMetaOmics · 2025
L1 77/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
77/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 50% of all assessed papers rank 572 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1 from shipped data. PRIMARY (Fig 2I/J differential metabolic pathways): re-ran the authors' own script F2IJ_Differential Pathway.R (metagenomeSeq fitFeatureModel, cumNorm) on the authors' own shipped PICRUSt2/EC_path_abun table (419 MetaCyc pathways x 144 samples). All four named EA-upregulated pathways reproduce in the EA-vs-Normal post-intervention contrast: DHGLUCONATE-PYR-CAT-PWY (adjP=0.019, AND the single significant pathway of 419 in the 6-month contrast, matching Fig 2I exactly), PWY-7295 (adjP=0.0098) and GALLATE-DEGRADATION-II-PWY (adjP=0.0105) significant & up in EA at 9 months; CHLOROPHYLL-SYN is up in EA but metagenomeSeq returns adjP=NA so significance is unconfirmable (partial). SECONDARY (Fig S7 Pearson correlations): 3 of 4 reported r-values reproduce within ~0.03 once the figure's stated log2-z transform is applied (e.g. -0.464 vs reported -0.46), but the WT/APP group labels are not keyed in the shipped N/S/EA tables so the exact subsets are inferred, and WT6 r=-0.64 is not regenerable (WT cohort not in the shipped table). NOT attempted: raw 16S->OTU/PICRUSt2 (upstream wet-pipeline; processed tables shipped, so reproduced downstream of them); Fig 1 LEfSe LDA/cladogram (un-shipped intermediates); Fig 2K ZiPi (sources a private vendor script); volcano/heatmap/boxplot/OPLS-DA/network-rendering Rmd (read un-shipped Desktop spreadsheets, plotting-only); Fig S5/S6 OFT (wet-lab behaviour). No reproduced value contradicts the paper; no fabrication indicated.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 77
    assessed: 2026-06-16 ⛓ 47f0a104be6b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Electroacupuncture ameliorates the behavioral and psychological symptoms of dementia (BPSD) in Alzheimer's disease by reshaping the gut microbial co-occurrence networks, driving keystone species, and regulating core microbiota composition and predictive metabolic pathways.

Core claims
  • Electroacupuncture reshapes microbial co-occurrence network topology and drives keystone species in AD-related BPSD, with R. gnavus emerging as a likely keystone species post-intervention. finding
  • Electroacupuncture ameliorates AD-related BPSD-like behavioral phenotypes (e.g., hyperactivity, anxiety) measured by the open field test. finding
  • Gut microbial keystone species and composition vary in a stage-specific (age-stratified) manner across the AD continuum, yielding distinct disease-discriminatory microbial markers for early-to-middle versus middle-to-late stage AD. finding
  • Electroacupuncture regulates predictive functional metabolic pathways, including glucose, L-arabinose, and gallate degradation and chlorophyllide a biosynthesis, in AD-related BPSD. mechanism
  • A Random Forest machine learning model identifies optimal microbial markers discriminating AD-related BPSD by age stage. method
  • 16S rRNA sequencing data deposited at NCBI BioProject PRJNA1201107 serves as a resource for the study. resource
Experimental setups
Assay System Perturbation Readout Platform
16S rRNA gene sequencing APPswe/PS1δE9 (APP/PS1) and wild-type mice, 6- and 9-month-old (fecal/gut microbiota) none (disease model vs WT) gut microbiota composition / relative abundance of bacterial taxa
16S rRNA gene sequencing APP/PS1 mice, 6- and 9-month-old electroacupuncture (vs Sham and Normal control) post-intervention microbial composition, co-occurrence network topology, keystone species HANS-JS502A electroacupuncture device
Open field test (OFT, behavioral assay) APP/PS1 and WT mice, 6- and 9-month-old electroacupuncture vs Sham vs Normal control total distance traveled, central zone crossing frequency, distance/duration in central zone
PICRUSt2 functional prediction gut microbiota of 6- and 9-month-old APP/PS1 mice electroacupuncture predicted metabolic pathway abundance (log2FC) PICRUSt2
Random Forest machine learning classification gut microbiota of APP/PS1 vs WT mice none importance scores of microbial markers for AD-related BPSD
LEfSe / Linear discriminant analysis (LDA) gut microbiota of 6- and 9-month-old APP/PS1 vs WT mice none differentially abundant taxa (LDA score)
Co-occurrence network analysis (Pearson correlation, ZiPi) gut microbiota of APP/PS1 and WT mice, 6- and 9-month-old none / electroacupuncture centrality measures (degree, closeness, betweenness, expected influence), keystone species, module connectivity
Key results
  • Bifidobacterium pseudolongum, Lactobacillus hamsteri, and Parabacteroides distasonis were significantly enriched in both 6- and 9-month-old APP/PS1 mice p_adj < 0.05
  • Akkermansia muciniphila and Clostridium cocleatum were significantly depleted in both 6- and 9-month-old APP/PS1 mice p_adj < 0.05
  • In 6-month-old APP/PS1 mice, Mucispirillum schaedleri and C. perfringens were negatively correlated and identified as keystone species r = -0.48, p_adj < 0.05
  • In 9-month-old APP/PS1 mice, B. pseudolongum and L. hamsteri were positively correlated keystone species r = 0.66, p_adj < 0.05
  • In 9-month-old APP/PS1 mice, B. pullicaecorum and C. celatum exhibited significant negative correlation as keystone species r = -0.46, p_adj < 0.05
  • OFT total distance traveled and central zone crossing frequency were significantly greater in APP/PS1 mice than age-matched WT mice; all such indices decreased significantly in EA groups vs Sham post-intervention p_adj < 0.05
  • Electroacupuncture significantly upregulated the DHGLUCONATE-PYR-CAT-PWY glucose degradation pathway in 6-month-old APP/PS1 mice
  • Electroacupuncture altered PWY-7295 (L-arabinose degradation IV), GALLATE-DEGRADATION-II-PWY, and CHLOROPHYLL-SYN (chlorophyllide a biosynthesis I) pathways in 9-month-old APP/PS1 mice
Key statistics
  • correlation r = -0.48 (M. schaedleri vs C. perfringens, 6-month-old APP/PS1 network)
  • correlation r = -0.64 (M. schaedleri vs C. perfringens, 6-month-old WT network)
  • correlation r = 0.66 (B. pseudolongum vs L. hamsteri, 9-month-old APP/PS1 network)
  • correlation r = -0.46 (B. pullicaecorum vs C. celatum, 9-month-old APP/PS1 network)
  • pvalue p_adj < 0.05 (significance threshold for differential taxa, behavioral indices, and correlations)
  • other LDA threshold = 2 (Log10 scale) (LDA score threshold for differentially abundant taxa)
  • other Pi = 2.5; Zi = 0.62 (ZiPi plot thresholds dividing microbes into network hubs, module hubs, connectors, peripherals)
  • count top 20 species; top 5 species (20 most abundant species in co-occurrence network; top 5 in ZiPi plot)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This animal study compared gut microbiota profiles between APPswe/PS1δE9 (APP/PS1) transgenic and wild-type (WT) mice at two age points (6 and 9 months) and then evaluated electroacupuncture (EA) effects across Normal, Sham, and EA groups at each age. Differential taxa were identified with LEfSe and visualized via volcano plots and heatmaps; co-occurrence networks were built from Pearson correlation coefficients; Random Forest ranked microbial marker importance; and PICRUSt2 predicted functional metabolic pathways. All significance thresholds were expressed as p_adj < 0.05, with LDA scores, Pearson r, and log2 fold-change serving as the primary effect measures.

Replicationbiological GroupsAPP/PS1 vs age-matched WT mice at 6 and 9 months (baseline); Normal vs Sham vs EA within each age group post-intervention (3 weeks) Pairingunpaired Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnot stated in main text (p_adj used throughout; method referenced in supplementary Table S5)
Statistical tests used
Test Applied to n Assumptions
LEfSe (Linear discriminant analysis Effect Size) with LDA threshold = 2 (log10 scale) Differential abundance of taxa between APP/PS1 and WT mice at 6 and 9 months (Figures 1A–B; Tables S1–S4) not stated
Pearson correlation coefficient on Z-score-standardized log2-transformed relative abundances Microbial co-occurrence network construction for all groups and age points (Figures 2E–H, S7) not stated
Random Forest (ensemble decision-tree importance scoring) Ranking of candidate microbial markers for BPSD in early-to-middle and middle-to-late AD stages (Figures S2C–D) na
OPLS-DA (Orthogonal partial least squares-discriminant analysis) Multivariate clustering of microbial composition across all groups post-intervention (Figure S4) not stated
PICRUSt2 functional pathway prediction (log2FC reported) Metabolic pathway differences between EA and control groups in 6- and 9-month APP/PS1 mice (Figures 2I–J) na
Unspecified inferential test (p_adj < 0.05 reported; specific test in supplementary Table S5) Open field test (OFT) behavioral comparisons between APP/PS1 and WT mice at baseline, and between EA, Sham, and Normal groups post-intervention (Figures S5–S6) not stated
Approaches that could also have been used
  • Co-occurrence networks were constructed using Pearson correlation coefficients applied to Z-score-standardized, log2-transformed relative abundances
    Could also: Spearman rank correlation or SparCC (Sparse Correlations for Compositional data) could also be used to build co-occurrence networks from 16S relative-abundance tables — Spearman is more robust to non-normality and outliers common in microbiome data; SparCC was specifically designed to address the compositionality constraint of relative abundances, where proportions are not independent across taxa and can inflate or deflate Pearson coefficients in ways that are difficult to anticipate
  • Differential taxa were identified using LEfSe with a fixed LDA score threshold
    Could also: DESeq2, edgeR, or ANCOM-BC could also be applied to 16S count data to identify differentially abundant taxa — These methods model integer count data directly with negative-binomial or related distributions and provide taxon-level FDR control with explicit statistical testing; they are widely used when the priority is formal inference alongside discovery
  • Microbial marker importance was evaluated with a Random Forest importance score
    Could also: LASSO or elastic-net penalized logistic regression could also be used for microbial biomarker selection — Penalized regression yields explicit feature selection with direct probabilistic interpretability, and stability-selection wrappers provide confidence estimates for included features, which can be useful in small-n animal study contexts where Random Forest importance scores may vary across runs
  • Functional metabolic pathway abundances were inferred from 16S amplicon data using PICRUSt2
    Could also: Shotgun metagenomic sequencing analyzed with HUMAnN3 or similar pipelines could also provide functional pathway profiles — Direct sequencing of the functional gene content bypasses the reference-database and phylogenetic-placement assumptions inherent in PICRUSt2, whose accuracy depends on how well the community is represented by sequenced reference genomes
  • Open field test behavioral outcomes were compared between groups with adjusted p-values; the specific inferential test is not named in the main text
    Could also: A two-way ANOVA (factors: genotype/treatment × age) followed by Tukey HSD or Dunnett post-hoc correction could also be used to test OFT indices across all groups simultaneously — A factorial ANOVA tests main effects and their interaction within a single model and controls the family-wise error rate across all pairwise comparisons in one step, making explicit whether age and treatment effects are additive or interactive
  • Results throughout are expressed as threshold-based adjusted p-values (p_adj < 0.05) without exact values, and without confidence intervals for behavioral or abundance comparisons
    Could also: Reporting exact adjusted p-values alongside effect-size estimates with 95% confidence intervals (e.g., fold-change ± CI for abundance, Cohen's d for behavioral endpoints) would also be standard practice — Exact values allow readers to apply their own evidentiary thresholds and to power future studies; confidence intervals convey both the magnitude and the precision of group differences, which is particularly informative in small-n animal experiments where uncertainty can be high
Software: LEfSe · PICRUSt2

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41676443

Paper: Su et al. 2025, iMetaOmics. "Electroacupuncture reshapes the microbial co-occurrence networks related to the behavioral and psychological symptoms of dementia in Alzheimer's disease." DOI 10.1002/imo2.70035 · PMCID PMC12806058. Repo: https://github.com/sufuyou2021/IMO-2025-0041 @ commit 7b72fdf0354dd00999eef9c50fe6b70e142480ee Raw data: SRA PRJNA1201107 (16S rRNA amplicon).

Study: 16S microbiome of APP/PS1 (AD model) vs WT mice at 6- and 9-months; an electroacupuncture (EA) intervention with Normal (N), Sham (S), EA groups, baseline (_6/_9) and 3-weeks-post (_6W3/_9W3), 12 mice/group.

What the repo ships (processed, human-auditable)

  • Data/Gut Microbes 6.xlsx, Gut Microbes 9.xlsxtop-20 species relative abundance, 20 species × 72 samples (N/S/EA × 12 × {baseline, W3}).
  • PICRUSt2/EC_path_abun_unstrat_norm_descrip.xls419 MetaCyc pathways × 144 samples (the .xls files are actually TSV). Also KO/EC/COG predictions.
  • R/Rmd/py scripts that render each figure, plus the authors' rendered .html.

In scope (pipeline-derived, reproducible from shipped data)

Result Pipeline Shipped input Status
Fig 2I/J Differential metabolic pathways altered by EA metagenomeSeq fitFeatureModel (cumNorm) — F2IJ_Differential Pathway.R EC_path_abun_unstrat_norm_descrip.xls PRIMARY — authors' own code + data
Fig S7 Pearson correlations between keystone species FS7_Correlation.py (pandas.corr) Gut Microbes 6/9.xlsx SECONDARY — group labels (WT/APP) not in shipped data

Out of scope / not attempted (and why)

  • Raw 16S → OTU/ASV table (SRA PRJNA1201107 → the 20-species & PICRUSt2 tables): upstream wet-pipeline; processed tables are shipped so we reproduce downstream of them. (The classic hard 20%.)
  • Fig 1 LDA / cladogram (LEfSe)F1AB_LDA.R needs LEfSe intermediates (lefse_table.txt, lefse_results2.txt, bmap) not shipped.
  • Fig 2K ZiPiF2K_ZIPI.R sources a private vendor script /PERSONALBIO/work/microbio/m13/bin/script/functions.community_similarity.R (not shipped) → env_unresolvable for that result.
  • Fig 1E–H / S2 / S4 volcano/boxplot/OPLS-DA, Fig 2A–H centrality/network rendering — the .Rmd read un-shipped per-figure spreadsheets (C:«path» N.xlsx). The underlying abundances overlap with Gut Microbes 6/9, but exact per-figure sheets/subsets are not shipped; these are plotting-only (no new computed statistic beyond what we test in S7).
  • Fig S5/S6 OFT behavior — wet-lab behavioral data, not a bioinformatic pipeline.

Reported values targeted (see original/claims.tsv)

  • Fig 2I: EA upregulated DHGLUCONATE-PYR-CAT-PWY (glucose degradation) in 6-mo.
  • Fig 2J: EA upregulated PWY-7295 (L-arabinose deg IV), GALLATE-DEGRADATION-II-PWY (gallate deg I), CHLOROPHYLL-SYN (chlorophyllide a biosynth I) in 9-mo.
  • Fig S7 r-values: −0.48 (M.schaedleri/C.perfringens APP6), −0.64 (WT6), +0.66 (B.pseudolongum/L.hamsteri APP9), −0.46 (B.pullicaecorum/C.celatum APP9).
Figures / tables: Figure 2IFigure 2JFigure S7AFigure S7BFigure S7C
C1
Reported
EA significantly upregulated DHGLUCONATE-PYR-CAT-PWY (glucose degradation) in 6-mo APP/PS1 (Fig 2I)
Reproduced
adjP=0.0189, higher in EA (2.27 vs N 0.52); the SOLE significant pathway of 419 in EA-vs-Normal 6W3
exact
C2
Reported
EA upregulated PWY-7295 (L-arabinose degradation IV) in 9-mo APP/PS1 (Fig 2J)
Reproduced
adjP=0.0098, higher in EA (1.27 vs 0.22)
exact
C3
Reported
EA upregulated GALLATE-DEGRADATION-II-PWY (gallate degradation I) in 9-mo APP/PS1 (Fig 2J)
Reproduced
adjP=0.0105, higher in EA (0.55 vs 0.13)
exact
C4
Reported
EA upregulated CHLOROPHYLL-SYN (chlorophyllide a biosynthesis I) in 9-mo APP/PS1 (Fig 2J)
Reproduced
higher in EA (4.79 vs 1.59) but metagenomeSeq returns adjP=NA (excess zeros) — direction only
partial
C5
Reported
r=-0.48 M.schaedleri~C.perfringens 6-mo APP network (Fig S7A)
Reproduced
r=-0.475 (Sham group, raw rel.abund, n=12)
within tolerance
C6
Reported
r=-0.64 M.schaedleri~C.perfringens 6-mo WT network (Fig S7B)
Reproduced
not regenerable from shipped N/S/EA data (WT cohort absent)
did not match
C7
Reported
r=+0.66 B.pseudolongum~L.hamsteri 9-mo APP network (Fig S7C)
Reproduced
r=+0.632 (Sham+EA AD mice, log2-z per caption, n=24)
within tolerance
C8
Reported
r=-0.46 B.pullicaecorum~C.celatum 9-mo APP network (Fig S7C)
Reproduced
r=-0.464 (Sham+EA, log2-z, n=24) — near-exact
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 77/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

The primary functional claim reproduces strongly: re-running the authors' own F2IJ_Differential Pathway.R (metagenomeSeq fitFeatureModel) on their own shipped PICRUSt2 table yields DHGLUCONATE-PYR-CAT-PWY (adjP=0.0189) as the sole significant pathway at 6-mo (matching Fig 2I) plus PWY-7295 (0.0098) and GALLATE-DEGRADATION-II-PWY (0.0105) up in EA at 9-mo. Secondary Fig S7 correlations reproduce 3/4 within ~0.03 once the caption's log2-z transform and inferred N/S/EA groupings are applied. The deviations are on our/data-availability side — inferred cohort subsets, a caption-only transform, one unconfirmable significance (CHLOROPHYLL-SYN adjP=NA), and the WT6 cohort being absent from the shipped table — not contradictions of the paper. Overall a solid, no-fabrication reproduction with explainable input-side caveats → yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

157.6 k
tokens (I/O) · 10.1 M incl. cache
16 min
runtime · 0.01 CPU-h
1.3 GB
peak RAM
4
HPC jobs
hummel
machine