Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Dental Plaque Microbial Resistomes of Periodontal Health and Disease and Their Changes after Scaling and Root Planing Therapy.

mSphere · 2021
L1 76/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
76/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 48% of all assessed papers rank 586 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (primary target) — third-party-tool-on-paper's-data per brief rule P16. All 112 public dental-plaque shotgun metagenomes (HS48/PS40/RS24 across 4 open SRA BioProjects, all retrievable + complete) were downloaded, human-decontaminated (bowtie2/hg19), and run through the cited ARGs-OAP/SARG protocol using the paper's EXACT SARG v2.2 database (24 types/1244 subtypes/12085 genes) at the paper cutoffs (E<=1e-7, id>=80%, len>=25aa). RESULT: pooled ARG TYPE count = 18 — EXACT match to the paper's headline 18 ARG types. Per-group shared-ARG abundance fraction reproduced within ~1% (HS98.8/PS98.1/RS99.4 vs 98.4/99.2/99.6). Total non-human read volume within ~5% (1.61e9 pairs vs 1.53e9). Top-6 ARG types 5/6 concordant (beta-lactam, tetracycline, multidrug, MLS, bacitracin shared; #6 aminoglycoside vs paper kasugamycin). Subtype count 227 vs 269 and per-group Venn (HS161/PS195/RS160; 116 shared vs HS209/PS223/RS181; 134) run ~15-20% lower — fully explained by the diamond-vs-NCBI-blastx engine substitution (forced by an OOM on the pooled 10.65M-read query) plus ARGs-OAP DB-version drift; faithful method on the identical SARG v2.2 DB universe, graded partial, never forced. Taxonomy/MRG/MGE/network were stretch goals and not attempted. No authors' analysis-code repo exists (the cited GitHub is a third-party MGE reference DB); reproduction is by re-running the described pipeline. All grades provisional pending human audit; no values fabricated.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ 4c70f3b0c468
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates how periodontitis and scaling and root planing (SRP) therapy affect the composition and abundance of antibiotic-resistant genes (ARGs) and metal-resistant genes (MRGs) in the dental plaque microbiota, and how these resistomes relate to microbial community structure and mobile genetic elements.

Core claims
  • Periodontitis significantly alters dental plaque microbial community diversity and structure compared to healthy and treated states finding
  • NetShift analysis identifies Fretibacterium fastidiosum, Tannerella forsythia, and Campylobacter rectus as key driver species in periodontitis progression finding
  • Periodontitis and SRP treatment increase the number of ARGs and MRGs in dental plaque and significantly alter ARG/MRG profile composition finding
  • Bacitracin, beta-lactam, MLS, tetracycline, and multidrug resistance genes are the main high-abundance ARG classes; multimetal, iron, chromium, and copper resistance genes are the primary MRG types in dental plaque finding
  • Cooccurrence of ARGs, MRGs, and mobile genetic elements (MGEs) indicates a coselection phenomenon in dental plaque resistomes finding
  • Microbial community composition significantly correlates with and shapes ARG distribution in dental plaque, per Procrustes analysis finding
  • Haemophilus parainfluenzae is a predicted host for the majority of ARG subtypes based on cooccurrence network analysis finding
  • Metagenomic shotgun sequencing (PCR-independent) was used to comprehensively profile the dental plaque resistome across 112 samples method
Experimental setups
Assay System Perturbation Readout Platform
metagenomic shotgun sequencing human dental plaque (112 samples: HS, PS, RS) periodontitis disease state / SRP treatment taxonomic composition and resistome profiles
taxonomic/species profiling dental plaque metagenome disease state (HS/PS/RS) species relative abundance, alpha/beta diversity MetaPhlAn
ARG abundance profiling against reference resistance gene database dental plaque metagenome periodontitis / SRP treatment ARG type/subtype relative abundance (copies/cell)
MRG abundance profiling against reference resistance gene database dental plaque metagenome periodontitis / SRP treatment MRG type/subtype relative abundance (copies/cell)
NetShift network analysis dental plaque microbiome network (control = HS+RS vs case = PS) periodontitis vs healthy/resolved driver species NESH score NetShift
Procrustes analysis of Bray-Curtis dissimilarity matrices dental plaque metagenome none correlation between ARG/MRG profiles and microbial community structure
Spearman correlation network analysis dental plaque metagenome (≥56 of 112 samples) none cooccurrence between ARG subtypes and microbial taxa
Key results
  • Shannon and Simpson diversity indices significantly higher in PS group than HS and RS groups; no significant difference between HS and RS
  • 13 species (e.g., Tannerella forsythia, Porphyromonas gingivalis, Prevotella intermedia, Filifactor alocis) significantly enriched in PS group vs HS/RS
  • NetShift identified 3 driver species (Fretibacterium fastidiosum, Tannerella forsythia, Campylobacter rectus) with higher NESH scores
  • Detected 18 ARG types and 269 ARG subtypes (of 1,244 reference subtypes/24 types) in dental plaque 269 subtypes / 18 types
  • Detected 18 MRG types and 240 subtypes; multimetal resistance genes most abundant (0.060-1.586 copies/cell), followed by iron, chromium, copper 0.060-1.586 copies/cell (multimetal)
  • Number of ARG and MRG subtypes significantly higher in PS and RS groups than HS group, with no significant difference between PS and RS in ARG count
  • Procrustes analysis showed significant correlation between ARG profiles and microbial community composition r=0.7649
  • Shared (core) ARG subtypes represented near-total abundance across groups while shared MRG subtypes represented a markedly lower proportion of total abundance ARG ~98-99%; MRG ~61-66%
Key statistics
  • correlation r=0.7649, M2=0.4149 (Procrustes test of ARG profile vs microbial community Bray-Curtis dissimilarity)
  • pvalue P=0.001 (999 permutations) (significance of Procrustes ARG-microbiome correlation)
  • count 48 HS, 40 PS, 24 RS samples (112 total) (study cohort composition)
  • count 269 ARG subtypes across 18 types (out of 1,244 subtypes/24 types in reference database) (ARG detection breadth)
  • count 240 MRG subtypes across 18 types (MRG detection breadth)
  • mean 98.39% ± 5.85% (HS), 99.23% ± 0.91% (PS), 99.55% ± 0.57% (RS) (proportion of total ARG abundance from 134 shared ARG subtypes)
  • mean 61.13% ± 9.58% (HS), 64.30% ± 10.34% (PS), 65.86% ± 11.60% (RS) (proportion of total MRG abundance from 160 shared MRG subtypes)
  • other Spearman's rho > 0.6, Q value < 0.01; 57 nodes (32 taxa, 25 ARG subtypes), 85 edges (cooccurrence network between ARG subtypes and microbial taxa)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This retrospective metagenomics study compared dental plaque resistomes across 48 healthy-state (HS), 40 periodontitis-state (PS), and 24 post-SRP resolved-state (RS) samples retrieved from the NCBI SRA database. Alpha-diversity (Shannon, Simpson) and beta-diversity (Bray-Curtis and Jaccard PCoA) were used to characterize microbial community and resistome differences among groups; significance of group separation was assessed but the specific multivariate permutation test (e.g., PERMANOVA) was not named. Cooccurrence networks were built on Spearman correlations with FDR-corrected Q-value thresholds, Procrustes permutation analysis linked ARG profiles to microbial structure, and NetShift analysis identified ecological driver taxa. Results were reported primarily as relative abundances, p-value threshold symbols, and Procrustes fit statistics.

Replicationbiological Sample sizeHS n=48, PS n=40, RS n=24 (total 112); samples retrieved from NCBI SRA; no a priori power calculation described GroupsThree groups: healthy-state (HS), periodontitis-state pre-treatment (PS), resolved-state post-SRP (RS) Pairingunclear Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionFDR correction implied by Q-value threshold (Q < 0.01) for Spearman cooccurrence network; correction method for alpha-diversity or species differential abundance comparisons not stated
Statistical tests used
Test Applied to n Assumptions
Alpha-diversity comparison (Shannon-Wiener index, inverse Simpson index) — specific pairwise test not named Fig. 1a and Fig. S1c: HS vs PS vs RS microbial community species-level diversity HS n=48, PS n=40, RS n=24 not stated
Beta-diversity significance testing — specific permutation test (e.g., PERMANOVA/ANOSIM) not named; PCoA visualisation of Bray-Curtis and Jaccard distances Fig. 1b, Fig. S1d, Fig. 2c, Fig. S2c, Fig. 4c, Fig. S4c: community and resistome structure among three groups HS n=48, PS n=40, RS n=24 not stated
Procrustes permutation test (999 permutations) Fig. 2g and Fig. 4g: correlation between ARG/MRG profiles and microbial community structure (Bray-Curtis dissimilarity matrices); M²=0.4149, r=0.7649, p=0.001 reported for ARGs n=112 total samples not stated
Spearman rank correlation (ρ > 0.6, Q < 0.01) Fig. S3 and Table S2: cooccurrence network between ARG subtypes and microbial taxa; applied to taxa/ARGs present in ≥56/112 samples n=112 samples; edges based on ≥56 co-occurring samples not stated
Pairwise group comparison of Bray-Curtis dissimilarity and Jaccard index values — specific test not named Fig. 2d, Fig. S2d, Fig. 4d, Fig. S4d: pairwise between-group beta-diversity distances for ARGs and MRGs HS n=48, PS n=40, RS n=24 not stated
Differential abundance testing for individual species — specific test not named; results reported as p<0.01, p<0.05, p≥0.05 thresholds Fig. 1c: top 45 most different species across HS, PS, RS groups HS n=48, PS n=40, RS n=24 not stated
NetShift analysis (NESH score-based ecological network shift) Fig. 1d: identification of driver taxa between control (HS+RS) and case (PS) microbiome networks HS n=48, RS n=24 as controls; PS n=40 as cases not stated
Approaches that could also have been used
  • The RS (post-SRP) group comprised 24 subjects while the PS (pre-SRP) group comprised 40 subjects, and the relationship between these two groups (whether subjects were matched) was not described; all three groups were analysed as independent cross-sectional samples
    Could also: If PS and RS samples come from the same individuals before and after treatment, a paired design (e.g., Wilcoxon signed-rank test for alpha diversity, or paired PERMANOVA for beta diversity) could be applied — Paired analysis accounts for within-subject correlation, increases statistical power, and directly quantifies the within-person treatment effect — which is the primary biological question addressed by the SRP comparison
  • Alpha-diversity indices (Shannon, Simpson) were compared across three groups, but the statistical test used to determine significance was not named
    Could also: A Kruskal-Wallis test with Dunn's post-hoc correction (or a one-way ANOVA with Tukey HSD if normality holds) is a standard named approach for three-group alpha-diversity comparisons — Naming the test and its correction for three pairwise comparisons allows readers to assess assumptions, reproduce the analysis, and interpret which specific pairs differ
  • Beta-diversity group separation was described as 'significant' across multiple distance matrices and resistome types, but the permutation test used (e.g., PERMANOVA/adonis, ANOSIM) was not named
    Could also: PERMANOVA (via adonis in vegan) with pairwise post-hoc tests and FDR correction is a standard named approach for testing multivariate group differences in microbiome studies — Reporting the test name, F-statistic, R², and corrected p-values for each pairwise comparison gives readers the effect size alongside the p-value and enables reproducibility
  • Species differential abundance across three groups was assessed and reported only as p-value threshold symbols (<0.01, <0.05, ≥0.05), without naming the statistical test or showing exact p-values
    Could also: Established metagenomics differential abundance methods such as DESeq2, MaAsLin2, or ANCOM-BC handle compositionality and provide effect estimates (fold-change or coefficient) with confidence intervals alongside adjusted p-values — These methods address the compositional nature of relative-abundance data, report exact adjusted p-values and effect sizes, and allow readers to judge biological relevance alongside statistical significance
  • Cooccurrence networks were built using a fixed Spearman ρ > 0.6 threshold with Q < 0.01, applied to features present in ≥ 50% of samples
    Could also: Compositionally-aware correlation methods such as SparCC, SPIEC-EASI, or CLR-transformed Pearson/Spearman correlations could also be used for relative-abundance metagenomics data — Standard Spearman correlations on relative-abundance data can be affected by compositionality (spurious correlations induced by the constant-sum constraint); compositionally-aware methods are designed to mitigate this issue
  • Dispersion around means was reported inconsistently: some values as mean ± SD, others as ranges, and group comparisons shown as box plots without explicit reporting of IQR or whisker definitions
    Could also: Consistently reporting median with IQR (for non-normal data displayed as box plots) or mean with 95% CI alongside exact p-values would also convey both central tendency and uncertainty — Consistent dispersion reporting, together with exact p-values rather than threshold symbols, allows meta-analyses and direct comparison with other studies reporting the same metrics
Software: MetaPhlAn · NetShift · Unspecified tool for ARG/MRG annotation (reference resistance gene database with 1,244 ARG subtypes)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34287005

Paper: Kang Y, Sun B, Chen Y, Lou Y, Zheng M, Li Z. Dental Plaque Microbial Resistomes of Periodontal Health and Disease and Their Changes after Scaling and Root Planing Therapy. mSphere 2021. PMID 34287005 · PMCID PMC8386447 · DOI 10.1128/msphere.00162-21.

Nature of the study

A re-analysis (meta-analysis) of 112 public dental-plaque shotgun metagenomes pooled from 4 NCBI BioProjects. The paper generated no new sequencing — all data are public SRA. Three groups:

  • HS = healthy, 48 individuals
  • PS = periodontitis, before treatment, 40 individuals
  • RS = periodontitis, after Scaling-and-Root-Planing (SRP), 24 individuals

Data: PRJNA255922 (no publication), PRJNA528558, PRJNA625082, PRJNA230363. Exact per-sample accessions + group labels are in Table S1.

The "code" link is a reference DATABASE, not a pipeline

github.com/KatariinaParnanen/MobileGeneticElementDatabase is a third-party MGE reference database (a FASTA + annotation table), used by the paper only for the MGE-annotation step. There is no authors' analysis-code repository. Per brief rule P16 this is fine: we reproduce by running the described third-party pipelines on the paper's own public data.

Pipeline-derived results (IN SCOPE)

step tool described reproducible?
QC: human removal (hg19), quality filtering (bowtie2/bwa + quality filter) yes
Taxonomy / species counts, alpha+beta diversity MetaPhlAn2 (default), vegan partial (DB-version sensitive)
ARG annotation + abundance (copies/cell) ARGs-OAP / SARG (Yin et al protocol; 24 types, 1244 subtypes; E≤1e-7, ≥80% id, ≥25 aa) yes — PRIMARY target
MRG annotation + abundance BLASTX vs BacMet v2.0 (E≤1e-7, ≥80% id, ≥25 aa) yes (stretch)
MGE annotation + abundance BLASTX vs KatariinaParnanen MGE DB (cited repo) yes (stretch)
diversity / PCoA / PERMANOVA / procrustes / networks vegan, Hmisc, Gephi downstream of above

OUT OF SCOPE

  • Wet-lab / clinical subject recruitment (none — public data).
  • NetShift driver-species network, Gephi visualizations (manual/visual, not a pinnable numeric pipeline output) — noted but not graded.

Primary reproduction target

Run ARGs-OAP (the exact cited SARG-based protocol) on the 112 metagenomes → reproduce the pooled detection counts (18 ARG types / 269 ARG subtypes), the top-6 ARG types, and the per-group Venn counts. This is the cleanest 1:1 third-party-tool-on-paper's-data reproduction. Then push to taxonomy (MetaPhlAn), MRG (BacMet), and MGE (cited DB) as stretch results.

Honesty caveats known up front

  • ARGs-OAP/SARG database version drift: paper used the 2018-era SARG (24 types / 1244 subtypes). Current bioconda args_oap may bundle a newer SARG → detected-subtype counts may differ; this is a faithful-method, different-DB situation and will be graded accordingly (within-tol/partial), never forced.
  • MetaPhlAn2 marker DB is frozen to a version; exact species counts (243/14/1) are DB-version sensitive.
Figures / tables: Fig 2eFig 3
data-N-total
Reported
112 dental plaque metagenomes
Reproduced
112 (all SRR runs downloaded + processed across 4 BioProjects)
exact
data-N-groups
Reported
HS=48, PS=40, RS=24
Reproduced
HS=48, PS=40, RS=24
exact
seq-total-nonhuman
Reported
1,529,804,857 non-human reads (mean 13,658,972)
Reproduced
1,611,069,871 read-pairs (mean 14,384,552); +5.3%
within tolerance
arg-types-detected
Reported
18 ARG types (pooled)
Reproduced
18
exact
arg-subtypes-detected
Reported
269 ARG subtypes (pooled)
Reproduced
227
partial
arg-top6-types
Reported
beta-lactam;tetracycline;multidrug;MLS;bacitracin;kasugamycin
Reproduced
tetracycline;MLS;multidrug;beta-lactam;bacitracin;aminoglycoside (5/6 overlap)
within tolerance
arg-shared-pct
Reported
HS 98.39%, PS 99.23%, RS 99.55%
Reproduced
HS 98.81%, PS 98.14%, RS 99.36% (within ~1%)
within tolerance
arg-venn
Reported
HS 209, PS 223, RS 181; 134 shared
Reproduced
HS 161, PS 195, RS 160; 116 shared
partial
taxa-counts-metaphlan
Reported
243 bacterial + 14 viral + 1 fungal
Reproduced
not attempted (stretch)
partial
mrg-bacmet
Reported
18 MRG types / 240 subtypes
Reproduced
not attempted (stretch)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 76/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

87.6 k
tokens (I/O) · 5.4 M incl. cache
11 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.