Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Metatranscriptomics of the human oral microbiome during health and disease.

mBio · 2014
L1 62/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
62/100
Reproducibility score
0.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 23% of all assessed papers rank 891 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Partial, honestly-graded reproduction. The gene-level pipeline (custom 59-organism metatranscriptome reference -> Flexbar trim -> Bowtie2 unique alignment -> HTSeq-count -> paired edgeR GLM/LRT DE) was executed end-to-end for all 6 biological samples (8 SRA runs) after debugging two self-introduced bugs in the GFF reference-building step (coordinate corruption from GenBank flatfile continuation-line misrouting; GFF3 attribute-delimiter corruption from unescaped ';' in gene product descriptions). Reproduced DE results show real but moderate concordance with the paper's SuppFile2 ground truth: only 51.9% of the paper's 161,685 reported gene loci are present in this reproduction's reference (a consequence of substituting the decommissioned legacy HOMD SEQF-ID lookup with a modern NCBI eutils assembly search for 26/59 organisms, plus this reproduction's own from-scratch GenBank flatfile parser for the other 33 vs. the original's presumably different/more complete extraction), and among the overlapping genes, logFC correlation is moderate (Pearson r=0.45) with 60.5% signconcordance -- directionally consistent with real biological signal but far from an exact reproduction. EC-level/pathway analysis (SuppFile3.xlsx) was NOT attempted -- it requires a separate EC-annotation sub-pipeline (KEGG/UniProt protein-to-EC mapping) that was out of reach within this reproduction pass's scope. No completeness is claimed beyond what is stated here; all comparisons are provisional pending human review.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-31
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Because microbiome-associated diseases such as periodontitis produce similar symptoms despite interpatient variability in microbial community composition, the authors hypothesized that human-associated microbial communities undergo conserved changes in metabolism (community-level metabolic gene expression) during the transition from health to disease.

Core claims
  • Disease-associated periodontal communities display conserved community-level metabolic gene expression profiles between patients, whereas the metabolic gene expression of individual species is highly variable between patients. finding
  • Conserved disease-associated metabolic changes are accomplished by a patient-specific cohort of interchangeable microbes, explaining the low conservation seen in metagenomic/marker-gene surveys. mechanism
  • Fusobacterium nucleatum acts as a keystone species in periodontitis, being the sole bacterium responsible for community lysine fermentation to butyrate in all patients, without a change in its relative rRNA abundance. finding
  • Disease-associated communities are significantly less diverse and less species rich than patient-matched health-associated communities, and are dominated by a few ribosome-rich organisms. finding
  • Pathways upregulated in all diseased sites include lysine fermentation to butyrate, histidine catabolism, nucleotide biosynthesis, and pyruvate fermentation. finding
  • Bacteria expressing known extracellular virulence factors (collagenases, proteases) differ between patients while the virulence functions themselves are conserved. finding
  • Binning >160,000 metagenome genes by Enzyme Commission (EC) number enables community-wide, gene-expression-based metabolic reconstruction of the microbiome in health and disease. method
  • A 60-organism reference 'metagenome' (completed and draft HOMD genomes) plus per-gene raw read counts and differential EC expression tables are provided as a community resource. resource
Experimental setups
Assay System Perturbation Readout Platform
rRNA (16S) sequencing for community composition and alpha/beta diversity Patient-matched healthy and diseased subgingival periodontal plaque from 10 aggressive periodontitis patients (pools of 3 teeth per condition; 30 health- and 30 disease-associated populations) none (disease state: health vs aggressive periodontitis) OTU counts, Shannon index, unweighted UniFrac beta diversity, relative rRNA abundance per taxon
Comparison of rRNA gene (DNA) abundance versus rRNA abundance Subset of the patient-matched healthy and diseased periodontal plaque samples none (health vs disease) Correlation between rRNA gene abundance and rRNA abundance per population
Metatranscriptomics (community RNA-seq) after human and bacterial rRNA depletion 9 health- and 9 disease-associated periodontal plaque populations from 3 patient-matched aggressive periodontitis patients none (health vs disease) mRNA read counts mapped to a 60-organism metagenome (>160,000 genes); differential expression per gene and per EC gene family Illumina HiSeq
Read alignment / database classification of RNA-seq reads RNA-seq reads from the 18 periodontal plaque samples none Fraction of reads aligning to HOMD (4.4 Gbp), RefSeq human RNA (135 Mbp), RefSeq viral genomes (121 Mbp), and to the 60-organism metagenome Human Oral Microbiome Database (HOMD), RefSeq huRNA, RefSeq virusDB
EC-based metabolic network reconstruction / pathway mapping 60-organism oral metagenome gene expression from 3 patients (health vs disease) none (health vs disease) Up-/down-/unchanged expression of enzyme-encoding gene families mapped onto metabolic pathways KEGG pathway database; Enzyme Commission (EC) annotation
Variance/dispersion analysis of expression (individual gene vs EC-binned vs randomly binned control) Metatranscriptomes of the 60-organism metagenome across patients in silico control: genes randomly binned into 1,137 groups Tagwise and common dispersion (biological coefficient of variance) as a function of expression level edgeR
Species-resolved expression of metabolic and virulence genes Diseased periodontal communities of patients 1, 2, and 3 (e.g., F. nucleatum, T. forsythia, P. tannerae, P. gingivalis) none (health vs disease) Normalized per-species expression (log2 reads per million) of lysine fermentation, histidine degradation/THF, pyruvate fermentation, collagenase, and protease genes
Clinical periodontal examination 10 aggressive periodontitis patients (ages 30-40; 4 male, 6 female) none Probing depth (mm), clinical attachment loss (mm), plaque index, percent bleeding, smoking status
Key results
  • Disease-associated communities were less species rich than patient-matched health-associated communities (Shannon index) P = 0.03
  • Health- and disease-associated populations segregated into distinct groups, with diseased populations less dispersed than healthy ones PERMANOVA P = 0.01; PERMDISP P = 0.008
  • Disease-associated populations were more related to the average disease state (centroid) than to the paired health-associated population from the same individual P = 0.0005
  • About 18% of ~1,100 unique enzyme-encoding gene families were differentially expressed at the microbiome level during disease ~18% of ~1,100 EC families, P < 0.05
  • Lysine fermentation to butyrate, histidine catabolism, nucleotide biosynthesis, and pyruvate fermentation showed enhanced gene expression in all diseased sites
  • F. nucleatum was the sole bacterium responsible for community lysine degradation to butyrate in all patients, while its proportional rRNA was identical in health and disease
  • Collagenase expression was dominated by T. forsythia in patient 1, augmented by P. tannerae in patient 2, and by P. gingivalis in patient 3; protease expression followed similar patient-specific patterns
  • Expression dispersion increased with expression level at the individual gene level and for randomly binned genes, but decreased at the EC-binned community level
Key statistics
  • pvalue P = 0.0005 (Paired two-tailed Student t test: Euclidean distance to condition centroid vs to paired sample from same patient)
  • pvalue P = 0.008 (PERMDISP: diseased populations less dispersed than healthy populations)
  • pvalue P = 0.01 (PERMANOVA: health- and disease-associated populations segregate into distinct groups)
  • pvalue P = 0.03 (Paired two-tailed Student t test on Shannon index: health-associated populations more species rich)
  • other ~18% of ~1,100 unique enzyme-encoding gene families differentially expressed (P < 0.05) (Community-level differential EC expression in disease)
  • count 1.5 billion RNA-seq reads total; 28 to 85 million reads mapped to the 60-organism metagenome per sample; 17.3 ± 2.05 million mRNA reads per sample (Metatranscriptome sequencing depth across 9 healthy and 9 diseased populations)
  • other 55 to 65% of total reads aligned to reference databases; >99% of these prokaryotic; 66 to 91% of HOMD/huRNA/virusDB-mapped reads mapped to the 60-species metagenome (Read alignment rates to HOMD, human RNA, and viral databases)
  • count >160,000 bacterial genes in a 60-organism metagenome representing 60 to 90% of total healthy or diseased rRNA; median 12 to 21 reads per mRNA, mean 75 to 156 reads per mRNA; 1,137 random gene bins used as control (Reference metagenome scale, per-gene coverage, and variance-analysis control)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study compared patient-matched healthy and diseased periodontal microbial communities using rRNA-based diversity metrics (10 patients) and RNA-seq-based community gene expression profiling (3 patients, 9 health- and 9 disease-associated samples). Group differences in diversity/distance measures were tested with paired two-tailed Student's t-tests, community composition differences were tested with PERMANOVA and PERMDISP, and differential gene/enzyme expression was analyzed with edgeR, reporting fold changes and dispersion estimates (tagwise and common biological coefficient of variation).

Replicationbiological Sample size10 patients used for rRNA-based diversity analyses (pooled samples from 3 teeth per condition per patient); 3 patients (9 health- and 9 disease-associated samples) used for RNA-seq gene expression analyses. No formal power calculation described. Groupspatient-matched healthy vs. disease-associated (aggressive periodontitis) periodontal microbial communities Pairingpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated in the provided text
Statistical tests used
Test Applied to n Assumptions
paired two-tailed Student's t-test Shannon diversity index, health vs. disease (Fig. 1B) n = 10 patients not stated
paired two-tailed Student's t-test mean Euclidean distance to centroid, health vs. disease (Fig. 1D) n = 10 patients not stated
PERMANOVA beta diversity (unweighted UniFrac) group separation, health vs. disease (Fig. 1C) n = 10 patients not stated
PERMDISP dispersion of beta diversity groups, health vs. disease (Fig. 1C) n = 10 patients not stated
edgeR differential expression analysis EC enzyme-encoding gene family and gene-level expression, disease vs. health (Fig. 2, Fig. 5) 3 patients (9 health- and 9 disease-associated samples) not stated
Approaches that could also have been used
  • Paired comparisons of diversity indices and distances (Fig. 1B, 1D) were made with paired two-tailed Student's t-tests.
    Could also: A paired Wilcoxon signed-rank test — This nonparametric alternative does not assume the differences between paired health/disease values are normally distributed, which can be a useful check with a modest sample size (n = 10).
  • Differential gene/enzyme expression between health and disease was analyzed with edgeR.
    Could also: DESeq2 (Wald or likelihood-ratio test) — DESeq2 uses a related negative-binomial framework with its own dispersion-shrinkage approach and is another widely used option for RNA-seq count data, offering a point of comparison for robustness of the differential expression calls.
  • About 18% of ~1,100 enzyme gene families were called differentially expressed at P < 0.05 without a stated multiple-testing adjustment in this section.
    Could also: Benjamini-Hochberg false discovery rate (FDR) correction — Applying an FDR correction across the full set of simultaneous EC/gene-level tests would help control the expected proportion of false positives among the genes flagged as differentially expressed.
  • Community-level compositional differences between health and disease were assessed with PERMANOVA and PERMDISP.
    Could also: ANOSIM (analysis of similarities) — ANOSIM is another distance-based, non-parametric method for testing group differences in community composition and can serve as a complementary check alongside PERMANOVA/PERMDISP.
  • Variability was summarized as SEM in Fig. 1A but as mean ± SD in Fig. 1C and 1D.
    Could also: Reporting a 95% confidence interval consistently across figures — A CI directly conveys the precision of the estimated mean and allows readers to compare the magnitude of effects across figures using a single, consistent measure of uncertainty.
  • Sample sizes (10 patients for diversity metrics; 3 patients for RNA-seq) were used without a stated power calculation.
    Could also: Including an a priori or post hoc power/sample-size justification — This would give readers additional context on the ability of the chosen cohort size to detect biologically meaningful differences, complementing the patient-matched design already used to control for interindividual variability.
Software: edgeR

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

reference_construction
Reported
Custom 59-organism 'metaorganism' reference metagenome built from HOMD + GenBank genome sequences for the oral community abundant in the sample set (per Jorth et al. 2014 Methods and metarna-seq/build_reference.py).
Reproduced
Rebuilt 59/59 organism reference (33 via NCBI GenBank efetch, 26 via NCBI eutils assembly search substituting for the now-decommissioned legacy HOMD SEQF-ID system). Final merged reference: 145,080 gene features, 0 bad coordinates, 0 malformed GFF3 attributes after fixing two self-introduced parser bugs (continuation-line misrouting corrupting start>end coordinates; unescaped ';' in /product= descriptions breaking GFF3 attribute parsing).
partial
trim_align_count_pipeline
Reported
Flexbar trimming -> Bowtie2 single-end unique alignment (-k 1) -> HTSeq-count (intersection-nonempty, gene-level) per sample, 6 biological samples (3 Healthy, 3 Disease) from 8 SRA runs.
Reproduced
All 6 samples (Healthy1-3, Disease1-3) processed end-to-end successfully on 3rd attempt after fixing the reference GFF bugs. htseq __not_aligned counts internally consistent with bowtie2's reported unaligned-read counts for spot-checked samples. Deviation: htseq-count run with -i ID (self-built GFF uses ID= attribute) instead of original's -i locus_tag.
within tolerance
paired_edgeR_differential_expression
Reported
Paired GLM/LRT differential expression (~Patient+Status design, glmFit+glmLRT, edgeR) between Disease and Healthy samples at gene level; results in SuppFile2.xlsx (161,685 gene rows: Locus ID, per-sample counts, Log2 FC, Log2 CPM, P-value, Product, EC#, Pathways).
Reproduced
Ran identical statistical design (Patient=1:3,1:3 blocking factor, Status Healthy/Disease, glmLRT coef=StatusDisease) on reproduced count matrix (145,069 gene-ID intersection across all 6 samples). Gene-ID overlap with SuppFile2: 83,982/161,685 GT IDs (51.94%). For overlapping IDs: Pearson r=0.4507 for logFC (n=83,922 comparable), 60.53% logFC sign concordance (vs 50% random baseline), raw-count correlations per sample ranged 0.24-0.87 across the 6 samples (weakest for Healthy1/Healthy2, strongest for Healthy3/Disease3).
partial
EC_level_pathway_analysis
Reported
EC-number-aggregated differential expression and KEGG pathway analysis (SuppFile3.xlsx, 1,136 EC-level rows) via the repo's ECcounter.pl/PullEC.pl/HOMD_GenomeMerge.pl EC-annotation pipeline.
Reproduced
NOT attempted. Would require rebuilding a separate EC-annotation pipeline (protein-to-EC mapping via KEGG/UniProt, EC-level count aggregation) not exercised in this reproduction pass given the scope already covered by the gene-level pipeline above. Documented gap, not a drop of the whole RU.
m.public.grade.error

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 62/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

Input data identity is excellent (all 8 SRP033605 runs → the paper's 6 samples), and the statistical layer was reproduced exactly (~Patient+Status paired edgeR glmFit/glmLRT, coef=StatusDisease). The deviation sits entirely upstream in the reference/annotation layer: the legacy HOMD SEQF-ID lookup the published pipeline relies on is decommissioned, so 26/59 organisms were resolved by a substitute NCBI eutils assembly search and 33/59 through a self-written GenBank parser, yielding a reference that matches only 83,982/161,685 (51.94%) of the paper's gene loci and giving moderate agreement (logFC r=0.4507, 60.53% sign concordance, per-sample count r=0.24–0.87). Responsibility is shared — our substitute reference build and inferred Flexbar parameters on one side, the authors' undeposited reference artifact and thin Methods description on the other — but there is no fabrication signal: the directional Healthy-vs-Disease signal replicates. Because the EC/KEGG pathway analysis (SuppFile3, 1,136 rows) was not attempted, the paper's functional conclusions remain untested, which caps q7 and q8 at yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.