A case study for large-scale human microbiome analysis using JCVI's metagenomics reports (METAREP).
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the paper's actual pipeline-derived statistical claims by running METAREP's own R scripts (not reimplementations) on «our HPC»: Metastats differential-abundance detection matched its shipped reference output exactly on all deterministic columns (mean/variance/stderr), with expected small stochastic variation in its permutation-based p/q-values (no random seed set in the original script); the two-sample Fisher's-Exact/Equality-of-Proportions test script reproduced independently-computed R ground-truth p-values exactly; and the paper's Morisita-Horn-distance/average-linkage clustering methodology was applied faithfully (same code, same parameters) to real paper-linked HMP HUMAnN example data, yielding a real, structured distance matrix as a partial/methodological (not numerically exact) reproduction, since the raw-hit-to-clustering-matrix aggregation step normally performed by HUMAnN itself had to be substituted. Figure 6's Solr query-benchmark and the full PHP/Solr/MySQL web-application deployment were correctly scoped out as infrastructure-dependent, not pipeline-derived, claims. All 10 paper-linked SRA datasets (the RU's designated SRS047225 plus the 9 SRS accessions shipped as real examples in the repo) were profiled and confirmed to deliver their promised dual public-16S/controlled-WGS HMP data structure; the repo's unrelated DeLong-et-al. ocean example data was correctly excluded from scope.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-28
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a scalable comparative-metagenomics framework (JCVI METAREP v1.3.1, with dynamic annotation weighting and distributed search) support interactive query and comparison of hundreds of millions of taxonomic and functional annotations from the Human Microbiome Project, and thereby reveal how taxonomic and functional profiles vary across body habitats and individuals? Specific scenarios test that (i) enzymatic processes of pyruvate metabolism and their taxonomic membership vary across body habitats and (ii) although composition varies between individuals and over time, samples from the same body habitat are more similar to one another.
- ★ METAREP version 1.3.1 is an open-source, scalable tool for querying, browsing and comparing extremely large volumes of metagenomic annotations, with an extended data model, dynamic weighting, distributed searches and advanced clustering. resource
- ★ The dynamic weighting feature scales to over 400 million weighted gene annotations derived from 14 billion HMP short reads as predicted by the HUMAnN pipeline. method
- ★ METAREP's data model was expanded to directly import and analyze results from two HMP annotation pipelines: JPMAP (assembly ORF annotation) and HUMAnN (short-read annotation). method
- ★ Clustering patterns of taxonomic abundance derived from functional (enzymatic) genes are not always consistent even for body habitats from similar body regions; physical proximity is not necessarily the best indicator of taxonomic profile similarity. finding
- ★ Relatively high abundances of Crenarchaeota and Euryarchaeota are associated with skin habitats as determined by the metabolic marker PFOR, which the authors state has not been previously reported using a metabolic marker. finding
- ★ For each of the three pyruvate-metabolism marker enzymes, the majority of relative abundance falls within only five to six phyla, while many low-abundance lineages (including at least one eukaryotic phylum and less-studied lineages such as Thermotogae and Archaea) represent a reservoir of genetic diversity. finding
- Cluster topologies for the enzymatic marker analyses were consistent across Euclidean, Bray-Curtis and Morisita-Horn distance metrics. finding
- Pyruvate metabolism marker enzymes (PDHC, PFOR, PFL) can serve as functional biomarkers whose taxonomic distributions differ by body habitat. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole genome shotgun (WGS) metagenomic sequencing, short-read annotation via HUMAnN pipeline | Human microbiome; 741 samples from up to fifteen body habitats of 108 healthy adult men and women (HMP) | none (observational survey of healthy 'normal' donors) | Weighted gene annotations: taxonomy (NCBI), EC, GO, KEGG Orthology, KEGG and MetaCyc pathway abundances | HUMAnN (HMP Unified Metabolic Analysis Network) pipeline; METAREP v1.3.1 |
| Assembly-based ORF annotation (WGS metagenome assemblies) | Human microbiome body-habitat communities (HMP); over 700 assemblies (705 assembly datasets) | none | ORF-based functional and taxonomic annotations | JCVI Prokaryotic Metagenomics Annotation Pipeline (JPMAP) |
| Enzymatic marker filtering and comparative abundance analysis (METAREP Compare page) | Pooled HUMAnN datasets from 13 body habitats (n = 493 datasets; 97 donors) | none (filter queries for PDHC, PFOR, PFL) | Absolute and relative weighted annotation counts of marker enzymes across phyla and body habitats | METAREP v1.3.1 Compare page (http://www.jcvi.org/hmp-metarep) |
| Hierarchical clustering and heatmap/dendrogram visualization | Pooled body-habitat HUMAnN datasets (13 body habitats) | none | Dendrogram topologies and relative abundance heatmaps of phyla versus body habitats | METAREP; Euclidean, Bray-Curtis and Morisita-Horn distance metrics with average linkage clustering |
| Multi-dimensional scaling, count summaries and non-parametric statistical tests | HMP WGS datasets across body habitats and individuals (oral habitats highlighted) | none | Differentially abundant taxa and pathways; distance matrices; statistical test results | METAREP Compare page (PDF/text export) |
| Software performance benchmarking of response time | METAREP HMP instance hosting HMP WGS annotations | distributed search / dynamic weighting configuration | Query response time / scalability of weighted frequency calculations | METAREP v1.3.1 distributed search architecture |
| Sample-variation analysis of taxonomic and pathway composition within/between individuals, body habitats and over time (Scenario 2) | HMP WGS samples from multiple individuals and body habitats sampled over time | none | Taxonomic and pathway composition similarity (hierarchical clustering) | METAREP v1.3.1 |
- – The HMP METAREP instance hosts over 400 million weighted gene annotations predicted from 14 billion short reads by HUMAnN, plus ORF annotations from over 700 assemblies (totals: 498 HUMAnN datasets, 14,613 million reads, 424 million weighted annotations, sum of weights 4,916.0 million, 705 assembly datasets). 424 million weighted annotations from 14,613 million reads
- – PDHC analysis recovered 39 phyla; 94% of total abundance came from five phyla: Actinobacteria, Firmicutes, Proteobacteria, Bacteroidetes and Fusobacteria. 39 phyla; 94% from 5 phyla (29%, 27%, 24%, 12%, 2%)
- – PFL analysis recovered 15 phyla; 97% of total abundance came from six phyla: Firmicutes, Proteobacteria, Bacteroidetes, Actinobacteria, Fusobacteria and Cyanobacteria. 15 phyla; 97% from 6 phyla (50%, 26%, 10%, 7%, 2%, 2%)
- – PFOR analysis recovered 14 phyla; 95% of total abundance came from seven phyla including Firmicutes (27%), Euryarchaeota (25%), Crenarchaeota (20%), Proteobacteria (10%), Thermotogae (9%), Actinobacteria (2%) and Dictyoglomi (2%), showing more variable habitat clustering than PDHC and PFL. 14 phyla; 95% from 7 phyla
- ▲ Left and right retroauricular crease (skin) samples were most distantly related to all other habitats in the PFOR analysis and were dominated by Crenarchaeota. Crenarchaeota 81% (left) and 73% (right)
- – In the PDHC analysis, stool was the most distantly related habitat due to high Bacteroidetes abundance, while anterior nares clustered with retroauricular crease driven by high Actinobacteria. Bacteroidetes 57% in stool; Actinobacteria 59%–84% in nares/skin cluster
- – In the PFL analysis, anterior nares and stool were the most distantly related habitats, separated by high Actinobacteria in anterior nares and high Bacteroidetes in stool despite similar Firmicutes abundance. Actinobacteria 32% (nares, highest); Bacteroidetes 26% (stool, highest); Firmicutes 45% nares vs 46% stool
- – Consistent dendrogram topologies were recovered across all three distance metrics for 13 PFOR, 12 PDHC and 10 PFL filtered body habitats. 13 / 12 / 10 body habitats
- count over 400 million weighted gene annotations from 14 billion short reads (Table 1 totals: 424 million weighted annotations; 14,613 million reads; 4,916.0 million sum of annotation weights) (Scale of HUMAnN annotations hosted in the HMP METAREP instance)
- count 741 samples from up to fifteen body habitats of 108 healthy adult men and women; approximately 38 billion short reads (3.5 Tbp) generated, of which over 14 billion were processed and analyzed (HMP WGS metagenomic survey scope)
- count n = 493 HUMAnN datasets from 13 body habitats; 97 donors (Datasets pooled for the Scenario 1 pyruvate enzyme marker comparison)
- other PDHC: 39 phyla recovered; 94% of abundance from Actinobacteria (29%), Firmicutes (27%), Proteobacteria (24%), Bacteroidetes (12%), Fusobacteria (2%); remaining 6% across 34 phyla each <1% (Taxonomic distribution of pyruvate dehydrogenase complex)
- other PFL: 15 phyla; 97% from Firmicutes (50%), Proteobacteria (26%), Bacteroidetes (10%), Actinobacteria (7%), Fusobacteria (2%), Cyanobacteria (2%); remaining 3% across nine phyla each <1% (Taxonomic distribution of pyruvate formate lyase)
- other PFOR: 14 phyla; 95% from Firmicutes (27%), Euryarchaeota (25%), Crenarchaeota (20%), Proteobacteria (10%), Thermotogae (9%), Actinobacteria (2%), Dictyoglomi (2%); remaining 5% across seven phyla each <1% (Taxonomic distribution of pyruvate:ferredoxin oxidoreductase)
- other Crenarchaeota 81% (left retroauricular crease) and 73% (right retroauricular crease) (PFOR marker dominance in skin habitats)
- count Stool: 68 HUMAnN datasets, 6,262 million reads, 78 million weighted annotations, 1,563.8 million weight sum, 151 assembly datasets; Supragingival plaque: 89 / 4,192 / 112 / 1,538.0 / 118 (Table 1, largest body-habitat datasets by read count)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper is primarily a software/methods report describing JCVI's METAREP tool, illustrated with case-study analyses of Human Microbiome Project shotgun metagenomic data across body habitats. Comparisons shown in the provided text rely on exploratory multivariate approaches (hierarchical clustering and heatmaps using multiple distance metrics, with plans for multidimensional scaling), and the introduction states that non-parametric statistical analyses were also used to detect differentially abundant taxa and pathways in oral habitats, though the specific test and its results are described later in the paper, past the point where the provided text ends.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Hierarchical clustering with multiple distance metrics (Euclidean, Bray-Curtis, Morisita-Horn) | Comparison of pyruvate-metabolism enzyme (PDHC, PFOR, PFL) abundance by phylum across 13 pooled body habitats (Figure 2, Figure S1) | 493 HUMAnN datasets from 97 donors, pooled across 13 body habitats | not stated |
| Non-parametric statistical test (specific test not named in the provided portion of text) | Detection of differentially abundant taxa and pathways in oral habitats (described as a downstream feature/scenario; details appear beyond the text provided) | — | not stated |
-
Differences in taxonomic/functional composition across body habitats were explored via hierarchical clustering and heatmaps built on several distance metrics (Euclidean, Bray-Curtis, Morisita-Horn), assessed by visual/topological consistency.↳ Could also: A permutation-based multivariate test such as PERMANOVA (Adonis) or ANOSIM applied to the same distance matrices — This would provide a formal significance value for whether community composition differs among body habitats or groups, complementing the descriptive clustering already shown.
-
The paper states that non-parametric analyses were used to identify differentially abundant taxa and pathways, without specifying the exact test in the portion of text available.↳ Could also: Count/compositional-data-aware methods developed for metagenomic or RNA-seq-like data, such as DESeq2, edgeR, or ALDEx2 — These approaches explicitly model overdispersion, differing library sizes, and compositionality in count data, which can complement a general-purpose non-parametric rank test.
-
Multiple enzymes, taxa, and pathways appear to be compared across many body habitats in parallel (e.g., PDHC, PFOR, PFL across 10-13 habitats and dozens of phyla).↳ Could also: Applying a false-discovery-rate procedure such as Benjamini-Hochberg across the full family of comparisons — This would help control the overall false discovery rate when many taxa, pathways, and habitats are examined simultaneously.
-
Consistency of dendrogram topology across distance metrics was assessed qualitatively by visual comparison of clustering results.↳ Could also: Quantitative clustering-stability measures such as bootstrap/jackknife resampling (e.g., pvclust) or cophenetic correlation coefficients — These would give a numerical confidence measure for cluster groupings in addition to the visual comparison already performed.
-
Phylum-level abundances within each body habitat are reported as single percentage values (e.g., means or totals) without an accompanying measure of variability across samples or donors.↳ Could also: Reporting a dispersion measure such as SD, IQR, or a 95% confidence interval alongside each abundance percentage — This would convey how much abundance varies across donors/samples within a habitat, in addition to the central estimate.
-
Body-habitat and individual relationships were visualized primarily through hierarchical clustering/heatmaps in the portion of text provided.↳ Could also: Ordination approaches such as PCoA or NMDS paired with a distance-based significance test (e.g., PERMANOVA) — This would offer a complementary low-dimensional visualization of sample relationships together with a statistical test of group separation.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a software/case-study paper, and the reproduction is honest about what that permits: the deterministic core of METAREP's statistics engine reproduces its own shipped reference bit-for-bit (all 13 taxa rows, columns 1-6 of Routput.diffAb), and two_way_sample_test.r matches independently computed R ground truth (featA fisher 0.10837 vs 0.1083696). The only numeric divergence — O(1e-3) in the p/q columns — is fully explained by an unseeded B=1000 permutation test, i.e. the method's own specification, so severity is negligible. The real limitation is on our/scope side, not the authors' side: the paper publishes dendrograms (Fig.2/3) and latency curves (Fig.6) rather than tabulated values, the HUMAnN aggregation step is not in the repo so we substituted our own, the Fisher/proportions test ran on synthetic input, and two claims (Solr benchmark against hardcoded «ip»:8989, full web-stack deployment) are genuinely infeasible outside JCVI. Nothing here is fabrication-suspect; the correct reading is solid but partial coverage of a paper whose central claims are largely non-numeric, hence yellow on derivability, core claim and overall.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.