Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A case study for large-scale human microbiome analysis using JCVI's metagenomics reports (METAREP).

PLoS One · 2012
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the paper's actual pipeline-derived statistical claims by running METAREP's own R scripts (not reimplementations) on «our HPC»: Metastats differential-abundance detection matched its shipped reference output exactly on all deterministic columns (mean/variance/stderr), with expected small stochastic variation in its permutation-based p/q-values (no random seed set in the original script); the two-sample Fisher's-Exact/Equality-of-Proportions test script reproduced independently-computed R ground-truth p-values exactly; and the paper's Morisita-Horn-distance/average-linkage clustering methodology was applied faithfully (same code, same parameters) to real paper-linked HMP HUMAnN example data, yielding a real, structured distance matrix as a partial/methodological (not numerically exact) reproduction, since the raw-hit-to-clustering-matrix aggregation step normally performed by HUMAnN itself had to be substituted. Figure 6's Solr query-benchmark and the full PHP/Solr/MySQL web-application deployment were correctly scoped out as infrastructure-dependent, not pipeline-derived, claims. All 10 paper-linked SRA datasets (the RU's designated SRS047225 plus the 9 SRS accessions shipped as real examples in the repo) were profiled and confirmed to deliver their promised dual public-16S/controlled-WGS HMP data structure; the repo's unrelated DeLong-et-al. ocean example data was correctly excluded from scope.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-28
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a scalable comparative-metagenomics framework (JCVI METAREP v1.3.1, with dynamic annotation weighting and distributed search) support interactive query and comparison of hundreds of millions of taxonomic and functional annotations from the Human Microbiome Project, and thereby reveal how taxonomic and functional profiles vary across body habitats and individuals? Specific scenarios test that (i) enzymatic processes of pyruvate metabolism and their taxonomic membership vary across body habitats and (ii) although composition varies between individuals and over time, samples from the same body habitat are more similar to one another.

Core claims
  • METAREP version 1.3.1 is an open-source, scalable tool for querying, browsing and comparing extremely large volumes of metagenomic annotations, with an extended data model, dynamic weighting, distributed searches and advanced clustering. resource
  • The dynamic weighting feature scales to over 400 million weighted gene annotations derived from 14 billion HMP short reads as predicted by the HUMAnN pipeline. method
  • METAREP's data model was expanded to directly import and analyze results from two HMP annotation pipelines: JPMAP (assembly ORF annotation) and HUMAnN (short-read annotation). method
  • Clustering patterns of taxonomic abundance derived from functional (enzymatic) genes are not always consistent even for body habitats from similar body regions; physical proximity is not necessarily the best indicator of taxonomic profile similarity. finding
  • Relatively high abundances of Crenarchaeota and Euryarchaeota are associated with skin habitats as determined by the metabolic marker PFOR, which the authors state has not been previously reported using a metabolic marker. finding
  • For each of the three pyruvate-metabolism marker enzymes, the majority of relative abundance falls within only five to six phyla, while many low-abundance lineages (including at least one eukaryotic phylum and less-studied lineages such as Thermotogae and Archaea) represent a reservoir of genetic diversity. finding
  • Cluster topologies for the enzymatic marker analyses were consistent across Euclidean, Bray-Curtis and Morisita-Horn distance metrics. finding
  • Pyruvate metabolism marker enzymes (PDHC, PFOR, PFL) can serve as functional biomarkers whose taxonomic distributions differ by body habitat. mechanism
Experimental setups
Assay System Perturbation Readout Platform
Whole genome shotgun (WGS) metagenomic sequencing, short-read annotation via HUMAnN pipeline Human microbiome; 741 samples from up to fifteen body habitats of 108 healthy adult men and women (HMP) none (observational survey of healthy 'normal' donors) Weighted gene annotations: taxonomy (NCBI), EC, GO, KEGG Orthology, KEGG and MetaCyc pathway abundances HUMAnN (HMP Unified Metabolic Analysis Network) pipeline; METAREP v1.3.1
Assembly-based ORF annotation (WGS metagenome assemblies) Human microbiome body-habitat communities (HMP); over 700 assemblies (705 assembly datasets) none ORF-based functional and taxonomic annotations JCVI Prokaryotic Metagenomics Annotation Pipeline (JPMAP)
Enzymatic marker filtering and comparative abundance analysis (METAREP Compare page) Pooled HUMAnN datasets from 13 body habitats (n = 493 datasets; 97 donors) none (filter queries for PDHC, PFOR, PFL) Absolute and relative weighted annotation counts of marker enzymes across phyla and body habitats METAREP v1.3.1 Compare page (http://www.jcvi.org/hmp-metarep)
Hierarchical clustering and heatmap/dendrogram visualization Pooled body-habitat HUMAnN datasets (13 body habitats) none Dendrogram topologies and relative abundance heatmaps of phyla versus body habitats METAREP; Euclidean, Bray-Curtis and Morisita-Horn distance metrics with average linkage clustering
Multi-dimensional scaling, count summaries and non-parametric statistical tests HMP WGS datasets across body habitats and individuals (oral habitats highlighted) none Differentially abundant taxa and pathways; distance matrices; statistical test results METAREP Compare page (PDF/text export)
Software performance benchmarking of response time METAREP HMP instance hosting HMP WGS annotations distributed search / dynamic weighting configuration Query response time / scalability of weighted frequency calculations METAREP v1.3.1 distributed search architecture
Sample-variation analysis of taxonomic and pathway composition within/between individuals, body habitats and over time (Scenario 2) HMP WGS samples from multiple individuals and body habitats sampled over time none Taxonomic and pathway composition similarity (hierarchical clustering) METAREP v1.3.1
Key results
  • The HMP METAREP instance hosts over 400 million weighted gene annotations predicted from 14 billion short reads by HUMAnN, plus ORF annotations from over 700 assemblies (totals: 498 HUMAnN datasets, 14,613 million reads, 424 million weighted annotations, sum of weights 4,916.0 million, 705 assembly datasets). 424 million weighted annotations from 14,613 million reads
  • PDHC analysis recovered 39 phyla; 94% of total abundance came from five phyla: Actinobacteria, Firmicutes, Proteobacteria, Bacteroidetes and Fusobacteria. 39 phyla; 94% from 5 phyla (29%, 27%, 24%, 12%, 2%)
  • PFL analysis recovered 15 phyla; 97% of total abundance came from six phyla: Firmicutes, Proteobacteria, Bacteroidetes, Actinobacteria, Fusobacteria and Cyanobacteria. 15 phyla; 97% from 6 phyla (50%, 26%, 10%, 7%, 2%, 2%)
  • PFOR analysis recovered 14 phyla; 95% of total abundance came from seven phyla including Firmicutes (27%), Euryarchaeota (25%), Crenarchaeota (20%), Proteobacteria (10%), Thermotogae (9%), Actinobacteria (2%) and Dictyoglomi (2%), showing more variable habitat clustering than PDHC and PFL. 14 phyla; 95% from 7 phyla
  • Left and right retroauricular crease (skin) samples were most distantly related to all other habitats in the PFOR analysis and were dominated by Crenarchaeota. Crenarchaeota 81% (left) and 73% (right)
  • In the PDHC analysis, stool was the most distantly related habitat due to high Bacteroidetes abundance, while anterior nares clustered with retroauricular crease driven by high Actinobacteria. Bacteroidetes 57% in stool; Actinobacteria 59%–84% in nares/skin cluster
  • In the PFL analysis, anterior nares and stool were the most distantly related habitats, separated by high Actinobacteria in anterior nares and high Bacteroidetes in stool despite similar Firmicutes abundance. Actinobacteria 32% (nares, highest); Bacteroidetes 26% (stool, highest); Firmicutes 45% nares vs 46% stool
  • Consistent dendrogram topologies were recovered across all three distance metrics for 13 PFOR, 12 PDHC and 10 PFL filtered body habitats. 13 / 12 / 10 body habitats
Key statistics
  • count over 400 million weighted gene annotations from 14 billion short reads (Table 1 totals: 424 million weighted annotations; 14,613 million reads; 4,916.0 million sum of annotation weights) (Scale of HUMAnN annotations hosted in the HMP METAREP instance)
  • count 741 samples from up to fifteen body habitats of 108 healthy adult men and women; approximately 38 billion short reads (3.5 Tbp) generated, of which over 14 billion were processed and analyzed (HMP WGS metagenomic survey scope)
  • count n = 493 HUMAnN datasets from 13 body habitats; 97 donors (Datasets pooled for the Scenario 1 pyruvate enzyme marker comparison)
  • other PDHC: 39 phyla recovered; 94% of abundance from Actinobacteria (29%), Firmicutes (27%), Proteobacteria (24%), Bacteroidetes (12%), Fusobacteria (2%); remaining 6% across 34 phyla each <1% (Taxonomic distribution of pyruvate dehydrogenase complex)
  • other PFL: 15 phyla; 97% from Firmicutes (50%), Proteobacteria (26%), Bacteroidetes (10%), Actinobacteria (7%), Fusobacteria (2%), Cyanobacteria (2%); remaining 3% across nine phyla each <1% (Taxonomic distribution of pyruvate formate lyase)
  • other PFOR: 14 phyla; 95% from Firmicutes (27%), Euryarchaeota (25%), Crenarchaeota (20%), Proteobacteria (10%), Thermotogae (9%), Actinobacteria (2%), Dictyoglomi (2%); remaining 5% across seven phyla each <1% (Taxonomic distribution of pyruvate:ferredoxin oxidoreductase)
  • other Crenarchaeota 81% (left retroauricular crease) and 73% (right retroauricular crease) (PFOR marker dominance in skin habitats)
  • count Stool: 68 HUMAnN datasets, 6,262 million reads, 78 million weighted annotations, 1,563.8 million weight sum, 151 assembly datasets; Supragingival plaque: 89 / 4,192 / 112 / 1,538.0 / 118 (Table 1, largest body-habitat datasets by read count)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper is primarily a software/methods report describing JCVI's METAREP tool, illustrated with case-study analyses of Human Microbiome Project shotgun metagenomic data across body habitats. Comparisons shown in the provided text rely on exploratory multivariate approaches (hierarchical clustering and heatmaps using multiple distance metrics, with plans for multidimensional scaling), and the introduction states that non-parametric statistical analyses were also used to detect differentially abundant taxa and pathways in oral habitats, though the specific test and its results are described later in the paper, past the point where the provided text ends.

Replicationbiological Sample sizeSample sizes are described as counts of datasets and donors per habitat (e.g., 493 HUMAnN datasets from 97 donors for Scenario 1; per-habitat dataset counts in Table 1) rather than via a formal power or sample-size calculation GroupsBody habitats (up to 13-15 sites), taxonomic phyla, and individuals/time points Pairingunclear Randomization/blindingna Dispersionnone
Statistical tests used
Test Applied to n Assumptions
Hierarchical clustering with multiple distance metrics (Euclidean, Bray-Curtis, Morisita-Horn) Comparison of pyruvate-metabolism enzyme (PDHC, PFOR, PFL) abundance by phylum across 13 pooled body habitats (Figure 2, Figure S1) 493 HUMAnN datasets from 97 donors, pooled across 13 body habitats not stated
Non-parametric statistical test (specific test not named in the provided portion of text) Detection of differentially abundant taxa and pathways in oral habitats (described as a downstream feature/scenario; details appear beyond the text provided) not stated
Approaches that could also have been used
  • Differences in taxonomic/functional composition across body habitats were explored via hierarchical clustering and heatmaps built on several distance metrics (Euclidean, Bray-Curtis, Morisita-Horn), assessed by visual/topological consistency.
    Could also: A permutation-based multivariate test such as PERMANOVA (Adonis) or ANOSIM applied to the same distance matrices — This would provide a formal significance value for whether community composition differs among body habitats or groups, complementing the descriptive clustering already shown.
  • The paper states that non-parametric analyses were used to identify differentially abundant taxa and pathways, without specifying the exact test in the portion of text available.
    Could also: Count/compositional-data-aware methods developed for metagenomic or RNA-seq-like data, such as DESeq2, edgeR, or ALDEx2 — These approaches explicitly model overdispersion, differing library sizes, and compositionality in count data, which can complement a general-purpose non-parametric rank test.
  • Multiple enzymes, taxa, and pathways appear to be compared across many body habitats in parallel (e.g., PDHC, PFOR, PFL across 10-13 habitats and dozens of phyla).
    Could also: Applying a false-discovery-rate procedure such as Benjamini-Hochberg across the full family of comparisons — This would help control the overall false discovery rate when many taxa, pathways, and habitats are examined simultaneously.
  • Consistency of dendrogram topology across distance metrics was assessed qualitatively by visual comparison of clustering results.
    Could also: Quantitative clustering-stability measures such as bootstrap/jackknife resampling (e.g., pvclust) or cophenetic correlation coefficients — These would give a numerical confidence measure for cluster groupings in addition to the visual comparison already performed.
  • Phylum-level abundances within each body habitat are reported as single percentage values (e.g., means or totals) without an accompanying measure of variability across samples or donors.
    Could also: Reporting a dispersion measure such as SD, IQR, or a 95% confidence interval alongside each abundance percentage — This would convey how much abundance varies across donors/samples within a habitat, in addition to the central estimate.
  • Body-habitat and individual relationships were visualized primarily through hierarchical clustering/heatmaps in the portion of text provided.
    Could also: Ordination approaches such as PCoA or NMDS paired with a distance-based significance test (e.g., PERMANOVA) — This would offer a complementary low-dimensional visualization of sample relationships together with a statistical test of group separation.
Software: METAREP 1.3.1 · HUMAnN · JPMAP (JCVI Prokaryotic Metagenomics Annotation Pipeline)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

metastats-differential-abundance
Reported
METAREP's shipped Metastats implementation (scripts/r/metastats/detect_DA_features.r + run_metastats.r), the differential-abundance engine underlying the paper's community comparison features, reproduces its own shipped reference output on the demo human/mouse-gut class matrix (jrw.manmouse.class.matrix).
Reproduced
All 13 taxa rows reproduced in identical order. Columns 1-6 (per-group mean/variance/stderr, deterministic arithmetic) match the reference BIT-FOR-BIT across all 13 rows. Columns 7-8 (p-value, q-value) differ by O(1e-3): these derive from Metastats' permutation-based Monte Carlo t-test (default B=1000 permutations) and Storey-Tibshirani FDR calc downstream of it; the original script contains no set.seed() call, so repeated runs are EXPECTED to differ stochastically by design of the method itself, not due to any implementation discrepancy. This is graded within-tolerance rather than exact/mismatch because the deterministic core of the algorithm is proven correct, while the stochastic component behaves as the method's own specification dictates.
within tolerance
two-way-sample-test-fisher-proportions
Reported
METAREP's two_way_sample_test.r script (implementing the Fisher's Exact Test [METAREP-421] and Equality-of-Proportions Test [METAREP-422] features listed in the repo README's release notes) executes correctly and produces statistically correct p-values, odds ratios and relative risks, with Bonferroni/FDR correction.
Reproduced
Script output p-values (rounded to the requested precision) match the independently computed ground-truth p-values exactly for all 4 features under both test functions (e.g. featA fisher: script 0.10837 vs direct 0.1083696; featC prop: script 0.019631 vs direct 0.01963066). Odds ratio, relative risk, and proportion columns were also manually spot-checked as arithmetically correct.
exact
morisita-horn-clustering-methodology
Reported
The paper's stated sample-clustering methodology (Morisita-Horn distance + average-linkage hierarchical clustering, as implemented in METAREP's scripts/r/plots.r Option 7) was applied to real, paper-linked HMP HUMAnN example data (the 7 SRS samples shipped in data/humann/, explicitly tied to this paper's DOI in the repo's data/humann/README).
Reproduced
This is a methodologically faithful application of the paper's own stated clustering method and code to genuinely paper-linked data, and it produced a real, structured, non-degenerate distance matrix (SRS012291 vs SRS024567 Morisita-Horn distance = 0.00092, near-identical functional profiles; SRS022092 vs SRS022545 = 0.327; SRS052988 and SRS057290 are outliers at 0.98-1.0 from all other samples). It is graded partial rather than exact/within-tol because (a) the gene-family aggregation step from raw HUMAnN per-gene hits to a clustering-ready matrix is normally performed by HUMAnN itself, whose source is not in this repo, so this reproduction substitutes an equivalent-in-spirit but not identical aggregation; and (b) there is no published ground-truth distance matrix or dendrogram from the paper's actual Figures 2/3 to diff against numerically -- only the clustering algorithm/parameters are directly verifiable, not a specific published number.
partial
solr-query-performance-benchmarks-fig6
Reported
Figure 6 style query-performance/scalability benchmarks (ApacheBench load test against a live Solr index at increasing concurrency).
Reproduced
Out of scope / infeasible, not a failed reproduction: the script targets a hardcoded JCVI-internal IP (http://«ip»:8989/solr/Manangatang-Managed/select?...) that is not reachable or reconstructable outside JCVI's infrastructure. Reproducing this claim would require deploying an equivalent multi-million-record Solr index and replicating JCVI's specific hardware/network conditions, which is an infrastructure-benchmarking exercise rather than a portable pipeline-derived computational result. Marked out-of-scope per the brief's own guidance on infeasibility being a valid, non-fabricated outcome.
m.public.grade.error
metarep-web-application-deployment
Reported
The full METAREP web application (PHP/CakePHP UI + Apache Solr/Lucene search + MySQL backend) as an integrated large-scale comparative-metagenomics browsing/search system.
Reproduced
Out of scope: deploying a full multi-service web stack (PHP app server + Solr + MySQL, with JCVI-specific data loading conventions) is an infrastructure/DevOps exercise, not a single bioinformatic-pipeline-derived computational claim. The paper's actual quantitative/statistical claims live in the R analysis scripts (Metastats, two-way tests, clustering), which were reproduced directly and are covered by the claims above; the web UI is a visualization/access layer over those same computations.
m.public.grade.error

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is a software/case-study paper, and the reproduction is honest about what that permits: the deterministic core of METAREP's statistics engine reproduces its own shipped reference bit-for-bit (all 13 taxa rows, columns 1-6 of Routput.diffAb), and two_way_sample_test.r matches independently computed R ground truth (featA fisher 0.10837 vs 0.1083696). The only numeric divergence — O(1e-3) in the p/q columns — is fully explained by an unseeded B=1000 permutation test, i.e. the method's own specification, so severity is negligible. The real limitation is on our/scope side, not the authors' side: the paper publishes dendrograms (Fig.2/3) and latency curves (Fig.6) rather than tabulated values, the HUMAnN aggregation step is not in the repo so we substituted our own, the Fisher/proportions test ran on synthetic input, and two claims (Solr benchmark against hardcoded «ip»:8989, full web-stack deployment) are genuinely infeasible outside JCVI. Nothing here is fabrication-suspect; the correct reading is solid but partial coverage of a paper whose central claims are largely non-numeric, hence yellow on derivability, core claim and overall.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.