metaGEM: reconstruction of genome scale metabolic models directly from metagenomes.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
metaGEM (NAR 2021) reconstructs GEMs from metagenomes. PRIMARY CLAIM REPRODUCES EXACTLY: the Zenodo deposit (10.5281/zenodo.4407746, md5-verified) contains exactly 14,087 SBML models = the paper's claimed 14,087 GEMs (ocean 9365/gut 4127/soil 269/plant 172/lab 154). FBA-readiness confirmed: 30/30 stratified-sampled models load in cobrapy and grow under FBA. Lab-environment count 154 == reported 154 MQ MAGs. Full-pipeline MAG-quality claims (3750 HQ / 10349 MQ / 483 samples) are out of scope (require reprocessing TBs of raw metagenomes + GTDB-Tk; not derivable from the deposit). CarveMe reconstruction step attempted as secondary demonstration.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetMetabolic modeling of microbial communities traditionally relies on reference genomes, which fails to capture intra- and inter-species genetic diversity; the paper tests whether genome-scale metabolic models (GEMs) can instead be reconstructed directly from metagenome-assembled genomes (MAGs) to better represent community metabolism and disease-associated metabolic exchanges.
- ★ metaGEM enables end-to-end reconstruction of FBA-ready GEMs directly from metagenomes without relying on reference genomes method
- ★ GEMs reconstructed from metagenomes have fully represented metabolism comparable to isolated reference genomes finding
- ★ metaGEM-derived GEMs capture intraspecies metabolic diversity not accessible via reference-genome-based modeling finding
- ★ Community-level metabolic exchange modeling identifies potential differences in the progression of type 2 diabetes finding
- ★ metaGEM provides a resource of >14,000 MAGs and corresponding GEMs across diverse biomes resource
- metaGEM consistently recovers more high-quality genomes per sample than other MAG reconstruction pipelines finding
- ★ Reference-genome-based modeling can cause false positive/negative pathway predictions due to unrepresented genetic diversity mechanism
- A mapping-based abundance estimation method (no marker genes/reference genomes) correlates strongly with marker-gene-based estimates finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| MAG binning and refinement | 5 metagenomic datasets (lab culture, human gut, plant-associated, soil, ocean) | none | bin completeness and contamination (quality-filtered MAG sets) | CONCOCT, MetaBAT2, MaxBin2, metaWRAP |
| Genome-scale metabolic model (GEM) reconstruction | prokaryotic MAGs derived from metagenomes | none | FBA-ready metabolic reaction networks | CarveMe v1.2.2 with CPLEX solver v12.8 |
| GEM quality assessment | reconstructed genome-scale metabolic models | none | model quality report metrics | MEMOTE v0.9.13 |
| Community-level flux balance analysis (FBA) simulation | gut microbiome GEM communities | simulation media (M3 minus aromatic amino acids, M11) | essential metabolic exchanges sustaining non-zero growth of all community members | SMETANA v1.2.0 with CPLEX solver |
| MAG abundance quantification (mapping-based vs marker-gene-based) | lab culture MAGs (7 gut microbiome species) | none | relative and absolute MAG abundance | bwa/SAMtools mapping and mOTUs2 |
| Taxonomic classification | reconstructed MAGs | none | taxonomic labels | GTDB-Tk v1.1.0 |
| Controlled multispecies lab culture metagenomics | 7 human gut microbiome species grown in vitro, 4 biological replicates x 12 time points (48 samples) | none (time-course growth) | MAG reconstruction quality and abundance dynamics over time | short-read shotgun sequencing |
| MAG growth rate estimation | medium- and high-coverage MAGs | none | in situ growth rate | GRiD v1.3 |
- ▲ Reconstructed over 14,000 MAGs across 5 metagenomic biomes, including 3750 high-quality (>90% completeness, <5% contamination) and 10,349 medium-quality (>50% completeness, <10% contamination) genomes 14,000+ MAGs
- ▲ Lab culture validation reconstructed 154 medium-quality MAGs (137 also high-quality) with high average completeness and low contamination, ~3.2 MAGs per sample reflecting known growth dynamics 95.4% avg completeness, 0.3% avg contamination
- ▲ Mapping-based MAG abundance estimates were highly correlated with marker-gene-based (mOTUs2) abundance estimates r=0.99, P<1e-16
- ▲ metaGEM recovered more high-quality genomes per sample than comparable pipelines in gut and ocean metagenome comparisons
- – GEMs reconstructed from metagenomic MAGs showed fully represented metabolism comparable to GEMs from isolated genomes
- – Metagenomic GEMs captured intraspecies metabolic diversity within communities
- – Community-level SMETANA simulations identified differences in gut bacterial metabolic exchange patterns associated with type 2 diabetes progression
- correlation r=0.99, P<1e-16 (mapping-based vs marker-gene-based (mOTUs2) MAG abundance estimates in lab culture dataset)
- count 3750 high-quality MAGs (>90% completeness, <5% contamination across all datasets)
- count 10,349 medium-quality MAGs (>50% completeness, <10% contamination across all datasets)
- mean 95.4% average completeness, 0.3% average contamination (medium-quality MAGs reconstructed from lab culture dataset)
- mean ~3.2 MAGs per sample (average MAG yield per sample in lab culture dataset)
- count 483 whole metagenome shotgun samples (total samples analyzed from 5 metagenomic studies)
- other <10% contamination, >50% completeness thresholds (metaWRAP bin_refinement medium-quality bin criteria (parameters -x 10 -c 50))
- count 154 MQ MAGs, of which 137 also met HQ standard (lab culture dataset MAG reconstruction counts)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is primarily a computational pipeline (tool) paper describing metaGEM, an end-to-end workflow for reconstructing genome-scale metabolic models from metagenomes. The paper's main outputs are descriptive/benchmarking metrics (MAG completeness, contamination, and abundance estimates) rather than formal hypothesis testing across experimental groups. The one explicit inferential statistic in the provided text is a Pearson correlation comparing two abundance-estimation methods, and pipeline performance relative to other tools is otherwise described narratively/visually (e.g., a 'simple comparison' referencing a supplementary figure).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson correlation | Comparison of mapping-based vs. marker-gene-based (mOTUs2) MAG abundance estimates in the lab culture dataset (Figure 2A) | 48 metagenomic samples (7 species, 4 biological replicates, 12 time points) | not stated |
-
Agreement between mapping-based and marker-gene-based (mOTUs2) abundance estimates was assessed with a Pearson correlation (r = 0.99).↳ Could also: Spearman's rank correlation or Lin's concordance correlation coefficient — These would also be standard choices for compositional/abundance data that may not meet linear-correlation assumptions, and CCC in particular directly assesses agreement (rather than just linear association) between two measurement methods.
-
The correlation p-value is reported as a bound (P-value < 1e-16) rather than an exact value, with no confidence interval on r.↳ Could also: Reporting the exact computed p-value alongside a confidence interval for r — This would convey the precision of the estimated association strength in addition to its statistical significance.
-
MAG completeness and contamination are summarized as averages (95.4% and 0.3%) without an accompanying measure of spread.↳ Could also: Reporting SD, IQR, or range alongside the averages — This would communicate how much variability existed across the reconstructed MAGs, which is useful when averages are based on samples of differing quality or depth.
-
The comparison of MAG recovery against other published pipelines/studies (Supplementary Figure S3) is described as a 'simple comparison' without a stated formal test.↳ Could also: A formal statistical comparison such as a paired or mixed-effects model with pipeline/study as a fixed or random effect — This could quantify differences in genome recovery per sample with an associated uncertainty estimate, and would account for heterogeneity arising from different source studies.
-
Abundance correlation pooled all samples across the four biological replicates and 12 time points into a single Pearson correlation.↳ Could also: A repeated-measures correlation (e.g., rmcorr) or mixed-effects model with replicate as a random effect — Given the repeated-measures structure (multiple time points per replicate), this approach could account for non-independence among samples from the same biological replicate.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.