MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
MISMATCH (provisional, human-review-flagged). MiGPC ships NO analysis code (repo = MiGPC_MetaData.xlsx with 14 accessions + a README pointing to a Google-Drive OUTPUT catalog), so the pipeline exists only as a Methods description -> reproduced via BRIEF rule 2 by re-running it on the paper's OWN public SRA reads for the 2 smallest Table-3 samples. FRESH FULL RE-RUN this room (self-contained SLURM «job»: env build + ENA download + Trimmomatic 0.39 -> MEGAHIT 1.2.9 with the paper's OWN params incl --kmin-1pass -> Prodigal[MetaGeneMark substitute] -> CD-HIT-EST, all on compute nodes). KEY RESULT: the reproduced assembly yields ~3x FEWER but ~2x LONGER contigs than Table 3 for BOTH samples (PS 181,216 vs 596,279; PET 334,039 vs 1,144,946; max/avg longer; min_len 300 exact; total bases 0.56-0.59x of the paper's implied total). This re-run reproduces the prior run's counts BIT-IDENTICALLY (independent determinism confirmation) and newly captures the PET NR-gene count (567,288 vs 1,632,568). A CONTROL adding --min-count 1 (the strongest benign explanation: keep low-coverage k-mers) gives 242,903 contigs -- still 2.45x below the paper, closing only ~15% of the gap -> benign explanation EXCLUDED. NR genes also mismatch (PS 340,916 / PET 567,288 vs 922,926 / 1,632,568), downstream of the smaller assembly + gene-caller substitution. Not asserting fabrication: either a materially different undocumented assembly/preprocessing, or non-faithfully reported counts; a human adjudicates. NOT attempted (out of scope): full catalog over all 15 samples, MeTarEnz, MAGs, AlphaFold3, logistic regression.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 28assessed: 2026-06-21 ⛓ c23ad4e18da7
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper addresses whether a systematic, genome-resolved metagenomic catalog integrating diverse environmental microbiomes can be built to capture the genetic diversity, functional characteristics, and ecological distribution of enzybiotics (enzyme-based antimicrobials), thereby enabling discovery of novel antimicrobial enzymes and insight into their environmental reservoirs.
- ★ MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics resource
- ★ MiGPC integrates 15 whole-metagenome datasets spanning diverse environments (ocean, soil, gut, vegetation, plastic-contaminated sites) resource
- ★ The catalog contains over 136,000 enzybiotic sequences, 7,654 MAGs, and ~100 million unique genes/proteins finding
- ★ Approximately 62% of genes in the catalog remain functionally uncharacterized finding
- ★ Glycoside hydrolases and glycosyl transferases are the most prevalent CAZyme families in the catalog finding
- ★ Dominant enzybiotic-producing taxa belong primarily to the Pseudomonadota and Bacillota phyla finding
- ★ Statistical/machine learning modeling identified two major ecological clusters distinguishing polluted from relatively pristine environments based on co-abundance of enzybiotic-producing genera finding
- 19 high-confidence reference antimicrobial enzymes spanning hydrolases and oxidoreductases were curated using strict functional and sequence-based criteria method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Metagenomic sequence QC and adapter/quality trimming | 15 environmental whole-metagenome datasets (seawater, soil, gut, vegetation, plastic-contaminated sites, etc.) | none | quality-filtered sequencing reads | FastQC v0.11.8; Trimmomatic v0.39; SOAPaligner2 |
| Metagenome assembly | quality-filtered reads from 15 environmental metagenomes | none | assembled contigs | MEGAHIT v1.2.9 |
| Gene/protein (ORF) prediction | assembled metagenomic contigs | none | predicted open reading frames and protein-coding sequences | MetaGeneMark |
| Non-redundant gene/protein catalog clustering | predicted genes/proteins from all 15 metagenomes | none | unified non-redundant gene/protein catalog (~100 million genes) | CD-HIT v4.8.1 (90% identity threshold) |
| Genome binning | assembled metagenomic contigs | none | metagenome-assembled genomes (MAGs, n=7,654) | — |
| Functional annotation (KEGG Orthology and COG classification) | unified metagenomic gene catalog | none | KO numbers, COG categories, proportion of uncharacterized genes (~62%) | KofamKOALA; eggNOG-mapper v2.0.1 |
| CAZyme domain annotation | unified metagenomic gene catalog | none | CAZy family assignments (e.g., glycoside hydrolases, glycosyl transferases) | run_dbCAN2 (HMMER, DIAMOND, Hotpep) |
| BLAST-based sequence similarity search and environmental/taxonomic mapping | curated enzybiotic sequences vs. NCBI protein database and metagenomic assemblies | none | homologous producer taxa and environmental context assignment of enzybiotic sequences | BLAST; NCBI Taxonomy Browser |
- – Catalog comprises over 136,000 enzybiotic sequences, 7,654 MAGs, and ~100 million unique genes/proteins 136,000+ sequences; 7,654 MAGs; ~100M genes
- – ~62% of predicted genes had no characterized function after KEGG/eggNOG annotation 62%
- – Glycoside hydrolases and glycosyl transferases were identified as the most prevalent CAZyme families
- – Pseudomonadota and Bacillota were identified as the dominant enzybiotic-producing phyla
- – Statistical modeling uncovered two major ecological clusters separating polluted from pristine environments 2 clusters
- – 19 antimicrobial enzymes were selected as a high-confidence reference set spanning hydrolases and oxidoreductases 19 enzymes
- – Redundant sequences were removed via CD-HIT clustering at a 90% identity threshold to build the non-redundant catalog 90% identity threshold
- – 15 metagenomic datasets were selected from an initial pool of 38 candidate environments based on data availability, quality, and ecological distinctiveness 15 of 38 environments
- count over 136,000 enzybiotic sequences (size of curated enzybiotic sequence catalog in MiGPC)
- count 7,654 MAGs (metagenome-assembled genomes recovered across the 15 metagenomes)
- count ~100 million unique genes/proteins (size of the unified non-redundant gene/protein catalog)
- other ~62% (proportion of catalog genes lacking characterized function via KEGG/eggNOG)
- count 19 antimicrobial enzymes (curated high-confidence reference enzybiotic set spanning hydrolases and oxidoreductases)
- count 15 whole-metagenome datasets (environmental metagenomic datasets integrated into MiGPC)
- other 90% identity threshold (CD-HIT clustering threshold used for redundancy removal)
- count 10 bacterial classes (bacterial classes with strong BLAST sequence similarity to known enzybiotic proteins, used to guide environment selection)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational metagenomics catalog paper describing the construction of MiGPC through bioinformatic pipelines: quality control, assembly (MEGAHIT), gene prediction (MetaGeneMark), and sequence-level clustering (CD-HIT at 90% identity) across 15 publicly available whole-metagenome datasets. Functional annotation was performed with KofamKOALA (KEGG), eggNOG-mapper (COGs), and run_dbCAN2 (CAZymes). Results are reported primarily as counts and proportions; the paper mentions 'statistical modeling and machine learning' to identify two ecological clusters, but the specific methods are not detailed in the provided text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| CD-HIT sequence clustering (90% nucleotide/protein identity threshold) | Non-redundant gene and protein catalog construction across all 15 metagenomes; also applied to enzybiotic reference dataset for redundancy removal | ~100 million unique genes/proteins in unified catalog; 136,000+ enzybiotic reference sequences | not stated |
| BLAST (sequence similarity search) | Mapping curated enzybiotic sequences to metagenomic assemblies; identifying bacterial taxa with high homology to known enzybiotic proteins for environment selection | — | not stated |
| Statistical modeling / machine learning for ecological clustering (specific method not stated in available text) | Identification of two major ecological clusters distinguishing polluted from relatively pristine environments based on co-abundance of enzybiotic-producing genera | 15 metagenomic datasets | not stated |
-
Ecological clustering to distinguish polluted from pristine environments was described as 'statistical modeling and machine learning' without specifying the method in the available text↳ Could also: Standard approaches include PCoA on Bray-Curtis or Jaccard dissimilarity matrices, PERMANOVA (e.g., adonis2 in R/vegan) for significance testing of group separation, or UMAP/t-SNE for visualization followed by k-means or hierarchical clustering with a reported cluster quality metric (e.g., silhouette score) — Naming the specific method and reporting a measure of cluster support (e.g., PERMANOVA R² and p-value, or silhouette width) would allow readers to evaluate the strength of the ecological separation and reproduce the analysis independently
-
Non-redundant catalog construction used CD-HIT at a fixed 90% identity threshold↳ Could also: MMseqs2 (linclust or cluster mode) at a comparable identity cutoff could also be used for large-scale sequence clustering — MMseqs2 is substantially faster and more memory-efficient at the scale of ~100 million sequences; benchmarking studies have shown comparable clustering outcomes at matched identity thresholds, making it a common alternative for very large catalogs
-
Metagenome assembly was performed with MEGAHIT↳ Could also: metaSPAdes (SPAdes in meta mode) is another widely used assembler for metagenomes — metaSPAdes sometimes yields longer contigs and higher N50 values, particularly for lower-coverage samples, at the cost of higher memory requirements; reporting an assembly quality metric (e.g., N50, total assembled length) for each dataset would allow comparison across tools
-
Gene prediction was performed with MetaGeneMark on assembled contigs↳ Could also: Prodigal (in meta mode) is another widely benchmarked prokaryotic gene caller for metagenomes — Prodigal is commonly used alongside or instead of MetaGeneMark in large-scale catalog studies; reporting predicted ORF counts from both tools on a subset of samples would characterize sensitivity differences
-
Dataset size and catalog completeness were not formally evaluated; 15 datasets were selected from 38 candidates based on ecological criteria↳ Could also: Rarefaction (accumulation) curves for unique gene or protein cluster counts as a function of the number of datasets added could also be reported — Gene rarefaction curves would allow readers to assess whether catalog content is approaching saturation or whether additional datasets would substantially expand the catalog, helping gauge coverage completeness—a standard practice in large metagenomic reference catalogs
-
Enzybiotic sequences were filtered with multi-database consistency (annotations consistent across at least two databases) and a 90% CD-HIT threshold, but no quantitative sensitivity/specificity evaluation of this curation pipeline was reported↳ Could also: A held-out benchmark against a labeled set of known positives and negatives (e.g., from SwissProt experimentally characterized entries) could also be used to estimate precision and recall of the curation strategy — Reporting precision and recall, or at minimum the number of sequences filtered at each step, would quantify how the curation criteria affect completeness versus specificity of the final enzybiotic reference set
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41888223 (MiGPC: catalog of enzybiotics from environmental metagenomes)
- Paper: Afshar Jahanshahi et al., Scientific Reports 2026;16:15153. DOI 10.1038/s41598-026-44250-9.
- Repo: https://github.com/kkavousi/MiGPC-Catalog — ships NO analysis code.
Only
MiGPC_MetaData.xlsx(14 SRA/ENA accessions) + a README pointing to a Google-Drive folder holding the output catalog (Gene_Catalog.fa, Protein_Catalog.fa, Enzybiotic_Catalog.tsv). The bioinformatic pipeline is described in Methods only (tools + a few exact parameters), not provided as runnable code. - Data: raw reads public on SRA/ENA (PRJNA428417 is only sample #1 of 14+).
Pipeline described in Methods (per the paper)
QC: FastQC v0.11.8, Trimmomatic v0.39 (-phred33 SLIDINGWINDOW:4:20 MINLEN:36).
Assembly: MEGAHIT v1.2.9 (--k-min 27 --k-max 127 --k-step 10 --min-contig-len 300 -t 40).
Gene prediction: MetaGeneMark (gmhmmp -a -m MetaGeneMark_v1.mod).
Clustering: CD-HIT v4.8.1 (-c 0.9 -M 0 -T 0) -> "non-redundant genes".
Annotation: KofamKOALA, eggNOG-mapper v2.0.1, run_dbCAN2. Binning: MetaBAT2, dRep, CheckM.
Enzybiotic screening: MeTarEnz (authors' own tool). Structure: AlphaFold3, TM-align. Stats: logistic regression, GMM.
IN SCOPE (clearly-specified, low-hanging, pipeline-derived)
Per-sample Table 3 statistics, reproduced by running the described QC -> MEGAHIT -> gene-prediction -> CD-HIT pipeline on the paper's own raw reads:
- number of assembled contigs (>300 bp) [pure MEGAHIT output — clean 1:1]
- maximum contig length (bp) [pure MEGAHIT output — clean 1:1]
- non-redundant gene count [gene-caller dependent — see note]
Targets = the 2 smallest samples (fastest honest run):
- SRR28167096 "PS Seawater": reported contigs=596,279, max=261,036 bp, NR genes=922,926
- SRR28167100 "PET_Cow Gut": reported contigs=1,144,946, max=212,414 bp, NR genes=1,632,568
Tool-substitution note (gene prediction)
MetaGeneMark requires a per-user GeneMark license key (web-form + emailed key),
not installable from bioconda. Per BRIEF rule 2 ("a third-party tool on the
paper's data is equally valid") and the 80/20 rule, gene prediction is run with
Prodigal -p meta (the standard open substitute), then CD-HIT-EST -c 0.9.
The gene-count claim is therefore graded conservatively (gene-caller choice
shifts counts by 10–30%); the contig statistics are the clean 1:1 comparison.
OUT OF SCOPE (the hard last ~20% — not attempted, with reason)
- Full ~100 M-gene catalog over all 15 samples (the Plastic catalog alone is 53.3 M genes from a prior study; ~1.2 TB raw reads) — compute-prohibitive and not needed to test reproducibility of the described per-sample method.
- MeTarEnz enzybiotic screening (19,756 putative AMEs) — authors' own tool, downstream of the full catalog; not shipped runnable.
- MAGs (7,654), CAZyme counts, KEGG/eggNOG annotation totals, AlphaFold3 structures, logistic-regression accuracy — each needs the full catalog and/or external tools/databases far beyond the low-hanging assembly check.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Re-running the paper's own documented MEGAHIT pipeline on the authors' own SRA reads reproduces ~3x fewer but ~2x longer contigs than Table 3 for both samples (PS 181,216 vs 596,279; PET 334,039 vs 1,144,946), with the default assembly bit-identical on re-run. The single strong benign explanation (undocumented --min-count 1) was tested and excluded — it closes only ~15% of the gap. Input data is identical and the endpoints are directly comparable, so the deviation sits in the core assembly computation and lands on the authors' side (undocumented steps or non-faithful Table 3 counts), compounded by the repo shipping no analysis code. Not asserting fabrication, but a substantive, human-adjudicated mismatch.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.