Unlocking the microbial studies through computational approaches: how far have we reached?
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
▸Reproduction agent’s raw note
DROP (non_pipeline). PMID 36920617 ('Unlocking the microbial studies through computational approaches: how far have we reached?', Kumar/Yadav/Kuddus/Ashraf/Singh, Environ Sci Pollut Res 2023) is a narrative REVIEW article. Independently re-verified this run (2026-07-15): PubMed lists publication type = 'Review'; the abstract and full text synthesize existing literature (metagenomics, in-silico protein/DNA interaction, ML/DL, molecular docking) with NO Materials & Methods section, NO original datasets, and NO author-generated result tables/benchmark numbers; the Data-availability statement is verbatim 'NA'. The RU's auto-extracted links are text-mining false positives: 'code' github.com/hyattpd/prodigal = Prodigal (Doug Hyatt's gene-prediction tool, only mentioned, not the authors' code) and 'data' zenodo:10.5281/zenodo.3678129 resolves (verified this run) to BIOCOM-PIPE v1.19 (2020-02-21) by Djemiel/Dequiedt/Mondy et al. at INRAE — an unrelated 16S/18S/23S metabarcoding pipeline by a different group, not data this review deposited or analyzed. There is therefore no pipeline-derived result to reproduce and no dataset the paper relies on; no «our HPC» compute was required or run. NOT ATTEMPTED: any reproduction, because none is defined. Honest, well-founded drop — nothing forced or fabricated.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-19 ⛓ 06abac1a1f84
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe review addresses how far in silico/computational approaches (metagenomics, machine learning, deep learning) have advanced the study of microbial genomics, proteomics, functional diversity, vaccine development, and drug design.
- ★ Metagenomics enables culture-independent study of microbial communities directly from their natural environments, bypassing the need for clonal isolation. method
- ★ Machine learning and deep learning enable analysis of large-scale sequencing and proteomics datasets to reveal functional and taxonomic diversity of microorganisms. method
- ★ De novo assembly followed by BLAST is the fastest and most reliable method for NGS-based identification of pathogens from 16S–23S rRNA data compared to OTU clustering and mapping. finding
- ★ MARVEL, a random forest-based tool, predicts double-stranded DNA bacteriophage sequences in metagenomic data using a reduced set of informative features. method
- ★ Variation boosting ML models trained on TCGA whole-genome/transcriptome data can discriminate cancer types and distinguish cancerous from normal tissue using microbial signatures. finding
- ★ A random forest model built on vaginal microbiome 16S rRNA data can predict cervical intraepithelial neoplasia (CIN) staging using bacterial marker species. finding
- ★ CASTOR, a machine learning-based virus classification tool simulating RFLP in silico, accurately classifies HPV, HBV, and HIV-1 subtypes/genotypes. method
- ★ DeepARG deep learning networks (DeepARG-LS and DeepARG-SS) accurately predict antibiotic resistance genes from metagenomic sequence data. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| NGS of 16S–23S rRNA regions with de novo assembly/BLAST vs. OTU clustering vs. mapping | patient clinical samples (bacterial pathogens) | none | turnaround time and diagnostic sensitivity | — |
| Random forest classification (MARVEL) of genomic sequences | bacteriophage and bacterial genomes (metagenomic dataset) | none | identification of phage sequences via selected genomic features | — |
| Whole-genome and transcriptome sequencing with variation boosting ML models | human tissue samples, 33 cancer types (TCGA) | cancer vs. normal tissue comparison | discrimination of cancer type/stage via microbial signatures | The Cancer Genome Atlas (TCGA) |
| 16S rRNA (V3 region) sequencing with random forest modeling | vaginal swabs from 66 human subjects | none | bacterial species differentiating CIN1 vs. CIN2 groups | — |
| Machine learning-based virus classification (CASTOR, in silico RFLP simulation) | HBV, HPV, and HIV-1 genomic datasets | none | classification accuracy for subtyping/genotyping | — |
| k-mer signature-based classification (VirFinder) | viral and host genomes (metagenomic data) | none | identification of viral sequences, true positive rate | — |
| Deep learning classification (DeepARG-LS/DeepARG-SS) | metagenomic sequence data (30 ARG categories) | none | prediction accuracy and recall for antibiotic resistance genes | — |
| In silico genomic analysis (ML-based gene identification) | human samples, rhinovirus infection | rhinovirus infection vs. noninfected state | OTOF and SOCS1 gene expression levels | — |
- ▲ De novo assembly followed by BLAST had the shortest turnaround time and highest sensitivity among NGS analysis methods for pathogen identification 2 h 5 min turnaround; 80% sensitivity
- – MARVEL narrowed six candidate features down to three that best identify bacteriophage sequences via random forest
- – ML models trained on TCGA data successfully discriminated different cancer types and distinguished cancer from normal tissue using microbiome signatures
- ▲ 33 bacterial marker species differentiated CIN1 from CIN2 groups using a random forest model AUC = 0.952
- ▲ CASTOR correctly classified HPV, HBV, and HIV-1 subtypes/genotypes with high accuracy 99% (HPV alpha), 99% (HBV genotyping), 98% (HIV-1 M subtyping)
- ▲ DeepARG-LS and DeepARG-SS showed high accuracy and recall in predicting antibiotic resistance genes across 30 ARG categories >0.97 accuracy; >0.90 recall
- ▲ VirFinder showed better true positive rates than the existing tool VirSorter for viral sequence identification, particularly for small contigs
- – OTOF and SOCS1 expression levels could distinguish rhinovirus-infected from noninfected individuals
- other 80% sensitivity (de novo assembly + BLAST for NGS 16S–23S rRNA pathogen identification)
- other 2 h 5 min turnaround time (de novo assembly + BLAST NGS analysis method)
- other AUC = 0.952 (random forest model predicting CIN staging from vaginal microbiome)
- other 99% accuracy (CASTOR classification of HPV alpha species)
- other 99% accuracy (CASTOR classification of HBV genotypes)
- other 98% accuracy (CASTOR classification of HIV-1 M subtypes)
- other >0.97 accuracy (DeepARG models predicting antibiotic resistance genes)
- other >0.90 recall (DeepARG models predicting antibiotic resistance genes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a narrative review article summarizing computational and in silico approaches (bioinformatics tools, machine learning, and deep learning methods) used in microbial genomics, proteomics, and disease prediction across multiple previously published studies. The paper itself does not report original experimental data, a study design, or its own statistical analysis; instead it describes methods and reported performance metrics (e.g., accuracy, recall, AUC, turnaround time) from the cited works.
-
The review reports point performance metrics from cited studies (e.g., accuracy, recall, AUC, sensitivity, turnaround time) without accompanying measures of variability or uncertainty.↳ Could also: Reporting these metrics alongside confidence intervals or standard errors, where available in the original studies — Including a measure of uncertainty around performance metrics like AUC or accuracy would help readers gauge how precisely those estimates were determined, particularly for models trained or tested on smaller datasets.
-
Comparative performance claims between methods (e.g., de novo assembly plus BLAST vs. OTU clustering vs. mapping in Peker et al., or VirFinder vs. VirSorter) are summarized as differences in speed or true positive rate without formal statistical comparison being described.↳ Could also: A formal statistical test (e.g., McNemar's test for paired classifier comparisons, or a paired t-test/Wilcoxon signed-rank test for runtime differences) applied within the original studies — Such tests would allow a quantified assessment of whether observed differences between methods exceed what might be expected by chance, complementing the descriptive comparison presented.
-
Machine learning model performance (e.g., random forest models in Lee et al. and Poore et al., MARVEL's random forest classifier) is described via single accuracy/AUC figures from the original studies.↳ Could also: Cross-validation with reported variance across folds, or bootstrapped confidence intervals for AUC — This approach is commonly used alongside single-split performance metrics to characterize how consistent model performance is across different data partitions, which is especially informative for modest sample sizes such as the 66 subjects in the cervical neoplasia study.
-
Several cited studies (e.g., DeepARG, CASTOR) report classification accuracy across multiple categories or genotypes as summary percentages.↳ Could also: A confusion matrix with per-class precision/recall/F1 and multiple-comparison-adjusted significance testing across categories — This would let readers assess whether performance is uniform across all classes/genotypes or driven by a subset, complementing the aggregate accuracy figures reported.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope analysis — PMID 36920617
Title: Unlocking the microbial studies through computational approaches: how far have we reached? Journal: Environmental Science and Pollution Research (Springer), April 2023 DOI: 10.1007/s11356-023-26220-0 · PMCID: PMC10016191
Verdict: OUT OF SCOPE — review article, no pipeline-derived results (DROP: non_pipeline)
Evidence
- Article type: Narrative review article. PubMed and the PMC full text both classify it as a literature review/survey, "synthesizing computational microbiology developments rather than presenting original experimental data."
- Methods section: None. The PMC full text has no Materials & Methods heading or any equivalent describing original procedures or data generation.
- Data availability statement (verbatim from the published article):
"NA". - Original data / analysis: The authors generated no datasets, ran no pipeline, and report no result tables, benchmark numbers, or reproducible computational outputs. The paper catalogs other researchers' tools (Tables 1–2), NGS workflow figures, ML/DL applications, and molecular-docking concepts — all citing third-party studies.
The auto-extracted "code" and "data" links are text-mining false positives
The RU was spawned with:
- Code:
https://github.com/hyattpd/prodigal— Prodigal is a gene-prediction tool merely mentioned in the review. It is not the authors' code and produces no result the paper reports. - Data:
zenodo:10.5281/zenodo.3678129— this resolves to BIOCOM-PIPE (Djemiel, Dequiedt, Ranjard et al., INRAE UMR Agroécologie; v1.19, 2020-02-21), a standalone 16S/18S/23S metabarcoding pipeline release by a different group. It is an unrelated software deposit that happens to be cited/indexed; it is not data the review analyzed or deposited.
Both links are tool/citation mentions extracted by the screening text-miner, not the paper's own reproducible artifacts. There is no author dataset, accession, or pipeline output to reproduce or compare against.
In-scope pipeline results
None. Nothing in this paper is pipeline-derived in the author's own hands.
Out-of-scope (everything in the paper)
The entire content — descriptions of metagenomics, NGS, in-silico protein structure, docking, ML/DL — is a survey of the published literature, not reproducible computation.
Conclusion
No claim is reproducible because no claim is generated by an author pipeline. This is a
controlled drop with drop_reason = non_pipeline. Honest outcome; nothing was
forced. «our HPC» compute was not required (and was transiently unreachable at drop time,
which is immaterial — there is no job to run).
No individual results have been recorded for this entry yet.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
PMID 36920617 is a narrative review ('Environ Sci Pollut Res' 2023) with no Methods section, Data-availability = 'NA', and no original datasets or pipeline outputs — there is nothing to reproduce. The auto-extracted code (Prodigal) and data (zenodo 3678129 = the unrelated BIOCOM-PIPE pipeline) links are text-mining false positives, an artifact of our intake, not an authors' defect. Accordingly q1/q2 are red (no comparable data/endpoint) but q5/q7 are only yellow — derivability and core-claim are undefined for a review, not fabrication-suspect. Overall this is an honest non_pipeline drop (q8 yellow: out-of-scope, no discrepancy, nothing forced).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.