geneshot: gene-level metagenomics identifies genome islands associated with immunotherapy response.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL. geneshot is a published Nextflow tool (authors' own code, P16 own-repo) applied to 3 public ENA WGS cohorts. (1) DATASET PROFILING done for all accessions: PRJEB22893 (25 WGS, complete), PRJNA397906 (44 WGS, complete), PRJNA399742 (paper 'n=172' is misleading - 172 runs but only 39 WGS + 133 16S amplicon, 49 biosamples). (2) The 108-specimen analysis design (C7) was VERIFIED exactly from the intact deposited Nextflow manifest AND independently corroborated by the Nextflow trace (cutadapt x108, bwa x108). (3) HEADLINE NUMERIC CLAIMS (catalog 7.2M genes, 1.23M CAGs, 3019 significant CAGs, read-assignment %, per-sample gene ranges) are UNVERIFIABLE: all three large result HDF5 files in the figshare deposit are corrupt - results.hdf5 truncated to 17.4%, corncob.hdf5 to 25.1% (both md5-match figshare, so truncated at deposit time, not download), and details.hdf5 returns HTTP 400. This is a data-integrity finding, NOT proof of fabrication; a human should request a corrected deposit. (4) Full de-novo re-run (assemble 108 deep WGS metagenomes -> 7.2M-gene catalog) is computationally out of scope. NOT attempted: full reassembly; AMGMA genome-island calling; specific-taxa interpretation. Pipeline-runnability test on «our HPC»: in progress.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 75assessed: 2026-06-18 ⛓ d354a19c859b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan gene-level metagenomic analysis of WGS microbiome data, using co-abundant gene groups (CAGs) as the unit of analysis, identify common, experimentally testable microbial genomic features consistently associated with response to immune checkpoint inhibitor (ICI) cancer therapy across independent cohorts?
- ★ geneshot is a gene-level metagenomic bioinformatics tool that clusters de novo assembled protein-coding genes into co-abundant gene groups (CAGs) to reduce dimensionality and generate testable hypotheses from WGS microbiome data method
- ★ Applying geneshot to two independent ICI cohorts identifies microbial genomic islands consistently associated with ICI response in culturable type strains finding
- ★ ICI response-associated CAGs map to contiguous strain-specific genomic islands (10-35 kb) rather than being spread throughout the genome finding
- ★ Odoribacter splanchnicus genomic islands encoding type II secretion and TonB-dependent transport are at higher abundance in ICI responders across both cohorts mechanism
- ★ Genomic islands in Clostridium sp. HMb25 and Faecalibacterium prausnitzii annotated as integrated prophages (plus a CRISPR defense island in F. prausnitzii) implicate bacteriophage growth in the ICI-microbiome interaction mechanism
- ★ geneshot identifies taxa associated with ICI response: Odoribacter splanchnicus and Gemmiger formicilis enriched in responders; Coprococcus comes enriched in progressors finding
- geneshot is implemented as an end-to-end Nextflow workflow using corncob beta-binomial modeling and betta for taxon/function-level association, available open-source resource
- The gene-level approach reduces reliance on incomplete reference databases by inferring genomic islands without requiring a reference genome for the actual source organism method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| whole-genome shotgun (WGS) metagenomic sequencing reanalysis (de novo assembly + gene-level CAG analysis) | human stool microbiome from metastatic melanoma patients receiving ICI therapy | none (observational; ICI responder vs progressor) | relative abundance of protein-coding genes/CAGs and association with ICI response | geneshot pipeline (MegaHit, Prodigal, linclust, DIAMOND, FAMLI, corncob, betta; Nextflow) |
| taxonomic annotation | deduplicated gene catalog from gut microbiome WGS | none | lowest common ancestor taxonomic assignment of genes | DIAMOND blastp vs NCBI RefSeq DIAMOND index (2020-01-15) |
| functional annotation | deduplicated gene catalog | none | functional annotation of genes | eggNOG-mapper v5.0 (database 2020-07-17) |
| reference genome alignment of CAG member genes (genome island mapping) | 1886 reference genomes from associated taxonomic families; Odoribacter splanchnicus DSM 220712, Clostridium sp. HMb25, Faecalibacterium prausnitzii | none | genomic coordinates/contiguity of CAG member genes and proportion of genome aligned | AMGMA tool using DIAMOND |
- – 3019 CAGs (4509 genes) significantly associated with ICI response at FDR 0.01 3019 CAGs / 4509 genes
- – 2634 CAGs associated with progression and 385 CAGs associated with response 2634 vs 385 CAGs
- – ICI-associated CAG genes localize to contiguous genomic islands ranging 10-35 kb 10-35 kb
- ▲ Odoribacter splanchnicus type II secretion and TonB-dependent transport islands enriched in ICI responders across both cohorts
- – Odoribacter splanchnicus and Gemmiger formicilis more abundant in responders; Coprococcus comes more abundant in progressors
- – Co-abundance clustering yielded 1,232,769 distinct CAGs from the gene catalog 1,232,769 CAGs
- – Deduplicated gene catalog of 7,209,758 unique protein-coding sequences generated from per-sample assemblies 7,209,758 genes
- count 7,209,758 unique protein-coding sequences in gene catalog (deduplicated gene catalog across all specimens)
- count 75,488–536,005 (median 280,065) genes per sample by de novo assembly (genes assembled per stool sample)
- count 380,202–2,071,835 (median 1,174,474) genes detected per sample (genes detected by alignment per sample)
- count 1,232,769 distinct CAGs (co-abundance clustering output)
- count 3019 CAGs (4509 genes) at FDR 0.01 (CAGs significantly associated with ICI response)
- other median 86.1% (proportion of raw WGS reads uniquely assigned to gene catalog)
- count 1886 reference genomes (genomes used for genomic island alignment)
- count n=25 (PRJEB22893), n=44 (PRJNA397906), n=172 (PRJNA399742) (sample sizes of the three analyzed BioProjects)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
geneshot is a Nextflow-based pipeline that groups co-abundant microbial protein-coding genes into co-abundant gene groups (CAGs) via average-linkage cosine-distance clustering, then tests each CAG for association with a clinical outcome. CAG-level relative abundances were modeled with a beta-binomial regression (corncob) that included cohort as a covariate, applied jointly across two ICI cohorts (n=44 and n=172). Taxon- and function-level evidence was synthesized by aggregating corncob coefficient estimates through the errors-in-response model betta (intercept-only). CAGs meeting an FDR threshold of 0.01 were declared significant, and their member genes were subsequently aligned to reference genomes to localize genomic islands.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Beta-binomial regression (corncob), modeling logit of expected relative CAG abundance | CAG-level association with ICI response (responders vs. progressors) across two cohorts, with cohort as covariate | n=216 (PRJNA397906 n=44 + PRJNA399742 n=172); PRJEB22893 n=25 excluded due to missing phenotype | not stated |
| Errors-in-response model (betta), intercept-only; tests whether the intercept (mean taxon/function-level coefficient) equals zero | Taxon-level and function-level aggregation of corncob CAG coefficients | null | not stated |
-
CAG-level differential abundance was modeled with a beta-binomial regression (corncob) that included cohort as a fixed covariate↳ Could also: DESeq2 (negative binomial GLM with size-factor normalization) or ANCOM-BC (bias-corrected log-ratio framework) are also widely used for count-based metagenomics differential abundance testing — DESeq2 provides variance-shrinkage and normalized log2-fold-change estimates enabling direct comparison of effect magnitudes across CAGs; ANCOM-BC explicitly accounts for the compositional constraint of relative-abundance data, which could be relevant given that CAG abundances are computed as fractions of total read depth
-
Taxon-level and function-level evidence was synthesized by aggregating corncob coefficient estimates through the errors-in-response model betta (intercept-only)↳ Could also: A random-effects meta-analysis (e.g., via the metafor R package) applied to CAG-level coefficient estimates per taxon could also pool evidence across CAGs — A random-effects framework would additionally quantify between-CAG heterogeneity (I² or τ²) within each taxon, indicating whether the association signal is broadly distributed across gene groups or concentrated in a subset, which could inform mechanistic interpretation
-
Multi-cohort integration was achieved by including cohort as a fixed covariate in the beta-binomial model fitted jointly across all specimens↳ Could also: A two-stage meta-analysis approach — fitting cohort-specific models first, then combining standardized effect estimates — or a mixed-effects model with cohort as a random effect could also integrate multiple cohorts — Explicit meta-analysis frameworks provide formal between-cohort heterogeneity statistics (I², Q-test) and cohort-specific effect estimates alongside the pooled result, which may be informative when cohorts differ in cancer subtype, treatment regimen, or sequencing protocol
-
An FDR threshold of 0.01 was applied across all CAGs, but the specific FDR algorithm is not specified↳ Could also: Storey's q-value procedure, or the Benjamini-Yekutieli FDR correction (which is valid under arbitrary dependence), are alternatives to the Benjamini-Hochberg procedure used in high-dimensional genomics — CAG abundances are correlated by construction (co-abundant genes are clustered together), so specifying an FDR method suited to dependent tests and reporting the estimated proportion of true nulls (π₀) would further characterize the signal landscape
-
Gene relative abundance was defined compositionally as sequencing depth per gene divided by total depth across all genes in each specimen↳ Could also: Centered log-ratio (CLR) transformation prior to modeling is an alternative compositional normalization used in microbiome analyses — CLR transformation maps compositional data into an unconstrained Euclidean space, which is a natural fit for linear models and co-abundance correlation; comparing CAG groupings derived from CLR-transformed versus raw relative-abundance cosine distances could assess the sensitivity of CAG membership to the normalization choice
-
Significance was summarized as counts of CAGs passing the FDR threshold, without reported effect size estimates or confidence intervals in the main text↳ Could also: Reporting the corncob-estimated coefficients (log-odds of relative abundance difference) with 95% confidence intervals for the top-ranked CAGs, or volcano-plot-style visualization of effect size versus significance, are also standard in metagenomics association studies — Effect size estimates alongside p-values allow readers to distinguish biologically large associations from statistically significant but small-effect ones, and confidence intervals convey uncertainty in the estimated differences — both are useful for prioritizing candidates for experimental follow-up
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33952321 (geneshot)
Paper: Minot SS, Barry KC, Kasman C, Golob JL, Willis AD. geneshot: gene-level metagenomics identifies genome islands associated with immunotherapy response. Genome Biol 2021;22:135. DOI 10.1186/s13059-021-02355-6.
Tool: geneshot — an end-to-end Nextflow pipeline for gene-level metagenomics of WGS microbiome reads. Repo: https://github.com/Golob-Minot/geneshot (paper used v0.6.2). Authors' own code → P16 "own repo" case.
Deposited outputs (figshare 14417120, ~4.7 GB): the actual geneshot run outputs —
ici.manifest.csv (sample sheet), ici.params.json (exact params),
2020-06-04-ICI.results.hdf5 (562 MB, gene catalog + CAGs + abundances),
2020-06-04-ICI.details.hdf5 (3.7 GB), ...corncob.hdf5 (789 MB, beta-binomial
association), plus Nextflow report.html/trace.txt.
Pipeline steps (Methods)
preprocess (adapter trim + human removal) → per-specimen assembly + gene prediction →
dedup to gene catalog (AA clustering) → align reads to catalog → relative abundance
(length-adjusted) → CAGs (average-linkage co-abundance clustering) → optional
taxonomic/functional annotation → corncob beta-binomial regression
(formula treatment_response + BioProject) → aggregate to HDF5. AMGMA tool maps
significant genes onto reference genomes to call "genome islands."
In scope (pipeline-derived results)
| id | claim | paper location | pipeline |
|---|---|---|---|
| C1 | gene catalog = 7,209,758 unique protein-coding seqs | Results | geneshot dedup |
| C2 | per-specimen genes assembled 75,488–536,005 (median 280,065) | Results | assembly/prediction |
| C3 | median 86.1% reads uniquely assigned to catalog | Results | alignment |
| C4 | genes detected per sample 380,202–2,071,835 (median 1,174,474) | Results | abundance |
| C5 | 1,232,769 CAGs | Results | co-abundance clustering |
| C6 | 3,019 CAGs (4,509 genes) FDR<0.01; 385 response, 2,634 progression | Results | corncob |
| C7 | n=108 WGS specimens analysed; 80 with ICI phenotype (40 resp/40 prog) | Methods/manifest | manifest |
Out of scope (not pipeline / not feasibly re-derivable here)
- Full de-novo re-run of catalog from raw reads on 108 deep WGS stool metagenomes (assembly of ~108 samples, many >100 GB-RAM metaSPAdes jobs, hundreds of node-hours): computationally infeasible within this study's budget → NOT attempted at full scale. Instead: (a) verify deposited outputs vs paper numbers; (b) demonstrate the pipeline is runnable (geneshot test/CI workflow on «our HPC»).
- Genome-island identification (AMGMA) onto reference genomes (10–35 kb operons, type-II secretion / TonB / phage) — downstream annotation + manual biological interpretation → out of scope.
- Specific taxa calls (Odoribacter splanchnicus, Gemmiger formicilis, Coprococcus comes) — downstream taxonomic annotation interpretation → out of scope.
Reproduction strategy (two prongs)
- Claim verification vs authors' deposit — download the figshare HDF5 outputs and read catalog size, CAG count, corncob significant-CAG counts, per-sample gene ranges, read-assignment %; compare to the paper. This is the strongest fabrication- detection check available (internal consistency: deposit ⟷ paper) short of a multi-week full reassembly.
- Pipeline runnability — run geneshot's bundled test workflow end-to-end on «our HPC» (env resolves, docs sufficient, produces a catalog+CAGs+HDF5 on toy data).
Datasets
- PRJEB22893 — 25 WGS (Frankel 2017); no ICI phenotype in public metadata → in catalog, omitted from model.
- PRJNA397906 — 44 WGS (Matson 2018).
- PRJNA399742 — paper says n=172; ENA = 172 runs but only 39 WGS + 133 16S-amplicon, 49 unique biosamples. geneshot used the 39 WGS (Gopalakrishnan 2018).
- Catalog/run used 108 WGS specimens total (per deposited manifest).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a data-integrity / unverifiable case, not a discrepancy or fabrication. The analysis design (C7: 108 WGS specimens, 40 response/40 progression, PRJEB22893 unlabeled) reproduced exactly from the intact manifest and Nextflow trace, but all six headline numeric claims (7.2M-gene catalog, 1.23M CAGs, 3,019 significant CAGs, 86.1% read assignment, per-specimen ranges) are unverifiable because every results HDF5 in the figshare deposit is truncated at upload time (md5 matches figshare). The defect sits on the authors'/deposit side (broken deposit), so q4 and q2 are red, but there is no evidence the values are wrong — raw data and the open pipeline exist — so q5/q7/q8 stay yellow rather than fabrication-red. Recommended action: request a corrected/complete figshare deposit before re-grading.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.