Environment and Co-occurring Native Mussel Species, but Not Host Genetics, Impact the Microbiome of a Freshwater Invasive Species (Corbicula fluminea).
The main results reproduced, with only marginal, non-material deviations.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL but the paper's CENTRAL claim is reproduced. 16S microbiome pipeline run end-to-end on «our HPC»: DADA2 1.30.0 (maxEE 2/5, truncQ=2, 243-263bp) -> de-novo DECIPHER+FastTree -> mixed-depth rarefaction (biv 4000/env 2000) -> W/U-UniFrac -> vegan adonis2 + pairwiseAdonis (paper's repo, commit cb190f7). RESULT: environment structures the C. fluminea microbiome FAR more than native mussel species - CF-vs-sediment R2=0.49 (paper 0.47, near-exact) and CF-vs-seston R2(U)=0.195 (paper 0.17) DWARF CF-vs-mussels R2=0.04 (paper U 0.03, matches); all P=0.001; ordering sediment>seston>>mussels preserved; pairwiseAdonis CF-vs-each-mussel R2 0.07-0.26 (paper 0.11-0.33). C2 post-rarefaction ASVs 30652 vs 31091 (98.6%). Exact UniFrac R2 (weighted/unweighted split) differs from paper, explained by the documented tree deviation (de-novo FastTree vs SEPP/GreenGenes-13.8; UniFrac is phylogeny-sensitive) + larger sample subset (378 vs 318) + DADA2 version. C1 un-rarefied 43085 vs 57556 (version/pooling). NOTE: brief data pointer PRJNA757758 is RAD-Seq (host genetics, out of scope); 16S is in PRJNA757734/740316/761344 (382 runs = superset of 318). Out of scope: host RADseq, COI barcoding, FEAST, alpha-diversity stats.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 55assessed: 2026-06-21 ⛓ 153165228ef1
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study investigates how intrinsic (host genetic variation) and extrinsic (environmental conditions, co-occurring native freshwater mussels) factors shape the gut microbiome of the invasive Asian clam Corbicula fluminea, and whether it reciprocally influences or reflects the microbiome of co-occurring native unionid mussels.
- ★ The gut microbiome of C. fluminea is diverse, differs with environmental conditions, and varies spatially among rivers, but is unrelated to host genetic variation finding
- ★ Microbial source tracking suggests the gut microbiome of C. fluminea may be influenced by the presence of co-occurring native mussels finding
- ★ PICRUST2-inferred functions show high prevalence and diversity of degradation functions in the C. fluminea microbiome, especially degradation of carbohydrates and aromatic compounds finding
- ★ The modularity and functional diversity of the C. fluminea microbiome may be an asset allowing acclimation to an extensive range of nutritional sources in invaded habitats, potentially aiding invasive success mechanism
- ★ Population genomic (RADseq) data were integrated to test whether C. fluminea clonal lineage or within-lineage genetic ancestry contributes to microbiome diversity method
- FEAST source tracking was used bidirectionally to test reciprocal influence between C. fluminea and native mussel microbiomes, using seston and sediment as additional sources method
- Unionid microbiomes show some degree of species-specificity, with co-occurring mussel species collected from the same environment having distinct microbiomes (prior literature) finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| 16S rRNA gene sequencing (V4 region, Illumina MiSeq) | Corbicula fluminea gut tissue | none (natural gradient of environment/co-occurring mussel density) | gut bacterial community composition and diversity | Illumina MiSeq |
| 16S rRNA gene sequencing (V4 region, Illumina MiSeq) | six native unionid mussel species (Lampsilis ovata, Cyclonaias pustulosa, Cyclonaias asperata, Fusconaia cerina, Tritogonia verrucosa, Amblema plicata) gut tissue | none | gut bacterial community composition and diversity | Illumina MiSeq |
| 16S rRNA gene sequencing | sediment samples from collection sites | none | environmental microbial community composition (source for FEAST) | Illumina MiSeq |
| 16S rRNA gene sequencing | filtered water/seston samples from collection sites | none | environmental microbial community composition (source for FEAST) | Illumina MiSeq |
| RADseq population genomics | Corbicula fluminea mantle tissue | none | genetic structure/ancestry across populations | — |
| PICRUST2 functional inference from 16S ASV data | C. fluminea and native mussel gut microbiome data | none | predicted metabolic/enzymatic pathway abundance (MetaCyc) | PICRUST2 |
| Physicochemical water/sediment measurement | site water and sediment | none | temperature, pH, conductivity, dissolved oxygen, DOC, SRP, NH4+, NO2-, NO3-, sediment granulometry | YSI DO Probe |
- – C. fluminea gut microbiome diversity and structure differed with environmental conditions and varied spatially among rivers
- – C. fluminea microbiome was unrelated to host genetic variation (ancestry/clonal lineage)
- – FEAST source tracking indicated presence of co-occurring native mussels may influence the C. fluminea gut microbiome
- ▲ PICRUST2 predicted high prevalence and diversity of degradation functions, notably carbohydrate and aromatic compound degradation, in C. fluminea microbiome
- – Sequence coverage (Chao's non-parametric indicator) was high across rarefied samples 0.98 ± 0.02
- ▼ ASV richness dropped substantially after rarefaction of the full dataset 57,556 to 31,091 ASVs
- – C. fluminea density on sites was generally correlated with native mussel density
- count 180 C. fluminea specimens and 144 Unionidae specimens (6 species) collected (total specimen collection across 16 sites)
- count 80 C. fluminea and 144 native mussels used for microbiome analysis (subset selected for DNA extraction/sequencing)
- mean 0.98 ± 0.02 (Chao's non-parametric coverage estimator across rarefied samples)
- count 57,556 ASVs (total ASVs in un-rarefied dataset across all samples)
- count 31,091 ASVs (ASVs remaining after rarefaction of bivalve and environmental samples)
- count C. fluminea density averaged 0.5-92 individuals/m2 (density range across study sites)
- count native mussel density ranged 0.6-23 individuals/m2 (density range across study sites)
- pvalue P < 0.05 (threshold for DESeq2-identified pathways enriched in C. fluminea vs. native mussels)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study characterized the gut bacterial microbiome (16S rRNA V4, DADA2-processed ASVs) of 80 invasive Corbicula fluminea and 144 native freshwater mussels (6 species) across 16 field sites in two river basins, with library sizes normalized by rarefaction before all diversity analyses. Alpha-diversity was compared across host taxa using Wilcoxon signed-rank tests, and beta-diversity was visualized via PCoA of weighted/unweighted UniFrac distances with significance assessed by PERMANOVA (pairwiseAdonis post-hoc). Microbial source tracking (FEAST), Mantel tests, Kruskal-Wallis tests, and Pearson correlations addressed environmental drivers and cross-species microbiome exchange, while PICRUSt2-inferred metabolic pathway enrichment was tested with a DESeq2 negative binomial model.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon signed-rank test | Comparison of alpha-diversity (Shannon, Faith's PD) between C. fluminea and native mussel species — pooled across all species and separately for each species | 80 C. fluminea and 144 native mussels (6 species) across 16 sites | not stated |
| PERMANOVA (adonis, vegan) | Overall gut microbiome structure differences among host species, sites, and rivers using U- and W-UniFrac dissimilarities; also for functional Bray-Curtis dissimilarities between C. fluminea and native mussels | 80 C. fluminea + 144 native mussels for bivalve analyses | not stated |
| Pairwise PERMANOVA (pairwiseAdonis) | Post-hoc pairwise comparisons of microbiome structure following PERMANOVA | null | not stated |
| Spearman correlation (via envfit, vegan) | Correlation between microbiome dissimilarities and physicochemical site characteristics (sites with complete measurements only) | null | not stated |
| Spearman correlation | Geographic variation within C. fluminea: U- and W-UniFrac dissimilarities versus geographic distances between sites along the same river | 80 C. fluminea across 16 sites in 6 rivers | not stated |
| LEfSe (Linear discriminant analysis Effect Size) | Identification of microbial taxa differing between C. fluminea and each native mussel species, computed separately per species pair across all sites | null | not stated |
| Mantel test (Spearman-based, 500 iterations on random equal-sized subsamples; median p-value reported) | Co-variation of native mussel and co-occurring C. fluminea gut microbiomes | null | not stated |
| FEAST (Fast Expectation-mAximization microbial Source Tracking) | Estimation of proportional contribution of co-occurring microbial communities (other mussels, C. fluminea, seston, sediment) to each recipient bivalve gut microbiome, run in both directions | 80 C. fluminea + 144 native mussels as sinks or sources; environmental samples as additional sources | na |
| Kruskal-Wallis test | Comparison of source-tracking contribution proportions (reciprocal influence) across recipient or source species and sites (separate tests) | null | not stated |
| Pearson's correlation | Correlation of source-tracking contribution proportions with physicochemical variables, native mussel density, C. fluminea density, and native mussel richness on scaled data (psych R-package) | null | not stated |
| DESeq2 Wald test (negative binomial GLM) | Identification of PICRUSt2-inferred MetaCyc metabolic pathways enriched in C. fluminea compared to native mussels (P < 0.05) | 80 C. fluminea + 144 native mussels | not stated |
-
Rarefaction to a fixed read depth (4,000 sequences per bivalve; 2,000 per environmental sample) was used to normalize library sizes before all alpha- and beta-diversity analyses↳ Could also: Variance-stabilizing transformation (DESeq2 VST), cumulative-sum scaling (CSS in metagenomeSeq), or methods that model library size directly (e.g., ANCOM-BC) could also be applied — Rarefaction discards a portion of sequencing data and can reduce statistical power, particularly for samples near the rarefaction threshold; normalization methods that retain all reads may improve sensitivity and are increasingly used alongside or instead of rarefaction
-
Wilcoxon signed-rank tests were used to compare alpha-diversity indices between C. fluminea and each native mussel species, with individual specimens treated as the unit of analysis↳ Could also: A linear mixed-effects model (e.g., lme4 in R) with site or river as a random effect, or a Wilcoxon rank-sum (Mann-Whitney U) test for two independent groups, could also be applied — Multiple individuals collected from the same site share environmental conditions and are not fully independent; a mixed model explicitly partitions within-site from between-site variance, which can affect standard errors and p-values; rank-sum tests are the standard non-parametric choice for comparing two independent (unpaired) groups
-
Separate LEfSe analyses were computed for each pairwise comparison between C. fluminea and each native mussel species↳ Could also: ANCOM-BC or MaAsLin2 could also be applied in a multi-group design with explicit FDR correction across all taxa and comparisons simultaneously — Running multiple independent LEfSe comparisons increases the number of tests performed without a unified correction scheme; methods that natively control FDR across all taxa and groups provide a more coherent framework for false discovery control in multi-group microbiome comparisons
-
Mantel tests were run on 500 random equal-sized subsamples of matched mussel–C. fluminea pairs per site, and the median of the 500 p-values was reported↳ Could also: A single Mantel test using all available pairwise distances with a large permutation count (e.g., 999 or 9,999 permutations) could also be used — Summarizing across 500 independently resampled Mantel tests is non-standard and the statistical properties of combining or summarizing such p-values differ from a single permutation-based test; a single Mantel test with permutation inference is more directly interpretable under standard null-hypothesis testing conventions
-
Pearson's correlation was used to relate source-tracking proportions to physicochemical and biotic variables on scaled data↳ Could also: Spearman's rank correlation could also be used, particularly given that source-tracking output proportions are bounded between 0 and 1 and may be skewed — Pearson's correlation assumes a linear relationship between normally distributed variables; source-tracking proportions are compositional and often right-skewed, and Spearman correlation is more robust to departures from normality and to outliers in bounded proportional data
-
DESeq2's negative binomial GLM was applied to PICRUSt2-inferred MetaCyc pathway abundances to identify pathways enriched in C. fluminea↳ Could also: Testing on CLR (centered log-ratio)-transformed inferred abundances using a linear model or limma-voom could also be applied to predicted functional data — DESeq2 is designed for raw sequencing count data and assumes a negative binomial distribution; PICRUSt2 pathway abundances are predicted from marker-gene data rather than directly observed counts, and may not conform to that distributional assumption, whereas approaches for CLR-transformed compositional data may align more closely with the nature of inferred functional profiles
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35444631
Paper: Environment and Co-occurring Native Mussel Species, but Not Host Genetics, Impact the Microbiome of a Freshwater Invasive Species (Corbicula fluminea). Front. Microbiol. 2022. DOI 10.3389/fmicb.2022.800061. PMID 35444631 / PMC9014210.
Code pointer: https://github.com/pmartinezarbizu/pairwiseAdonis (third-party
PERMANOVA tool, commit cb190f7). Per P16 this is a fully valid reproducible
artifact: we run the tool on the paper's own data with the described design.
Data pointers: brief lists PRJNA757758 — that is RAD-Seq host genomics,
NOT the microbiome. The 16S amplicon data are in PRJNA757734 (80 CF),
PRJNA740316 (217: native mussels + seston + sediment), PRJNA761344 (85
bivalve). 382 amplicon runs deposited = superset of the 318 the paper analysed.
In scope (pipeline-derived, attempted)
The paper's central computational result is a 16S amplicon microbiome pipeline:
| Result | Pipeline | Claim |
|---|---|---|
| Un-rarefied ASV count | DADA2 (filterAndTrim maxEE c(2,5), truncQ=2) + 243-263bp length filter | C1: 57,556 ASVs |
| Post-rarefaction ASV count | mixed-depth rarefaction (bivalves 4000, environmental 2000) | C2: 31,091 ASVs |
| Sample count | sample inventory | C3: 318 (80 CF + 144 native mussels + env) |
| PERMANOVA CF vs native mussels | W/U-UniFrac + vegan adonis2 | C4: P=0.001, R2=0.09/0.03 |
| PERMANOVA CF vs seston | W/U-UniFrac + adonis2 | C5: P=0.001, R2=0.17/0.17 |
| PERMANOVA CF vs sediment | W/U-UniFrac + adonis2 | C6: P=0.001, R2=0.47/0.50 |
| Pairwise CF vs each mussel sp. | W-UniFrac + pairwiseAdonis (the repo) | C7: R2=0.11-0.33 |
Phylogeny for UniFrac in the paper: SEPP insertion into GreenGenes 13.8 (via QIIME2). Deviation: our primary tree is de-novo (DECIPHER align + FastTree) because (a) the gg-13-8 SEPP path is extremely heavy on ~57k fragments and (b) the QIIME2 SEPP env failed to build cleanly; group-level UniFrac PERMANOVA R2 is robust to tree-construction method. SEPP attempted as an optional faithfulness check.
Version deviation (documented): paper used DADA2 in R 3.6.3. The bioconda
dada2=1.14.0 / R 3.6.3 build is broken (S4 SRFilterResult/Mnumeric class
defect — fails in filterAndTrim). We therefore denoise with DADA2 1.30.0 / R
4.3.3 using the paper's exact parameters. ASV counts are version-sensitive, so
C1/C2 are expected to be close but not identical.
Out of scope (not attempted)
- Host genetics / RAD-Seq (PRJNA757758) — population-genomic / STACKS-type analysis; the paper's point is that host genetics does NOT structure the microbiome, but that is a separate non-amplicon pipeline.
- COI / mussel species barcoding (GenBank OK047371–OK047417) — Sanger/manual.
- FEAST source-tracking, Shannon alpha-diversity stats, taxonomy bar plots — secondary; not the pinned quantitative claims (may be added if time permits).
- Wet-lab steps (DNA extraction, library prep).
Compute
All heavy compute on «our HPC» («infra») via SLURM (std, 32 cpus, ≤12h wall). Data +
envs + intermediates on «infra»:
«path». «host» holds small results only.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.