Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Machine learning reveals microbial interactions driving plastic degradation across plastisphere environments.

Front Microbiol · 2026
L1 87/100 3/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

salvaged by watchdog from agreement.json (agent omitted ROOM_RESULT.json)

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-07-29

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether different aquatic environments (ocean, surface water, wastewater) harbor distinct potential plastic-degrading bacteria (PDBs) with unique degradation capabilities, and whether these potential PDBs are associated with specific non-plastic-degrading bacteria (NDBs) that facilitate plastic degradation through ecological interactions, using 16S rRNA sequencing combined with machine learning.

Core claims
  • Wastewater plastispheres harbor the most diverse and compositionally even microbial communities among the three habitats. finding
  • Potential PDBs are habitat-specific: Pseudomonas, Acinetobacter, and Aquabacterium in wastewater; Flavobacterium and Alteromonas in ocean; Psychrobacter and Novosphingobium in surface waters. finding
  • Potential PDBs show consistent co-occurrence with diverse NDB taxa (e.g., Clostridium_sensu_stricto_5, Lachnospiraceae_UCG-001, Cloacibacterium), suggesting facilitative interactions such as redox modulation, nutrient exchange, and biofilm support. mechanism
  • Combining Pearson's correlation networks with Random Forest modeling effectively identifies key taxa and potential PDB-NDB ecological interactions. method
  • A unified QIIME2-based pipeline applied to publicly available 16S rRNA datasets enables cross-habitat comparison of plastisphere composition, matched against PlasticDB to define potential PDBs at genus level. resource
  • ML application to plastisphere data is limited by taxonomic resolution, lack of functional validation, and insufficient integration of environmental metadata. finding
Experimental setups
Assay System Perturbation Readout Platform
16S rRNA gene amplicon sequencing (public NCBI SRA reanalysis) Ocean plastisphere biofilm communities (marine plastics) none (observational; plastic colonization in situ) OTU-level microbial community composition and relative abundance NCBI SRA datasets; QIIME2 v2023.2; DADA2; SILVA 138 classifier
16S rRNA gene amplicon sequencing (public NCBI SRA reanalysis) Wastewater plastisphere biofilm communities none (observational; plastic colonization) OTU-level microbial community composition and relative abundance NCBI SRA datasets; QIIME2 v2023.2; DADA2; SILVA 138 classifier
16S rRNA gene amplicon sequencing (public NCBI SRA reanalysis) Surface water plastisphere biofilm communities (72 samples after QC) none (observational; plastic colonization) OTU-level microbial community composition and relative abundance NCBI SRA datasets; QIIME2 v2023.2; DADA2; SILVA 138 classifier
Alpha diversity analysis (Shannon, Simpson, Chao1, Species Richness) Plastisphere communities across ocean, surface water, wastewater none Diversity indices compared across environments R; ANOVA
Beta diversity analysis (Bray-Curtis dissimilarity, PCoA) Plastisphere communities across ocean, surface water, wastewater none Community composition differences among environments R; vegan package (2.6-10); PERMANOVA with 999 permutations (adonis)
Potential PDB identification via database matching OTU table from all habitats matched to PlasticDB none Genus-/phylum-level identification of potential plastic-degrading taxa PlasticDB
Correlation network and machine learning analysis Potential PDB vs NDB relative abundance profiles across environments none Pairwise correlations and key NDB features influencing PDB abundance Pearson's correlation; Random Forest
Key results
  • Wastewater plastispheres were the most diverse and compositionally even communities relative to ocean and surface water.
  • Habitat-specific potential PDB genera identified (Pseudomonas, Acinetobacter, Aquabacterium in wastewater; Flavobacterium, Alteromonas in ocean; Psychrobacter, Novosphingobium in surface waters).
  • Consistent co-occurrence patterns between potential PDBs and NDB taxa (Clostridium_sensu_stricto_5, Lachnospiraceae_UCG-001, Cloacibacterium) detected by network and Random Forest analyses.
  • ANOVA revealed significant variation in microbial alpha diversity among the three environments.
Key statistics
  • count 824 microbial species reported with plastic-degrading capabilities; 230 proteins involved in plastic breakdown (Known plastic degraders from literature (Gambarini et al., 2022))
  • other up to 99.1% accuracy (ML classifiers predicting plastic-degrading microbes (Hemalatha et al., 2021))
  • other 97% similarity threshold (OTU clustering with qiime vsearch cluster-features-de-novo)
  • count 999 permutations (PERMANOVA (adonis) for beta diversity)
  • count 72 samples (Surface water dataset after removal of low-quality sequences)
  • other 460 million tons; 20-fold increase since 1964; ~91% unrecycled (Global plastic production and waste statistics (Introduction))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This secondary analysis pooled nine publicly available 16S rRNA sequencing datasets (three per environment) to compare plastisphere microbial communities across ocean, surface water, and wastewater. Alpha diversity metrics were compared among environments using ANOVA, and beta diversity was assessed with Bray–Curtis dissimilarity followed by PCoA ordination and PERMANOVA. Associations between potential plastic-degrading bacteria (PDB) and non-plastic-degrading bacteria (NDB) were explored with Pearson correlation and Random Forest modeling. Results were visualized with box plots and jitter plots; the full results section was not included in the supplied text.

Replicationunclear Sample sizeNine publicly available datasets were selected (three per environment); surface water retained 72 samples after QC; exact per-group n used in statistical tests not stated in the available text Groupsocean vs. surface water vs. wastewater plastisphere communities Pairingunpaired Randomization/blindingnot stated DispersionIQR Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
one-way ANOVA comparison of alpha diversity metrics (Shannon, Simpson, Chao1, Species Richness) across ocean, surface water, and wastewater plastisphere communities not stated
PERMANOVA (adonis, vegan, 999 permutations) beta diversity differences in Bray–Curtis dissimilarity among ocean, wastewater, and surface water plastisphere communities not stated
Pearson's correlation linear associations between potential PDB and NDB relative abundances not stated
Random Forest identification of key NDB features influencing potential PDB abundance across environments na
Principal Coordinates Analysis (PCoA) ordination of Bray–Curtis dissimilarity matrix for beta diversity visualization na
Approaches that could also have been used
  • Alpha diversity metrics were compared across three environments using ANOVA
    Could also: Kruskal-Wallis test followed by post-hoc Dunn's test with FDR correction could also compare alpha diversity across groups — Diversity indices from 16S rRNA data are often zero-inflated and non-normally distributed, particularly with unequal group sizes; a non-parametric approach makes no distributional assumption and is widely used for this comparison in microbiome studies
  • Multiple alpha diversity metrics (Shannon, Simpson, Chao1, Species Richness) were each tested separately without a stated multiplicity correction
    Could also: Applying a Benjamini-Hochberg FDR correction across the family of ANOVA tests would also control the false discovery rate — Testing four correlated diversity metrics simultaneously increases the family-wise type I error rate; a multiplicity correction clarifies which comparisons remain significant after accounting for the number of tests performed
  • PDB–NDB associations were quantified with Pearson's correlation on relative abundance data
    Could also: Spearman's rank correlation could also quantify monotonic associations between PDB and NDB taxa — Relative abundance data from amplicon sequencing are typically zero-inflated and non-normally distributed; Spearman's correlation requires no normality assumption and is less sensitive to extreme values, and is widely used for this type of co-occurrence analysis
  • Community composition was visualized and tested using Bray–Curtis dissimilarity with PCoA and PERMANOVA
    Could also: Non-metric multidimensional scaling (NMDS) combined with PERMANOVA could also serve as the ordination approach — NMDS does not assume a linear relationship between dissimilarity and ordination distance, which makes it broadly applicable to community data; it is a common complement or alternative to PCoA in microbiome studies
  • Random Forest was used to identify key NDB features associated with PDB abundance
    Could also: Sparse penalized regression (e.g., LASSO or elastic net) could also identify predictive NDB taxa — Penalized regression provides coefficient estimates with a natural effect-direction interpretation and can handle high-dimensional compositional data; it offers a complementary approach to Random Forest variable importance, which does not distinguish positive from negative associations
  • Datasets from nine separate studies were combined into a single analysis without a stated batch-correction step or mixed-effects modeling of study origin
    Could also: A linear mixed-effects model or PERMANOVA with study/batch as a covariate could also account for between-study technical heterogeneity — Pooling 16S rRNA datasets generated with different primers, sequencing platforms, and laboratory protocols introduces technical variation; explicitly modeling study as a random effect or blocking factor is a standard approach to reduce confounding from inter-study batch effects
Software: QIIME2 2023.2 · R / vegan 2.6-10 · DADA2 · SILVA classifier (qiime feature-classifier classify-sklearn) 138

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig 6Fig 4Fig 5Fig 6B
C1
Reported
75.15%
Reproduced
75.1468%
exact
C2
Reported
26.88%
Reproduced
26.8781%
exact
C3
Reported
24.02%
Reproduced
24.0205%
exact
C4
Reported
r>0.50
Reproduced
r>=0.50
exact
C5
Reported
9 named genera
Reproduced
same 9 genera = pipeline top set
exact
C6
Reported
more than five (avg r>0.65)
Reproduced
exactly 5 each; avg r 0.665 / 0.646
partial
C7
Reported
>6 distinct PDB each
Reproduced
exactly 4 distinct PDB each
partial
C8
Reported
shipped correlation.csv
Reproduced
72/76 exact (max abs diff 0.0)
exact
C9
Reported
threshold R2>0.5 (no value)
Reproduced
50/76 PDB genera R2>0.5 (lolopy)
within tolerance
C10
Reported
qualitative only
Reproduced
0.9595
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This is a strong reproduction: every concrete reported number (three PCA variances, the correlation pipeline, the 9 named top-interactor genera) reproduces from the authors' own shipped tables to printed/machine precision, so data identity and endpoint comparability are clean (q1/q2 green). The substantive issue is authors'-side: the Fig 4 text overstates interaction-partner counts — 'more than five' is exactly 5 (C6) and 'more than six' is exactly 4 (C7) per the authors' own regenerated PDB_NDB_corr.csv, with Stenoxybacter's avg r=0.646 also below the stated >0.65. These overstated counts are not derivable from the shared data (q5 yellow, fabrication-check flagged) and the deviation sits in the results/claim layer rather than input preprocessing (q3/q4 red). Severity is moderate — magnitude inflated but the qualitative direction (these genera are the top interactors) and the central thesis hold — so the core claim is confirmed in limited form and overall quality is solid-with-explainable-deviation (q6/q7/q8 yellow).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

154.9 k
tokens (I/O) · 18.2 M incl. cache
70 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.