Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Shifts from cooperative to individual-based predation defense determine microbial predator-prey dynamics.

ISME J · 2023
L1 84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1). In scope = the authors' own deterministic ODE predator-prey model (R + rodeo tabular->Fortran codegen + deSolve; 3 experiments F0_50/F3_50/F5_50 x 37 daily transfers; all params shipped in repo @88a96e8). The model builds, compiles, and integrates, and reproduces every published Fig.5/abstract signature: predator (flagellate) near-extinction (~100x crash) then recovery, filamentous-phenotype dominance by day37 (fraction 0.978), predator-free control with F=0, and the cooperative(toxin)->individual(filament) defense shift. Beyond the 80% floor the simulated predator dynamics were cross-checked against the repo's independently-shipped observed data: the F5_50 crash minimum (model 1143 vs observed 1250, same day 14) and recovery agree. This is a FRESH re-run on «our HPC» (SLURM «job», node n093) that reproduces the earlier «job» bit-for-bit (deterministic, no RNG); the conda env + rodeo 0.7.7 + repo were rebuilt/re-staged after the janitor reclaimed the «infra» workdir. NO fabrication concern: fully deterministic, every value derivable from shipped tables and consistent with observed data. KEY ENV FIX: the repo pins no rodeo version; current CRAN rodeo changed compile() default to fortran=FALSE, breaking the shipped compile('functions.f95') (-> 'functions.f95:1:8: unexpected symbol'). Pinning the contemporaneous rodeo 0.7.7 (current on CRAN at the 2023-02-19 commit) makes it run unmodified. NOT attempted (out of scope): 16S amplicon reanalysis of SRA PRJNA830374 (no shipped pipeline), all wet-lab quantities, parameter re-estimation, pixel-level figure overlay.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 84
    assessed: 2026-06-16 ⛓ 01470f27e58d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study investigates why and how bacterial predation-defense strategies succeed one another over many generations of co-culture with a bacterivorous flagellate, testing whether an initially dominant cooperative (toxin-based) defense is superseded by an individual-based defense due to maximization of individual rather than population-level benefits.

Core claims
  • Across all 16 replicate co-cultures of P. putida and P. lacustris, bacteria show a consistent succession of defense strategies over five weeks (~35 predator generations). finding
  • An initial cooperative defense based on toxic secondary metabolites is highly effective, bringing flagellate predators close to extinction. finding
  • The cooperative toxin-based defense is consistently superseded by a second, individual-based defense (filamentation) that arises via de novo mutations. finding
  • The second defense is inferior to the first in terms of reducing predator abundance. finding
  • The succession of defenses is not caused by predator adaptation invalidating the original (metabolite-based) defense. finding
  • The succession from cooperative to individual defense is driven by maximization of individual rather than population-level fitness benefits, and rapid evolution undermines social cooperation. mechanism
  • An ODE-based semi-continuous mathematical model (7 state variables, 9 processes) reproduces and explains the observed predator-prey and defense-succession dynamics. method
  • Data and model code are made publicly available via a git repository. resource
Experimental setups
Assay System Perturbation Readout Platform
Semi-continuous co-culture with daily dilution/resource replenishment Pseudomonas putida KT2440 + Poteriospumella lacustris JBM10 (16 replicates) varying initial flagellate/bacteria densities (high, low, absent-control) daily predator and prey cell densities epifluorescence microscopy, NSI-elements AR 5.11.01 software, Neubauer improved counting chambers
Morphometric analysis of bacterial filaments (length/size distribution) Pseudomonas putida (single-celled and filamentous forms), October 2020 replicates none/other (grazing pressure over time) filament length and size distribution of 100 individuals per sample/day NSI-elements AR 5.11.01, SYBR Green I staining, epifluorescence microscopy
Whole-genome sequencing 9 filamentous P. putida isolates + 1 non-filamentous control isolate none (comparative genomics of evolved isolates) nucleotide mismatches/mutations associated with coding sequences (synonymous vs non-synonymous) Illumina NovaSeq 6000 (2x150 bp paired-end), BWA-MEM, Ococo, Prokka, EMBOSS Transseq, Galaxy pipeline
Flagellate growth inhibition/exposure assay using culture filtrates axenically grown Poteriospumella lacustris flagellates exposure to sterile-filtered filtrates from co-cultures vs bacteria-only control cultures, sampled at different experimental phases (day 2-3, 8-12, 27-29) flagellate growth rate after 24 h exposure syringe filtration (0.2 µm), Neubauer counting chambers
Key results
  • Consistent succession of bacterial defense strategies observed in all 16 replicate co-cultures n=16
  • Initial cooperative toxin-based defense brought flagellate predator populations close to extinction
  • Toxin-based defense was superseded by filamentation defense arising from de novo mutations, which was less effective at suppressing predators
  • Coefficient of variation among replicate observations never exceeded 0.5 for any time point after day six CV ≤ 0.5
  • Mutations identified in filamentous isolates relative to non-filamentous control, located in coding sequences
  • Flagellate growth was inhibited by filtrates from co-cultures, verifying metabolite-mediated growth inhibition
  • Succession of defenses occurred independent of predator adaptation to the original toxin-based defense
Key statistics
  • count n = 16 (number of replicate co-culture experiments)
  • other coefficient of variation ≤ 0.5 (variability among replicate abundance observations after day 6)
  • count ~35 predator generations (generations of P. lacustris over the 5-week experiment)
  • count ~5 million paired reads, minimum 100-fold coverage (whole-genome sequencing depth for filamentous isolates)
  • count 1 × 10^5 flagellates mL^-1 and 1 × 10^5 bacteria mL^-1 (initial densities, high predation pressure line)
  • count 1 × 10^3 flagellates mL^-1 and 1 × 10^4 bacteria mL^-1 (initial densities, low predation pressure line)
  • count 1 × 10^4 bacteria mL^-1, 0 flagellates (initial densities, predator-absent control line)
  • other dilution rate 0.5 day^-1 (daily semi-continuous culture transfer rate)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study paired a 5-week semi-continuous co-culture experiment (n = 16 replicate vials across three starting-date blocks and two initial-density conditions) with a deterministic ODE-based mathematical model. Analysis was primarily descriptive: daily cell abundances were tracked by microscopy, replicate consistency was characterised with the coefficient of variation, and flagellate growth rates were computed arithmetically from 24-hour exposure assays. No formal inferential hypothesis tests are reported; mechanistic conclusions are supported principally by the calibrated ODE model and the qualitative consistency of patterns across all replicates.

Replicationbiological Sample size16 total co-culture replicates; duplicate runs per condition per start date; three independent start dates (Sep 2020, Oct 2020, Apr 2021) plus additional Apr/Jun 2021 high-density runs; triplicate batches for filtrate exposure assays; 9 filamentous isolates + 1 non-filamentous control for WGS GroupsHigh vs. low initial flagellate density vs. predator-absent control; co-culture vs. control filtrates; naive vs. previously exposed flagellates Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Coefficient of variation (descriptive summary statistic) Assessment of between-replicate consistency of abundance time-series (Fig. 1 caption) 16 co-culture replicates; 4-replicate subset for filament measurements not stated
Arithmetic growth-rate computation from density measurements Flagellate growth rates during 24-hour filtrate exposure assays (Methods: Verification of flagellate growth inhibition) Triplicate 3-mL batches per condition not stated
ODE numerical integration (deterministic mechanistic model) Simulation of predator-prey dynamics across all experimental phases (Methods: Mathematical modeling) not stated
Approaches that could also have been used
  • Dispersion across replicates is shown only graphically (individual dots) and summarised with a single CV threshold; no numerical spread metric (SD, IQR, CI) is reported for abundance data at any time point
    Could also: Report SD or IQR at each sampled time point, or add shaded uncertainty bands to time-series plots — Explicit dispersion metrics allow readers to numerically gauge between-replicate variability and facilitate comparison with other predator-prey studies reporting similar summaries
  • Flagellate growth rates from the 24-hour filtrate exposure assays (triplicate batches per condition) are computed arithmetically without a formal inferential comparison between filtrate origins or time points
    Could also: A one-way ANOVA or Kruskal-Wallis test with pairwise post-hoc comparisons (e.g., Tukey HSD or Dunn) could formally compare growth rates across filtrate types and harvest phases — Formal testing would yield effect-size estimates and p-values, making it easier to evaluate whether inhibition differences between co-culture and control filtrates, or across harvest time points, exceed what would be expected by chance with n = 3
  • Replicate consistency is assessed qualitatively via a single CV threshold (≤ 0.5 after day 6) rather than a formal reproducibility metric
    Could also: Intraclass correlation coefficient (ICC) computed across replicates for each daily time point would provide a bounded, confidence-interval-bearing measure of reproducibility — ICC separates within-replicate measurement error from between-replicate biological variability and is interpretable on a standardised 0–1 scale, supporting stronger reproducibility claims
  • ODE model parameters are reported as fixed point estimates with no stated uncertainty quantification or formal goodness-of-fit statistic
    Could also: Markov chain Monte Carlo (MCMC) sampling or a bootstrap resampling scheme could propagate parameter uncertainty into model trajectories and yield credible or confidence intervals around simulated dynamics — Reporting uncertainty bands around model predictions would allow readers to assess how robust the mechanistic conclusions are to the inherent uncertainty in parameter estimation from noisy experimental data
  • Mutation identification relied on whole-genome sequencing of 9 individually isolated filamentous clones and 1 non-filamentous control, providing a snapshot of genotypic diversity at a single endpoint
    Could also: Metagenomic (population-level) sequencing at multiple experimental time points could track allele frequencies longitudinally across the 5-week experiment — Longitudinal allele frequency data would provide direct empirical evidence of selective sweep dynamics and complement the ODE model's inference about mutation emergence timing
  • The two initial-density conditions (high and low flagellate density) and the predator-absent control are treated descriptively, with no formal statistical test comparing trajectory shapes or endpoints across conditions
    Could also: Linear mixed models or generalised additive mixed models (GAMMs) with condition as a fixed effect and replicate as a random effect could formally test whether abundance trajectories differed by initial condition — Such models accommodate the repeated-measures, non-linear nature of the time-series data and would quantify whether initial density systematically altered the timing or magnitude of the observed defense succession
Software: R/Fortran (rodeo package) · NSI-elements AR 5.11.01 · SICKLE (windowed adaptive trimming) · BWA-MEM · Ococo · Prokka · EMBOSS Transseq · Galaxy platform

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
15
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GCA_000007565.2 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PRJNA830374 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig.5
C1
Reported
ODE model builds (rodeo->Fortran) and integrates all 3 experiments x 37 transfers (deterministic, all params shipped)
Reproduced
Builds and integrates all 3 experiments x 37 transfers with contemporaneous rodeo 0.7.7 (deSolve 1.42, R 4.5.3, rtol=atol=1e-10). Fresh «job» reproduces original «job» bit-for-bit. Prior build failure was version drift only (current CRAN rodeo changed compile() default to fortran=FALSE).
exact
C2
Reported
predators driven 'close to extinction' early (F3_50, F5_50)
Reproduced
Flagellate F crashes ~100x from peak (F5_50 peak 1.156e5 -> min 1.143e3 @day14.12; F3_50 min 1.015e3 @day15.12). Observed F5_50 min 1250 @day14 -> model within ~9%, same crash day.
within tolerance
C3
Reported
predators 'always recovered' by day ~37
Reproduced
F recovers to ~4.1e4 by day37 in both predator experiments (F3_50 40953, F5_50 40943); observed recovers to ~7.4-7.6e4 (same order, clear recovery from the ~1e3 minimum).
within tolerance
C4
Reported
filamentous phenotype 'became dominant' by end
Reproduced
day37 filament fraction Bff/(Bff+singlecell)=0.978 in both F3_50 & F5_50 (Bff=2.061e7 vs singlecell 4.683e5); observed filaments (Bf_len) rise to ~1.49e7.
within tolerance
C5
Reported
control F0_50 has no predators throughout
Reproduced
F identically 0 at every timepoint; bacteria stay single-cell (Bo=4.206e7, Bff=0, X=0).
exact
C6
Reported
shift from cooperative (toxin) to individual-based (filament) defense (central thesis)
Reproduced
Fig.5 reproduction shows early toxin pulse (X) + toxin-producer transient (Bx) giving way to filament dominance (Bff) after day~20 (qualitative/structural; no numeric table in paper).
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

Clean, faithful reproduction. The authors' deterministic ODE predator-prey model (R + rodeo->Fortran + deSolve, all params shipped @88a96e8) builds, integrates all 3 experiments x 37 transfers, and reproduces every Fig.5/abstract signature: ~100x flagellate crash (model min 1143 vs observed 1250, same day14), recovery to ~4.1e4 by day37, filament dominance (fraction 0.978), predator-free control (F≡0), and the cooperative->individual defense succession. The only deviation was a resolvable dependency-version drift (rodeo 0.9.2 vs era-correct 0.7.7) on our side — not a result discrepancy and not an authors' defect. The one honest limitation is q2: the paper's claims are qualitative and shipped as figures (no numeric table), so exact-value comparison was indirect/structural, backed by the repo's own observed data. No fabrication signal.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

624.4 k
tokens (I/O) · 54.4 M incl. cache
170 min
runtime · 0 CPU-h
0.2 GB
peak RAM
2
HPC jobs
hummel
machine