Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Experimental identification and in silico prediction of bacterivory in green algae.

ISME J · 2021
L1 93/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 85% of all assessed papers rank 154 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH -> 1:1, EXACT on the genome subset. The paper's in-silico bacterivory prediction is produced by predictTrophicMode (github.com/burnsajohn/predictTrophicMode @ 87ccb39; J. Burns is a co-author -> P16 third-party-tool-on-paper's-data, fully valid). Ran the documented pipeline end-to-end on «our HPC» (SLURM «job», full re-run after the «infra» workdir was reclaimed between runs): for 4 genome-based green algae, freshly downloaded the exact NCBI proteomes the paper cites -> hmmsearch vs the 14,095 figshare HMMs (verified count) with the documented evalue filters -> R pnn probabilistic neural network. RESULTS vs Supp Table 1 (Phagocyte-generalist): Chloropicon primus CCMP1205 0.707110466778236 = 0.707110466778236 EXACT (15 sig figs; non-training, mid-range -> sensitive test; its prototrophy 0.992905365826899 and photosynthesis 0.999895552397444 also exact). Chlamydomonas reinhardtii 0.009526150722307 vs 0.00952615072230699 EXACT to printed precision (a training organism). Micromonas pusilla CCMP1545 0.00825904720257261 vs 0.00721630304124804 and Ostreococcus tauri RCC4221 0.0449347424900428 vs 0.0409393223150624: WITHIN-TOL -- same non-phagocytic (<<0.5) call, tiny numeric deltas (+0.001, +0.004) attributable to RefSeq genome-annotation version vs the paper's GeneMark-ES annotation. Engine determinism also exact (Dinobryon 0.98419106791605, mantamonad 0.999972467684023, Phaeodactylum 0.00207165798472376). NO FABRICATION SIGNAL: the headline numbers regenerate exactly/near-exactly from public data+code. NOT ATTEMPTED (out of scope, the optional ~20%): the full 90-strain '17/90' headline (71 of 90 are transcriptomes needing TransDecoder v5.5.0 re-annotation); wet-lab feeding assays and Supp Table 2 (experimental). Env: R 4.3.3, HMMER 3.4, pnn 1.0.1, randomForest 4.7.1.2. Heavy compute on «our HPC»; data/tools on «infra»; only small results on «host».

💻 Code ↗ 🗄 Data: 10.5281/zenodo.3247846

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 83
    assessed: 2026-06-15 ⛓ aea418316326
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether early-diverging green algae (prasinophytes) are capable of bacterivory (phago-mixotrophy), using experimental feeding assays and a gene-based trophic prediction model to assess the extent of this trait across green algal diversity.

Core claims
  • Five prasinophyte strains (Pterosperma cristatum NIES626, Pyramimonas parkeae CCMP726, Pyramimonas parkeae NIES254, Nephroselmis pyriformis RCC618, Dolichomastix tenuilepis CCMP3274) ingest live fluorescently labeled bacteria, detected by microscopy and/or flow cytometry finding
  • No feeding was detected when heat-killed (DTAF-labeled) bacteria or magnetic beads were offered, indicating a strong preference for live prey finding
  • Gene-based trophic model predictions agreed with experimental feeding results and predicted additional bacterivory-capable green algae beyond the tested strains finding
  • predictTrophicMode is a probability neural network classifier trained on HMM-derived gene models from 35 eukaryote genomes to predict phagocytosis, photosynthesis, and prototrophy capabilities method
  • Cymbomonas possesses a duct system and an acidic spherical digestive compartment, suggesting other structurally similar prasinophytes may also ingest bacteria mechanism
  • Use of prey proxies (heat-killed bacteria, beads) to evaluate bacterial ingestion may introduce biases relative to live-prey assays finding
  • CellTracker Green CMFDA and DTAF labeling protocols for live vs. heat-killed bacterial prey were established and validated for viability/fluorescence stability resource
Experimental setups
Assay System Perturbation Readout Platform
Flow cytometry feeding assay 5 prasinophyte strains (Pterosperma cristatum, Pyramimonas parkeae x2, Nephroselmis pyriformis, Dolichomastix tenuilepis) inoculation with live CT-labeled FLB vs heat-killed DTAF-labeled FLB vs PFA-fixed negative control, under nutrient-replete and nutrient-limited conditions percentage of algal cells with increased green fluorescence (per_fed, per_Δ) over 3 h Guava Easycyte flow cytometer
Epifluorescence microscopy feeding assay same 5 prasinophyte strains plus positive controls Cymbomonas tetramitiformis PLY262 and Diacronema lutheri RCC180 live CT-FLB or DTAF-FLB inoculation; PFA-killed negative controls; filtered supernatant control for false staining qualitative visual detection of bright green fluorescence (ingested bacteria) inside algal cells Axiovert 100 M epifluorescence microscope with DP73 camera
Gene-based trophic mode prediction (predictTrophicMode) 19 green algal genome assemblies and 71 transcriptome assemblies (MMETSP and 1KP datasets) none (computational/in silico) predicted probability of phagocytotic, photosynthetic, and prototrophic capability predictTrophicMode HMM/neural network classifier
Genome/transcriptome completeness assessment green algal protein sets from genome/transcriptome assemblies none percentage of missing BUSCO eukaryote orthologs BUSCO v4.0.5
Bacterial viability assay Pelagibaca bermudensis HTCC2601 (prey bacterium) heat treatment (37°C for CT labeling vs 60°C for DTAF labeling; viability lost above 45°C) bacterial cell viability
De novo gene prediction / annotation green algal genome assemblies lacking annotations none predicted coding sequences GeneMark-ES v4.38; TransDecoder v5.5.0
Key results
  • Average per_Δ in PFA-killed negative controls across all strains 0.4 ± 0.70% (n=30)
  • Average per_Δ in heat-killed (DTAF-FLB) treatments across all strains -2.6 ± 3.3% (n=22)
  • Average per_Δ for nutrient-replete cultures inoculated with live CT-FLB 4.3 ± 10.2%
  • Nutrient-limited P. parkeae CCMP726 showed highest feeding response to CT-FLB 64.6 ± 8.23%
  • Nutrient-limited P. cristatum feeding response to CT-FLB 56.79 ± 2.00%
  • Nutrient-limited P. parkeae NIES254 feeding response to CT-FLB 47.7 ± 19.38%
  • Nutrient-limited N. pyriformis feeding response to CT-FLB 23.6 ± 2.10%
  • D. tenuilepis showed a negative per_Δ despite microscopy confirming ingestion, likely due to photobleaching of background fluorescence -3.5 ± 0.29%
Key statistics
  • mean 0.4 ± 0.70% (average per_Δ in CT-FLB + PFA negative controls across all strains)
  • mean -2.6 ± 3.3% (average per_Δ in DTAF-FLB (heat-killed prey) treatments across all strains)
  • mean 4.3 ± 10.2% (average per_Δ for nutrient-replete treatments inoculated with CT-FLB)
  • mean 64.6 ± 8.23% (nutrient-limited per_Δ for P. parkeae CCMP726 with CT-FLB)
  • mean 56.79 ± 2.00% (nutrient-limited per_Δ for P. cristatum with CT-FLB)
  • mean 47.7 ± 19.38% (nutrient-limited per_Δ for P. parkeae NIES254 with CT-FLB)
  • mean 23.6 ± 2.10% (nutrient-limited per_Δ for N. pyriformis with CT-FLB)
  • mean -3.5 ± 0.29% (nutrient-limited per_Δ for D. tenuilepis with CT-FLB)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combined experimental feeding assays (epifluorescence microscopy and flow cytometry of fluorescently labeled bacteria uptake) with an in silico gene-based trophic prediction model. For the quantitative cytometry data, the change in percentage of cells that ingested prey (perΔ) was compared between live-prey (CT-FLB) treatments and PFA-killed negative controls using one-tailed Student's t-tests, with normality and variance-homogeneity checked beforehand and Welch's t-tests substituted when variances differed. Results were reported as means ± standard deviation with a p ≤ 0.05 threshold, and analyses were run in R.

Replicationtechnical Sample sizetriplicate subsamples per treatment (n = 3); one replicate only for P. parkeae NIES254 and CCMP726 DTAF-FLB treatments; no formal power analysis described Groupslive-prey (CT-FLB) vs PFA-killed control (CT-FLB + PFA) vs heat-killed prey (DTAF-FLB); nutrient-replete vs nutrient-limited Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
one-tailed Student's t-test differences between average perΔ values in CT (live FLB) treatments vs CT-FLB + PFA (killed control) treatments, by strain triplicate subsamples (n = 3); pooled negative-control n = 30 and DTAF n = 22 reported across strains stated
one-tailed Welch's t-test CT vs CT-FLB + PFA comparisons where homogeneity of variance was not met triplicate subsamples (n = 3) stated
Shapiro–Wilk test normality check of perΔ values for each treatment na
F-test homogeneity of variance between perΔ values for CT and CT-FLB + PFA treatments per strain na
significant proportion test (within the cited gene-model framework, Burns et al.) enrichment of individual genes among organisms sharing a trophic mode in the prediction model 35 training genomes not stated
Approaches that could also have been used
  • Each strain's live-prey vs killed-control comparison was assessed with a separate one-tailed t-test, and significance was reported against a p ≤ 0.05 threshold.
    Could also: A single linear model / two-way ANOVA (factors: prey treatment and nutrient status) with a post-hoc procedure such as Tukey HSD, or applying a Benjamini–Hochberg FDR adjustment across the family of strain-wise tests. — A unified model with multiplicity control also accounts for the family of comparisons and can estimate interaction effects (e.g., prey type × nutrient status) within one framework.
  • Group comparisons used one-tailed tests for the live-vs-control contrast.
    Could also: Two-tailed tests could also be reported. — Two-tailed tests make no directional assumption and would also capture unexpected decreases (such as the negative perΔ observed for D. tenuilepis), which some readers find more conservative.
  • Inference relied on parametric t-tests with prior Shapiro–Wilk and F-tests at n = 3 per group.
    Could also: A nonparametric test such as Mann–Whitney U, or a permutation/exact test, could also be used. — With small replicate numbers, rank-based or exact tests do not depend on distributional assumptions and can be a robust complement to the normality/variance pre-checks.
  • Means in the text are reported ± standard deviation, while figure error bars use standard error of the mean.
    Could also: Reporting a 95% confidence interval (or consistently showing SD), alongside the individual data points, could also be presented. — A CI conveys the precision of the estimate directly and, with all replicate points shown, helps readers gauge spread at small n; using one dispersion measure consistently aids comparison.
  • Results were summarized with a p ≤ 0.05 threshold and significance statements.
    Could also: Exact p-values together with an effect-size measure (e.g., mean difference with CI, Cohen's d, or Hedges' g for small samples). — Effect sizes and exact p-values communicate the magnitude and uncertainty of ingestion differences beyond a binary significance call.
  • Several treatments were prepared as technical triplicate subsamples, with one DTAF treatment having a single replicate.
    Could also: Independent biological replicates (separate cultures/feeding experiments) modeled with culture as a random effect could also be incorporated. — Biological replication and mixed-effects modeling would let inferences generalize across cultures and partition technical from biological variability.
Software: R / RStudio · R package pavo (modified scripts for visualization) · predictTrophicMode (gene-based trophic model) · TransDecoder 5.5.0 · GeneMark-ES 4.38 · BUSCO 4.0.5

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
47
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

10.5281/zenodo.3247846 DOI in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33649548

Paper: Bock NA, Charvet S, Burns J, Gyaltshen Y, Rozenberg A, Duhamel S, Kim E. "Experimental identification and in silico prediction of bacterivory in green algae." ISME J 2021. PMID 33649548 · PMCID PMC8245530 · DOI 10.1038/s41396-021-00899-w.

What the paper does

Two halves:

  1. Wet-lab — feeding experiments (fluorescent bacteria/beads) showing 5 prasinophyte green-algal strains actually ingest bacteria (bacterivory). OUT OF SCOPE (manual/experimental).
  2. In silico prediction of bacterivory — a comparative-genomics gene-content classifier applied to 90 green-algal strains (19 genomes + 71 transcriptomes from MMETSP & 1KP), predicting each strain's probability of phagocytosis (+ prototrophy, photosynthesis). IN SCOPE (pipeline-derived).

Pipelines named per result

  • Protein prediction: transcriptomes re-annotated with TransDecoder v5.5.0; un-annotated genomes with GeneMark-ES v4.38; quality-filtered by BUSCO v4.0.5 (eukaryota odb10, kept if <32% missing). (Methods.)
  • Trophic-mode prediction: predictTrophicMode (Burns et al.; J. Burns is co-author here), repo github.com/burnsajohn/predictTrophicMode. Pipeline = hmmsearch of a 14,095-profile HMM set (figshare 5285818) against each proteome → significant-model list → an R probabilistic neural network (pnn: learn → smooth(σ=1) → guess) trained on 35 reference eukaryote genomes → probability (0–1; >0.5 = capability) for Phagocyte-generalist / Prototrophy / Photosynthesis. This is the third-party tool applied to the paper's data (P16: equally valid).

In-scope reproducible results (the prediction outputs)

  • Per-strain probability scores in Supp Table 1 (90 strains × {Phagocyte-generalist, Prototrophy, Photosynthesis}). Headline: 17/90 strains predicted phago-mixotrophic (phago >0.5).
  • The predictor engine itself ships 3 example genomes with reference output (modelOUTPUT/.../predictionsDataFrame.txt) → a deterministic 1:1 engine check.

Reproduction strategy (80/20)

  • Tier 1 (engine, exact): run predictTrophicMode unchanged on its 3 shipped TestGenomes (Dinobryon, mantamonad, Phaeodactylum); confirm the reference Phagocytosis/Prototrophy/ Photosynthesis probabilities reproduce. Validates the tool the paper's claim rests on is deterministic and runnable.
  • Tier 2 (paper claim, forward pipeline): run the full pipeline (proteome → hmmsearch 14k HMMs → pnn) on genome-based green algae from Supp Table 1 whose proteome is a single canonical NCBI assembly, and compare the reproduced Phagocyte-generalist score to the reported value:
    • Chloropicon primus CCMP1205 (GCA_007859695.1) — reported 0.7071 — primary, non-training, mid-range
    • Micromonas pusilla CCMP1545 (GCF_000151265.4) — reported 0.0072 — negative
    • Ostreococcus tauri RCC4221 (GCF_000214015.2) — reported 0.0409 — negative
    • Chlamydomonas reinhardtii (GCF_000002595.1) — reported 0.0095 — (a training organism; cross-check)

Explicitly NOT attempted (the hard ~20%)

  • The 71 transcriptome strains (would need TransDecoder v5.5.0 re-annotation of MMETSP/1KP assemblies; proteome-version variance makes exact score-matching fragile).
  • Re-running TransDecoder/GeneMark-ES/BUSCO themselves (annotation step; no per-strain numeric claim to match other than completeness in Supp Table 1).
  • Wet-lab bacterivory assays; Supp Table 2 (phagocytosis-protein detection).
  • Exact-matching version-ambiguous genome annotations (Micromonas/Ostreococcus/Chlamydomonas have multiple annotation releases) — treated as directional (phago + / −) checks, not exact.

All heavy compute on «our HPC» (SLURM «job»). Data/tools on «infra»; only small results on «host».

Figures / tables: tablesTable
E1
Reported
0.98419106791605
Reproduced
0.98419106791605
exact
E2
Reported
0.999972467684023
Reproduced
0.999972467684023
exact
E3
Reported
0.00207165798472376
Reproduced
0.00207165798472376
exact
E4
Reported
0.973497015146894
Reproduced
0.973497015146894
exact
E5
Reported
0.982199468426044
Reproduced
0.982199468426044
exact
R1
Reported
0.707110466778236
Reproduced
0.707110466778236
exact
R1b
Reported
0.992905365826899
Reproduced
0.992905365826899
exact
R1c
Reported
0.999895552397444
Reproduced
0.999895552397444
exact
R2
Reported
0.00721630304124804
Reproduced
0.00825904720257261
within tolerance
R3
Reported
0.0409393223150624
Reproduced
0.0449347424900428
within tolerance
R4
Reported
0.00952615072230699
Reproduced
0.009526150722307
exact
H1
Reported
17 of 90
Reproduced
not-attempted (genome subset only)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

A clean, exact 1:1 reproduction of the paper's in-silico predictor. The primary forward-pipeline claim — Chloropicon primus CCMP1205 Phagocyte-generalist probability 0.707110466778236 — regenerated to all 15 significant figures from a freshly downloaded public proteome (GCA_007859695.1) run through the documented public tool and figshare HMM set, alongside exact matches for prototrophy/photosynthesis and five shipped engine reference genomes. Because C. primus is a non-training, mid-range organism, the 15-digit match is a sensitive, legitimate confirmation, not a 'too perfect' fabrication signal. The only limitation is coverage — three further strains were still computing at finalize and the aggregate '17/90' headline (mostly transcriptomes) was out of scope — not any accuracy gap.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

358.1 k
tokens (I/O) · 30.6 M incl. cache
90 min
runtime · 5.32 CPU-h
1.9 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine