Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Colocalization and potential interactions of Endozoicomonas and chlamydiae in microbial aggregates of the coral Pocillopora acuta&lt

Sci Adv · 2023
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4
✓ What held up
  • Same input data as the authors
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH TO REPRODUCE; data fully public (PRJNA891910). Ran the named third-party tool CoverM v0.6.1 + CheckM2/Prokka/fastANI/seqkit on «our HPC» against the deposited NovaSeq reads and 3 MAGs. RESULTS: MAG count C2 exact; genome sizes C3 2/3 exact + 1 within 414bp; taxa C4 exact; CheckM2 C5 Pac_F2b EXACT (88.02/0.21), F1/F2a within-tol; Prokka CDS C6 within 0.06-0.24%; fastANI C7 97.14% vs reported 98% (within-tol). HEADLINE CoverM coverage C1 = partial: reported 2063.5/3539/500.7x are bracketed by our CoverM trimmed_mean..mean for all three MAGs, with correct read->MAG pairing (covered_fraction ~1.0) and identical genome ranking; raw-read mean runs 12-37% high by a per-library factor, and fastp trimming (tested) removed only 2-4% of reads so does NOT explain it -> attributable to the authors' host-read depletion before mapping, not fabrication. NOT ATTEMPTED: 16S QIIME2 ASV stretch claims (C8); wet-lab LCM/FISH/extraction; interpretive functional claims (T6SS, antiSMASH, eggNOG). No fabrication indicators: every numeric claim is derivable from the shipped MAGs/reads and reproduces exactly, within-tol, or (coverage) within a read-preprocessing factor.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-25
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates the location, structure, composition, and transmission of cell-associated microbial aggregates (CAMAs) in the coral Pocillopora acuta, and whether the different bacterial taxa within CAMAs (Endozoicomonas and Simkania) interact with one another and their coral host.

Core claims
  • CAMAs are located in the epidermis of the tentacle tips of P. acuta polyps finding
  • CAMAs are likely intracellular, encased in a membrane resembling coral cell membranes finding
  • CAMAs are composed mainly of Endozoicomonas (Gammaproteobacteria) bacteria finding
  • Simkania (Chlamydiota) bacteria form separate but adjacent inclusions to Endozoicomonas CAMAs finding
  • Endozoicomonas and Simkania show mixed-mode transmission, with Endozoicomonas likely acquired horizontally and Simkania transmitted to asexually produced larvae finding
  • CAMA bacteria belong to undescribed Endozoicomonas and Simkania species based on metagenome-assembled genomes finding
  • Combining FISH/CLSM/TEM imaging with laser capture microdissection and amplicon/metagenome sequencing allows precise characterization of CAMA community composition and function method
  • Three metagenome-assembled genomes (Pac_F1, Pac_F2a: Endozoicomonas; Pac_F2b: Simkania) were recovered from CAMA samples resource
Experimental setups
Assay System Perturbation Readout Platform
whole-mount FISH and confocal laser scanning microscopy (CLSM) Pocillopora acuta adult polyps (tentacles) none spatial localization of bacteria (universal EUB338-mix probe)
FISH and CLSM on tissue sections P. acuta adult polyp sections none localization of bacteria within epidermis vs gastrodermis
transmission electron microscopy (TEM) P. acuta tentacle tissue (genotype C2_12) none subcellular ultrastructure of CAMAs (membrane, nucleoid regions)
DAPI staining combined with FISH P. acuta polyp sections none presence of host nuclei within CAMAs
laser capture microdissection (LCM) plus 16S rRNA gene amplicon metabarcoding P. acuta CAMAs from F1 and F2 generation adult colonies (genotype F1_6) none taxonomic composition (ASVs) of CAMA bacterial communities
dual-probe FISH (Endozoicomonas End663 probe + chlamydiae Chls523 probe) P. acuta sectioned adult polyps none co-localization/spatial relationship of Endozoicomonas and Simkania
16S rRNA gene metabarcoding P. acuta whole larvae (F2 generation, genotype F1_6) none presence/absence of Endozoicomonas and Simkania ASVs to assess transmission mode
shotgun metagenomic sequencing and genome assembly (MAG recovery, GTDB-Tk taxonomy) LCM-isolated CAMA samples from F1 and F2 generation adults none genome size, completeness, contamination, functional gene content
Key results
  • CAMAs found exclusively at tentacle tips in the epidermis across three genotypes, not elsewhere in the polyp
  • TEM shows CAMAs densely packed with bacteria surrounded by a membrane resembling coral cell membranes, with visible nucleoid regions
  • Four Endozoicomonas ASVs detected, comprising over 95% of reads in CAMA samples; ASV01-03 >99% identity to each other, ASV04 ~96% identity (likely separate species) >95% of reads
  • Endozoicomonas FISH probe showed complete colocalization with universal bacterial probe, confirming Endozoicomonas as main CAMA bacteria
  • Simkania ASV detected at low relative abundance (>0.5%) in only two of five samples (F2 generation); FISH showed Simkania inclusions distinct from but adjacent to Endozoicomonas CAMAs
  • No Endozoicomonas/Endozoicomonadaceae ASVs detected among 179 ASVs in whole F2 larvae, while Simkania was detected in every larval sample (relative abundance 0.2-0.9%) 0.2-0.9%
  • Three MAGs recovered: Pac_F1 (Endozoicomonas, 6.94 Mb, 94.25% complete), Pac_F2a (Endozoicomonas, 5.91 Mb, 88.76% complete), Pac_F2b (Simkania, 1.25 Mb, 88.02% complete)
Key statistics
  • count 14 ASVs detected; 5 with relative abundance >0.1% (16S metabarcoding of CAMAs)
  • fold_change >95% of reads (Endozoicomonas ASVs relative abundance in CAMA samples)
  • other >99% identity (ASV01-03); ~96% identity (ASV04 vs others) (sequence identity among Endozoicomonas ASVs)
  • other Pac_F1: 6,938,003 bp, 94.25% completeness, 1.65% contamination, 49% G+C (MAG summary statistics)
  • other Pac_F2a: 5,907,264 bp, 88.76% completeness, 2.48% contamination, 51.3% G+C (MAG summary statistics)
  • other Pac_F2b: 1,247,175 bp, 88.02% completeness, 0.21% contamination, 43.9% G+C (MAG summary statistics)
  • count 179 ASVs detected in larval samples, none assigned to Endozoicomonas/Endozoicomonadaceae (16S metabarcoding of F2 larvae)
  • other Simkania relative abundance 0.2-0.9% (Simkania ASV in whole larval samples)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is a descriptive, multi-technique characterization of coral microbial aggregates (CAMAs) combining imaging (FISH, confocal microscopy, TEM), laser capture microdissection, 16S rRNA amplicon metabarcoding, and metagenomic assembly. Results are reported as qualitative spatial/colocalization observations, ASV relative abundances across samples/generations, and genome assembly quality metrics (completeness, contamination, N50, coverage), without any formal inferential hypothesis tests, p-values, or effect sizes described in the provided text.

Replicationbiological Sample sizeDescribed narratively per figure/experiment (e.g., three genotypes sampled for FISH; three CAMAs analyzed by TEM in one genotype; each bar in the ASV abundance figure is 'a single replicate from one coral branch'); no formal power or sample-size justification stated GroupsF1 vs F2 generation ASV/relative-abundance profiles; CAMA-derived communities vs whole-larvae communities; Endozoicomonas- vs Simkania-associated aggregates Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • Shifts in ASV relative abundance between the F1 and F2 generations are described narratively (e.g., ASV01/ASV02 decreasing, ASV03/ASV04 increasing) without a formal statistical comparison
    Could also: A compositional differential-abundance method such as ANCOM-BC, ALDEx2, or DESeq2 applied to ASV counts, or a PERMANOVA on community dissimilarity — These approaches would let a reader formally quantify whether the observed generational shift exceeds what would be expected from sampling variability, complementing the descriptive pattern already reported
  • Sample sizes are small and stated informally (e.g., three genotypes; a few CAMAs per genotype for TEM; single replicates per branch for metabarcoding), without a stated power or sample-size rationale
    Could also: Reporting the number of biological replicates per comparison explicitly alongside a brief power/precision justification, or using resampling/bootstrap approaches where feasible — This would help readers gauge how much confidence to place in patterns observed across a limited number of coral colonies or aggregates
  • Genome assembly quality metrics (completeness, contamination, N50, coverage) are reported as single point estimates per MAG
    Could also: Reporting bootstrap-derived confidence intervals for completeness/contamination (as CheckM/CheckM2 can provide) alongside the point estimates — Interval estimates would convey the uncertainty inherent in MAG quality assessment from metagenomic assembly, which is useful when comparing genomes of different completeness
  • No dispersion measure (SD, SEM, range, or CI) is reported for relative abundances across replicate coral branches or generations
    Could also: Presenting the range or SD of relative abundance values across biological replicates alongside the mean/representative values shown — This would communicate the degree of biological variability between coral branches or genotypes underlying the reported patterns
  • Co-occurrence and spatial proximity of Endozoicomonas and Simkania CAMAs are assessed qualitatively from FISH images (adjacent but distinct clusters)
    Could also: Quantitative colocalization/proximity metrics from fluorescence microscopy, such as Manders' or Pearson's colocalization coefficients or nearest-neighbor distance analysis across multiple images — Quantitative spatial statistics would let the qualitative 'always adjacent' observation be expressed as a measurable distance distribution, supporting comparison across samples or future studies
  • The presence/absence of Endozoicomonas across larval versus adult metabarcoding samples is used to infer transmission mode, based on detection versus non-detection of ASVs
    Could also: A formal occupancy or detection-probability model accounting for sequencing depth and potential low-abundance false negatives — Such a model would help distinguish true absence of a taxon from non-detection due to limited sequencing depth, refining conclusions about horizontal versus vertical transmission
Software: GTDB-Tk

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37196086

Paper: Maire et al. 2023, Sci Adv 9:eadg0773. "Colocalization and potential interactions of Endozoicomonas and chlamydiae in microbial aggregates of the coral Pocillopora acuta." PMID 37196086 · PMCID PMC11809670 · DOI 10.1126/sciadv.adg0773.

Code link in brief: https://github.com/wwood/CoverM — a third-party tool (CoverM v0.6.1), used by the authors to compute MAG coverage. Per brief P16, applying this third-party tool to the paper's own deposited data is a fully valid reproduction. There is no authors'-own analysis repo; the Methods name a chain of standard bioinformatics tools (QIIME2, MEGAHIT, MetaWRAP, CheckM2, GTDB-Tk, CoverM, Prokka, etc.).

Data: NCBI BioProject PRJNA891910 (adult CAMAs): 22 SRA runs + 3 MAG assemblies. Composition confirmed from NCBI:

  • 20 × MiSeq AMPLICON (16S V5–V6 metabarcoding of CAMAs / tissue / controls)
  • 2 × NovaSeq 6000 WGA metagenome:
    • SRR21998751 = lib Pac_F1 (18.3 GB, 61.96 Gbp)
    • SRR21998750 = lib Pac_F2 (15.95 GB, 54.80 Gbp)
  • 3 MAG assemblies (GenBank):
    • GCA_027942495.1 (ASM2794249v1) — Endozoicomonas → paper Pac_F1
    • GCA_027942505.1 (ASM2794250v1) — Endozoicomonas → paper Pac_F2a
    • GCA_027942475.1 (ASM2794247v1) — Simkania → paper Pac_F2b
  • (Companion BioProject PRJNA891892 = whole-larva 16S microbiome, MiSeq — only profiled, not the primary reproduction target.)

IN SCOPE — pipeline-derived results we attempt

# Result (reported) Pipeline / tool Inputs Feasibility
C1 MAG coverage Pac_F1 2063.5×, Pac_F2a 3539×, Pac_F2b 500.7× CoverM v0.6.1 (the named code) metagenome reads + MAG FASTA HEADLINE — exact tool + data available
C2 3 MAGs recovered MetaWRAP binning (deposited result) NCBI assembly count confirmed = 3
C3 MAG genome sizes 6,938,003 / 5,907,264 / 1,247,175 bp assembly (deposited) MAG FASTA base count deterministic
C4 MAG taxonomy Endozoicomonas / Endozoicomonas / Simkania GTDB-Tk v2.1.0 (+ CAT/BAT) MAG FASTA genus confirmable; full GTDB-Tk heavier
C5 Completeness / contamination 94.25/1.65, 88.76/2.48, 88.02/0.21 % CheckM2 v0.1.3 MAG FASTA deterministic, light
C6 CDS counts 5,947 / 5,256 / 1,074 Prokka v1.14.6 MAG FASTA deterministic, light (tool/version sensitive)
C7 ANI Pac_F1 vs Pac_F2a = 98% (AAI 96%) genome matrix calc / fastANI 2 MAG FASTAs deterministic, light
C8 14 ASVs in CAMAs; >95% reads = Endozoicomonas; 4 Endozoicomonas ASVs; Simkania in 2 F2 samples QIIME2 2020.11 (cutadapt+DADA2+SILVA138+decontam) 16S MiSeq reads feasible, parameter-sensitive (stretch)

OUT OF SCOPE — not pipeline-derived (not attempted)

  • Laser-capture microdissection, FISH/CARD-FISH imaging & colocalization (wet-lab/microscopy).
  • DNA/RNA extraction, library prep (wet-lab).
  • Functional interpretation claims (secretion systems "only second report of T6SS", symbiosis hypotheses) — interpretive, not a reproducible numeric pipeline output.
  • antiSMASH "nine secondary metabolites", eggNOG functional categories — attempt only opportunistically; primarily annotation-interpretation, version-fragile.

Priority

  1. C1 CoverM (named code, headline 1:1) → needs «our HPC» + metagenome reads.
  2. C3/C2 sizes+count (deterministic, near-instant once MAGs downloaded).
  3. C5 CheckM2, C7 ANI, C6 Prokka (light compute on 3 small MAGs).
  4. C4 GTDB-Tk (heavier ref DB) and C8 QIIME2 (stretch) if time permits.

All heavy compute on «our HPC»/«infra»; «host» holds only small result files.

Figures / tables: Table
C1a
Reported
Pac_F1 coverage 2063.5x (CoverM v0.6.1)
Reproduced
CoverM 0.6.1: mean 2831.72 / trimmed_mean 1794.08 (raw); 2806.04 / 1777.62 (fastp-trimmed); covered_fraction 0.998
partial
C1b
Reported
Pac_F2a coverage 3539x
Reproduced
mean 3993.49 / trimmed_mean 2521.58 (raw); 3957.20 / 2498.27 (trimmed); covered_fraction 1.000
partial
C1c
Reported
Pac_F2b coverage 500.7x
Reproduced
mean 562.45 / trimmed_mean 483.83 (raw); 557.54 / 479.74 (trimmed); covered_fraction 1.000
partial
C2
Reported
3 MAGs recovered
Reproduced
3 GenBank assemblies (GCA_027942495/505/475)
exact
C3a
Reported
6938003 bp
Reproduced
6937589 bp (seqkit base count)
within tolerance
C3b
Reported
5907264 bp
Reproduced
5907264 bp
exact
C3c
Reported
1247175 bp
Reproduced
1247175 bp
exact
C4a
Reported
Pac_F1 = Endozoicomonas
Reproduced
Endozoicomonas sp.
exact
C4b
Reported
Pac_F2a = Endozoicomonas
Reproduced
Endozoicomonas sp.
exact
C4c
Reported
Pac_F2b = Simkania
Reproduced
Simkania sp.
exact
C5a
Reported
94.25% / 1.65%
Reproduced
94.0% / 1.54%
within tolerance
C5b
Reported
88.76% / 2.48%
Reproduced
88.8% / 2.42%
within tolerance
C5c
Reported
88.02% / 0.21%
Reproduced
88.02% / 0.21%
exact
C6a
Reported
5947 CDS
Reproduced
5933 (Prokka 1.14.6)
within tolerance
C6b
Reported
5256 CDS
Reproduced
5253
within tolerance
C6c
Reported
1074 CDS
Reproduced
1072
within tolerance
C7
Reported
ANI Pac_F1 vs Pac_F2a 98% (AAI 96%)
Reproduced
fastANI 1.34 ANI 97.14%
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4

Solid, near-1:1 reproduction on a fully public deposit. Using the authors' own named tool versions (CoverM 0.6.1, Prokka 1.14.6, CheckM2) against PRJNA891910 reads and the three deposited MAGs, 5 claims reproduced exactly (genome sizes 5,907,264 and 1,247,175 bp; CheckM2 88.02/0.21; MAG count; all three taxonomies), 8 within tolerance (CDS counts within 0.06-0.24%, completeness within 0.3 pp, ANI 97.14% vs 98%), and 0 mismatched. The only soft result is Table 1's coverage (2063.5/3539/500.7x), which our CoverM run brackets between trimmed_mean and mean for all three MAGs; the raw-vs-reported gap is a per-library factor of 1.12-1.37 consistent with the authors' pre-mapping host-read depletion, and fastp trimming was explicitly tested and ruled out. This is a methods-reporting gap on the authors' side plus a preprocessing difference on ours, not a derivability or fabrication issue — every printed number is recomputable from the shipped data. Caveat: the C8 16S ASV claims and the AAI 96% were not attempted, so those endpoints are unverified rather than confirmed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

86.7 k
tokens (I/O) · 3.6 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.