Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A novel and dual digestive symbiosis scales up the nutrition and immune system of the holobiont Rimicaris exoculata.

Microbiome · 2022
L1 64/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
64/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 25% of all assessed papers rank 854 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL (in progress). Described well enough to reproduce: YES — the authors ship a complete anvi'o workflow (gitlab.ifremer.fr/rimicaris/..., not the figtree pointer in the registry, which is only the Fig.3 tree viewer) plus an OPEN companion deposit (Ifremer dataref REX_Velo) containing the raw reads, the 6 Megahit assemblies, and the 21 MAGs. Two headline claims already reproduced from metadata/deposit listings WITHOUT compute: C1 (677,134,974 reads ~ reported 677M / 113M avg, exact) and C2 (21 MAGs deposited == reported 21, exact). Remaining claims (C3 dRep->20, C4 GTDB-Tk taxonomy, C5/C6 per-class GC+genome-size, C7 contig counts) require «our HPC» compute on «infra»; scripts staged. BLOCKER: central «our HPC» VPN tunnel is currently down («host» ssh times out) so no «infra» download/compute has run yet; retrying periodically per HARD-RULE-1d (not touching the VPN). NOT attempted: wet-lab/microscopy (FISH/TEM, chromosome spacing) = out of scope. Grades provisional, human-checkable.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 64
    assessed: 2026-06-19 ⛓ 19676afb7d57
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What are the functional roles, genetic potential, and host-symbiont interactions of the bacterial symbionts residing in the foregut and midgut of the deep-sea hydrothermal vent shrimp Rimicaris exoculata, whose digestive symbionts had unknown roles?

Core claims
  • Genome-resolved metagenomics of separated foregut and midgut reconstructed 20 MAGs including novel lineages of Hepatoplasmataceae (foregut) and Deferribacteres (midgut). finding
  • Hepatoplasmataceae symbionts have streamlined, reduced genomes capable of using mostly broken-down complex molecules. finding
  • Deferribacteres can degrade complex polymers, synthesize vitamins, and encode numerous flagellar and chemotaxis genes for host-symbiont sensing. mechanism
  • Both symbionts harbor a diverse set of immune system genes favoring holobiont defense. finding
  • Deferribacteres colonize the bacteria-free ectoperitrophic space in direct contact with the host, elongating but not dividing despite possessing the complete genetic machinery for division. finding
  • Digestive symbionts have key communication and defense roles contributing to overall fitness of the Rimicaris holobiont. mechanism
  • A reproducible genome-resolved metagenomics workflow and a collection of MAGs/profiles for Rimicaris gut symbionts are provided as a resource. resource
Experimental setups
Assay System Perturbation Readout Platform
Shotgun genome-resolved metagenomics (assembly + binning into MAGs) Rimicaris exoculata foregut and midgut, three Mid-Atlantic Ridge vent sites (Rainbow, TAG, Snake Pit) none Metagenome-assembled genomes, taxonomy, functional/metabolic gene content (COG/KEGG/CAZymes/CRISPR-Cas) Illumina HiSeq3000, 2×150 bp, TruSeq Nano kit
Fluorescence in situ hybridization (FISH) Rimicaris exoculata midgut and foregut tissue sections (ectoperitrophic space) none Localization/identification of Deferribacteres symbiotic lineage Probe Def1229-Cy3; Zeiss Imager.Z2 microscope with ApoTome.2 and Colibri.7
YOYO-1 DNA-specific fluorescence labeling / chromosome counting Rimicaris exoculata midgut symbionts (sections) none Number of bacterial chromosomes per cell (cell division status) Zeiss LSM 780 confocal microscope, photon counting
Small-subunit rRNA taxonomic profiling (phyloFlash) Rimicaris exoculata foregut and midgut metagenomes none Taxonomic composition of samples / host DNA contamination phyloFlash v3.4, SILVA 138.1
Differential abundance analysis Foregut vs midgut MAGs none Differentially abundant MAGs between organs DESeq2; GeTMM normalization
Phylogenomic / ANI analysis Deferribacteres and Hepatoplasmataceae MAGs and related taxa none Phylogenetic placement and genome similarity GTDB-Tk v1.5.0, IQ-TREE v2.0.3, pyANI/dRep
Key results
  • Reconstructed 21 MAGs (≥60% completion, ≤10% redundancy); dereplication to a final collection of 20 MAGs after removing one Hepatoplasmataceae MAG. 21 then 20 MAGs
  • Shotgun sequencing recovered 677 million reads total across six metagenomes. 677 million reads
  • Average reads per metagenome. 113 million reads/metagenome
  • Per-sample assembly yielded contigs longer than 1 kbp. 103K–147K contigs
  • High-quality reads recruited to assembled contigs. average 57.60%
  • High-quality reads mapping to the final MAG collection were low, attributed to host DNA contamination. 1.75–2.61%
  • Deferribacteres symbionts elongate in the ectoperitrophic space but do not divide despite complete division machinery.
  • Hepatoplasmataceae abundant in foregut and Deferribacteres abundant in midgut, with little prior genetic differentiation between organs.
Key statistics
  • count 677 million reads (total shotgun reads across six digestive-tract metagenomes)
  • mean 113 million reads (average reads per metagenome)
  • other 57.60% (average high-quality reads recruited to assembled contigs)
  • other 1.75–2.61% (high-quality reads mapped to final MAG collection)
  • count 103K to 147K contigs (contigs >1 kbp per sample assembly)
  • other padj < 0.01; |log2FC| 1.5 (significance thresholds for differential MAG abundance (DESeq2))
  • other ≥99% ANI, 50% coverage threshold (dRep dereplication criteria for MAGs)
  • count 20 MAGs (final dereplicated MAG collection (from 21 reconstructed))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a genome-resolved metagenomics workflow (anvi'o, Megahit assembly, CONCOCT binning) to reconstruct MAGs from pooled digestive-tract samples (foregut and midgut) collected at three hydrothermal vent sites, with nine to ten specimens' organs pooled per site per organ type due to small tissue size. Differences in MAG relative abundance between organs were assessed statistically using GeTMM-normalized counts analyzed with DESeq2, applying adjusted p-value and log2-fold-change thresholds to call significance. Most other results (genome statistics, phylogenomics, metabolic pathway completeness, CAZyme and CRISPR annotation) were reported descriptively rather than through inferential hypothesis testing.

Replicationtechnical Sample size9-10 shrimp specimens' organs pooled per site (Rainbow, TAG, Snake Pit) into one foregut and one midgut sample per site; no formal power/sample-size justification given Groupsforegut vs. midgut MAG abundance across three hydrothermal sites Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionDESeq2 adjusted p-values (Benjamini-Hochberg-type FDR, as implemented by default in DESeq2)
Statistical tests used
Test Applied to n Assumptions
DESeq2 (Wald test with GeTMM-normalized counts) differential abundance of MAGs between foregut and midgut samples one pooled metagenomic sample per organ per site (3 sites), i.e., samples derived from pooling 9-10 specimens per site not stated
Approaches that could also have been used
  • Because of tissue scarcity, organs from 9-10 specimens were pooled into a single metagenomic sample per site per organ, yielding one composite sample per condition rather than multiple independent biological replicates.
    Could also: A design with multiple independently sequenced (unpooled) biological replicates per site/organ, or explicit reporting of pooling as a limitation alongside site-level replication — Independent biological replicates would allow variance between individuals to be estimated directly, which can strengthen confidence in differential abundance calls beyond what pooled composite samples permit.
  • MAG differential abundance between organs was assessed with DESeq2 on GeTMM-normalized counts, using combined padj and log2FC thresholds as the significance criterion.
    Could also: Alternative differential-abundance frameworks for compositional metagenomic/microbiome count data, such as ALDEx2, ANCOM-BC, or edgeR — These methods use different normalization and compositional-data assumptions and can be used as complementary approaches to cross-check the robustness of differential abundance calls in metagenomic count data.
  • Significance for MAG abundance differences is summarized via adjusted p-value and fold-change cutoffs, without reporting exact p-values, confidence intervals, or a dispersion/variance measure for the estimates.
    Could also: Reporting exact adjusted p-values alongside effect size estimates with confidence intervals (e.g., DESeq2's log2FC standard errors or shrinkage estimates) — Providing effect size with an interval estimate conveys the precision of the fold-change estimate in addition to whether a fixed threshold was crossed.
  • Only three sites (Rainbow, TAG, Snake Pit) were sampled, each contributing one pooled sample per organ, and the paper notes that site-level differences could not be statistically investigated due to limited organ samples per site.
    Could also: A mixed-effects or hierarchical model treating site as a random effect, applied if additional per-site or per-individual samples were available — Such models can partition variance attributable to site versus organ when sufficient replication exists, which the current pooled single-sample-per-site design does not support.
  • Taxonomic composition per sample was profiled descriptively with phyloFlash/SILVA rather than through a formal statistical comparison across organs or sites.
    Could also: A community-level statistical comparison such as PERMANOVA on beta-diversity distances (e.g., Bray-Curtis) across organ/site groups — This would allow a formal test of whether community composition differs significantly between foregut and midgut or between sites, complementing the descriptive taxonomic summaries presented.
Software: anvi'o v6 · bbduk/bbmap v38.57 · illumina-utils v1.4 · Megahit v1.2.9 · Prodigal v2.6.3 · HMMER v3.2.1 · Bowtie2 v2.4.2 · samtools v1.7 · CONCOCT v1.1.0 · dRep v2.3.2 · DESeq2 · GTDB-Tk v1.5.0 · TrimAl v1.4.1 · IQ-TREE v2.0.3 · KEGG Decoder v1.2.2 · dbCAN2 v2.0.6 · CRISPRCasFinder release 4.2.20 · phyloFlash v3.4

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Figure 3A
C1
Reported
677 million reads total, avg 113 M/metagenome (6 metagenomes)
Reproduced
677,134,974 reads summed over 6 ENA runs ERR7998031-36; avg 112,855,829
exact
C2
Reported
21 MAGs recovered (completion>60%, contamination<10%)
Reproduced
21 MAG FASTA files in the open Ifremer deposit 09_REDUNDANT_MAGs
exact
C3
Reported
20 MAGs after dRep dereplication
Reproduced
partial
C4
Reported
class composition Bacilli7/Deferribacteres4/Campylobacteria3/Clostridia2/Paceibacteria2/Kiritimatiellae1/Gammaproteobacteria1
Reproduced
partial
C5
Reported
Hepatoplasmataceae 5 MAGs, GC 22.42-26.56%, size 0.48-0.83 Mbp
Reproduced
partial
C6
Reported
Deferribacteres 4 MAGs, GC 46.67-47.6%, size 1.25-1.36 Mbp
Reproduced
partial
C7
Reported
103K-147K contigs >1kbp/metagenome; 57.60% avg read recruitment
Reproduced
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 64/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

121.7 k
tokens (I/O) · 7.4 M incl. cache
28 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.