Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Horizontal transmission enables flexible associations with locally adapted symbiont strains in deep-sea hydrothermal vent symbioses.

Proc Natl Acad Sci U S A · 2022
L1 66/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
66/100
Reproducibility score
0.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 28% of all assessed papers rank 830 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL (strong). Reproduction target = Dataset S2 host coding transcriptome (sd02.xlsx). Confirmed S2 reports the TransDecoder CDS (median 447-513bp, GC 51-53%), not raw transcripts. Recomputed count/size/GC/median + BUSCO(mollusca_odb10) on the TransDecoder CDS of the 3 deposited Alviniconcha TSA assemblies (GJGJ/GJGH/GJGI). RESULT for all 3 species: size_Mb (within -4.7..-6.5%), GC% (within +0.6..+0.7 pts) and BUSCO complete% (within -1.2..-1.7 pts) ALL within tolerance — a tight, independent match (9/15 metrics within-tol). n_transcripts is 14-17.5% LOWER and median_len 17-21% HIGHER; both are mechanistically explained by the single documented step we did not run — the authors' --retain_pfam_hits/--retain_blastp_hits homology rescue (needs uniref90/Pfam/nt), which adds short DB-supported ORFs (raising count, lowering median). Dataset S2 is fully derivable from the deposited data along the documented pipeline; residual deltas point to the un-run homology step, NOT fabrication — fabrication check PASSES. NOT attempted: nautilei (Ifremeria TSA unretrievable), de-novo Trinity reassembly, blobtools euk-filter, and the downstream FST/PCA/RDA population-genetic inferences (out of scope). Verdict is PROVISIONAL — must be independently checked.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 99ff0d0d1aa3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests the long-standing but previously untested hypothesis that horizontal transmission allows hydrothermal vent snail hosts (Alviniconcha and Ifremeria) to flexibly associate with locally adapted strains of their chemosynthetic bacterial symbionts, conferring fitness advantages across spatially variable vent habitats.

Core claims
  • Host populations of Alviniconcha and Ifremeria are not genetically differentiated across an ~800-km geographic gradient finding
  • Symbiont populations (Epsilon, Gamma1, Ifr-SOX) are clearly structured by geography/vent location, unlike their hosts finding
  • Host mitochondrial and symbiont core-genome phylogenies are topologically discordant, indicating symbionts are frequently environmentally (horizontally) acquired rather than vertically transmitted finding
  • A subset of symbiont genetic variants shows signatures of local adaptation (genotype-environment associations) to vent geochemistry, depth, and year finding
  • The Gamma1 symbiont exhibits host-species-specific strains, with the effect of host taxon on symbiont genetic variation exceeding the effect of geography by about a factor of three finding
  • Environment is a stronger driver of strain-level symbiont composition than host genetics finding
  • Symbiont populations differ in gene content (1.33-3.84% of protein-coding regions) between geographic locations and, for Gamma1, between host species finding
  • Population genomic methods (transcriptome/pangenome assembly, variant calling, RDA, Mantel tests) were applied to jointly assess host and symbiont genetic structure and its environmental correlates method
Experimental setups
Assay System Perturbation Readout Platform
RNA-seq-based transcriptome assembly and variant calling (host population genomics) gill tissue of Alviniconcha boucheti, A. kojimai, A. strummeri, and Ifremeria nautilei none SNPs, FST, principal component ordination of host population structure
Metagenomic sequencing and metagenome-assembled genome (MAG)/pangenome reconstruction chemosynthetic bacterial symbionts (Sulfurimonas 'Epsilon', Gamma1, GammaLau, Thiolapillus 'Ifr-SOX', Methylomonadaceae 'Ifr-MOX') associated with the four snail species none symbiont pangenome content, SNPs, FST, population structure
Phylogenomic tree reconstruction with ultrafast bootstrap and SH-aLRT support host mitochondrial genomes vs. symbiont core genomes none tree topology concordance/discordance between host and symbiont lineages
Redundancy analysis (RDA) for genotype-environment association Epsilon, Gamma1, and Ifr-SOX symbiont populations environmental gradient (hydrothermal fluid chemistry, depth, year; not an experimental manipulation) percent of variants significantly correlated with environmental predictors, r2 adjusted variance explained
Gene content/pangenome presence-absence analysis symbiont populations from different vent localities and host species none presence/absence of genes by functional category, percent differentially preserved genes PanPhlAn
Mantel tests and FST/ordination analyses paired host and symbiont population genetic datasets none correlation between genetic distance and geography/environment, corrected for host genetics
Key results
  • Host populations show low genetic differentiation among vent localities across all four species FST = 0.0400 to 0.1579
  • Symbiont populations are strongly structured between vent locations/regions FST = 0.1012 to 0.8682
  • Host mitochondrial and symbiont phylogenies are discordant across all species
  • Genotype-environment associations (RDA) significant for most symbiont-host pairs, indicating local adaptation 3.42-7.57% of variants, P ≤ 0.05; models explain 8.88-13.66% of variance
  • Gamma1 symbiont strain composition is strongly associated with host species, exceeding the geographic effect r = 0.6844-0.7424, P ≤ 0.0002; ~3-fold stronger than geography
  • Genotype-site associations for symbionts remain significant after correcting for host genetic variation P ≤ 0.0285, r = 0.2861 to 0.8512
  • Symbiont populations differ in gene content between geographic regions and host species 1.33-3.84% of protein-coding genes differentially preserved
  • Only 4 SNP sites in A. kojimai showed FST outlier signal in hosts; no adaptive variation detected in other host species q ≤ 0.05
Key statistics
  • fold_change FST = 0.0400 to 0.1579 (host population differentiation across vent localities)
  • fold_change FST = 0.1012 to 0.8682 (symbiont population differentiation across vent regions)
  • correlation r = 0.2861 to 0.8512, P ≤ 0.0285 (genotype-site association for symbionts corrected for host genetic variation)
  • correlation r = 0.6844 to 0.7424, P ≤ 0.0002 (Gamma1 symbiont genetic variation associated with host species (A. kojimai vs A. strummeri))
  • pvalue P = 0.001, r2adj = 38.54% (combined RDA model for Gamma1 symbiont with host species and environment as predictors)
  • other 3.42% to 7.57% of variants significant (P ≤ 0.05) (proportion of symbiont genetic variants correlated with environmental predictors in RDA models)
  • count 1,655 to 9,185 SNPs per host species; 1,716 to 7,741 variants per symbiont species (number of SNPs/variants used in host and symbiont population genetic analyses)
  • other 1.33% to 3.84% of protein-coding regions differentially preserved (gene content variation between symbiont populations from different vent localities or host species)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study applied population genomic methods to characterize host and symbiont genetic structure in four deep-sea snail species sampled across six hydrothermal vent locations. Host structure was assessed via FST and principal-component analyses of transcriptomic SNPs, while symbiont structure was additionally evaluated using redundancy analyses (RDA) conditioned on geography and constrained by environmental predictors. Phylogenetic concordance between host mitochondrial and symbiont core-genome trees was evaluated to infer transmission mode, and signatures of selection were investigated through FST outlier detection and pN/pS ratios. Gene content variation among symbiont populations was quantified using PanPhlAn.

Replicationbiological Sample size11–25 gill samples per host species for transcriptomics; 4–59 high-quality metagenome-assembled genomes per symbiont phylotype; no formal power analysis described GroupsHost and symbiont populations from six vent locations (Kilo Moana, Tow Cam, Tahi Moana, ABE, Tu'i Malila, Niua South) across four snail species (A. boucheti, A. kojimai, A. strummeri, I. nautilei) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFDR via q-value (q ≤ 0.05) for FST outlier detection; permutation-based p ≤ 0.05 for RDA axes and marginal environmental effects without stated family-wise correction across models
Statistical tests used
Test Applied to n Assumptions
FST (fixation index) Population differentiation among host and symbiont populations across vent localities 11–25 host individuals per species (1,655–9,185 SNPs); 4–59 metagenome-assembled genomes per symbiont phylotype (1,716–7,741 variants) not stated
Principal-component analysis (PCA) / ordination Host population genetic structure (Fig. 1B) 11–25 individuals per species not stated
Mantel correlation Effect of geography on host population structure; genotype–site and genotype–host-species associations for symbionts not stated
Redundancy analysis (RDA) / constrained ordination (CAP) Genotype–environment associations for symbiont populations (Fig. 2); variance partitioning across hydrothermal fluid chemistry, depth, and year not stated
Ultrafast bootstrapping and Shimodaira–Hasegawa approximate-likelihood ratio test (SH-aLRT) Phylogenetic node support for host mitochondrial and symbiont core-genome trees (Fig. 1C) not stated
pN/pS ratio analysis Characterization of candidate adaptive loci in symbiont species (Dataset S12) not stated
FST outlier detection with q-value threshold Identification of putatively adaptive variants in host species (Dataset S5) 1,655–9,185 SNPs per host species not stated
Approaches that could also have been used
  • Mantel tests were used to assess associations between pairwise genetic, geographic, and environmental distance matrices
    Could also: Multiple matrix regression with randomization (MMRR) or partial Mantel tests explicitly controlling for a third distance matrix could also be applied — Partial Mantel tests and MMRR allow simultaneous statistical control of one distance matrix (e.g., geographic) while testing another (e.g., environmental), enabling finer separation of isolation-by-distance from isolation-by-environment contributions to symbiont structure
  • RDA conditioned on geography was used to identify genotype–environment associations and candidate adaptive loci
    Could also: Latent factor mixed models (LFMM) or BayPass, which explicitly model neutral population structure as a random effect, could also be used for genotype–environment association testing — Methods that incorporate a neutral population-structure covariate directly into the association model can reduce false positives from shared demographic history; their results could complement RDA-identified candidates through cross-validation
  • FST outlier detection with a q-value threshold was applied to identify candidate loci under selection in host populations
    Could also: OutFLANK or BayeScan, which derive neutral FST expectations empirically from the observed locus distribution, could also serve as complementary outlier methods — Different FST-outlier approaches make distinct assumptions about demographic history; running multiple methods and reporting the intersection or union of candidates is a common practice for robustness assessment
  • Concordance between host mitochondrial and symbiont core-genome phylogenies was assessed qualitatively by visual inspection of tree topologies
    Could also: Formal cophylogenetic tests (e.g., PACo, ParaFit) could also quantify the degree of topological concordance statistically — Statistical cophylogenetic methods provide a formal null-hypothesis test of random host–symbiont associations and can estimate the proportion of codivergence events versus host switching, adding a quantitative layer to the qualitative comparison
  • pN/pS ratios were used to characterize selection pressure on candidate adaptive loci within symbiont populations
    Could also: McDonald-Kreitman tests or branch/site dN/dS models (requiring an outgroup) could also be applied to detect positive selection — The McDonald-Kreitman framework leverages within-population polymorphism alongside between-lineage divergence to formally test for adaptive evolution and can distinguish positive selection from relaxed purifying constraint, complementing pN/pS-based inference
  • Multiple RDA models and Mantel tests were conducted across symbiont species and host–symbiont pairs; no family-wise correction was reported across this set of tests
    Could also: A Benjamini-Hochberg FDR or Bonferroni correction applied across the full family of permutation-based RDA and Mantel p-values could also be reported — When many related hypothesis tests are conducted across species and environmental predictors, a family-wise correction reduces the expected number of spurious significant results; reporting corrected alongside uncorrected thresholds is a common practice in multi-species genomic comparisons
Software: PanPhlAn

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35349333

Paper: Breusing et al. 2022, PNAS 119(14):e2115608119. "Horizontal transmission enables flexible associations with locally adapted symbiont strains in deep-sea hydrothermal vent symbioses." Authors' analysis code: https://github.com/cbreusing/Provannid_host-symbiont_popgen (default branch main, last push 2024-04-09, no license file). Shell/R/Python scripts, no own data shipped — wraps standard third-party tools (Trinity, metaSPAdes, Panaroo, IQ-TREE, Angsd, FreeBayes, BUSCO, CheckM, fastANI, PanPhlAn ...). Third-party tool named in BRIEF: https://github.com/harvardinformatics/TranscriptomeAssemblyTools — used only for read-QC (FilterUncorrectabledPEfastq.py / RemoveFastqcOverrepSequenceReads.py) after Rcorrector. Reproducing a third-party tool on the paper's data is explicitly valid (P16).

Datasets the paper relies on

accession what it actually is (ENA-confirmed) role
PRJNA523619 130 Illumina WGS metagenomic runs + 3 Oxford Nanopore + 4 genomic symbiont metagenomes (DNA-seq)
PRJNA526236 66 RNA-seq metatranscriptomic runs (A. boucheti 21, A. kojimai 27, A. strummeri 18) host gill RNA-seq
PRJNA741492 "Ifremeria nautilei gill transcriptome" (ENA read_run empty; TSA/genome record) host (Ifremeria) transcriptome
TSA: GJGJ/GJGH/GJGI/DXJZ 01 deposited host transcriptome assemblies (Trinity CDS) the assembled product of the host pipeline

Note: the BRIEF lists only PRJNA523619 (the metagenome project). The host RNA-seq lives in PRJNA526236; the assembled host transcriptomes are deposited as TSA accessions.

Pipeline-derived results (candidate reproduction targets)

IN SCOPE — attempted

id reported result pipeline tractability
S2-transcripts per-species transcript counts (24,176–35,654) Trinity→CD-HIT→TransDecoder→filter; stats recomputable from deposited TSA HIGH — recompute stats on deposited assembly (no de-novo assembly needed)
S2-size per-species assembly size 20.68–28.68 Mb same HIGH
S2-gc per-species GC% (51.21–53.53) same HIGH
S2-medlen per-species median contig length (447–513 bp) same HIGH
S2-busco per-species BUSCO completeness 30.3–55.4% (mollusca_odb10, tran) BUSCO on deposited TSA MEDIUM — one SLURM job
meta-N metagenomic libraries 192 prepared / 66 excluded / 126 retained counting HIGH — metadata concordance (vs ENA 130 Illumina)

IN SCOPE — harder, attempt if compute/time permit (80% is a floor)

id reported result pipeline blocker
S2-denovo re-assemble ≥1 host transcriptome from raw RNA reads full Trinity pipeline multi-day; needs custom symbiont ref genomes for bbsplit + uniref90/Pfam DBs
S7-checkm per-MAG CheckM completeness/contamination (~115 MAGs) metaSPAdes→binning→CheckM heavy; needs MAGs deposited as downloadable genomes
S7-ani per-MAG fastANI to reference (e.g. 78–99.9%) fastANI depends on MAG availability
S8-pangenome symbiont pangenome size + core/accessory gene counts Panaroo very heavy

OUT OF SCOPE (not a pipeline reproduction of a pinnable computational value)

  • Wet-lab: DNA/RNA extraction, library prep, sequencing, sample collection (Dataset S1 geochemistry).
  • FST / PCA / RDA / Mantel / pN-pS population-genetic inferences: derived, multi-stage, depend on full upstream assembly+variant-calling stack; recorded as reported but not attempted unless the floor targets complete and compute remains.

Strategy

«our HPC» is the only compute («host» forbidden). All downloads on front node → «infra». Primary 1:1 evidence: recompute the published host-transcriptome statistics (Dataset S2) directly from the deposited TSA assemblies — this both reproduces the reported numbers and acts as a fabrication check (recomputed vs tabulated). BUSCO independently reproduces the completeness column. De-novo re-assembly is the stretch goal.

S2-size-boucheti
Reported
20.68 Mb
Reproduced
19.71 Mb (-4.7%)
within tolerance
S2-gc-boucheti
Reported
53.53%
Reproduced
54.12% (+0.59)
within tolerance
S2-busco-boucheti
Reported
45.4% complete
Reproduced
44.0% (-1.4)
within tolerance
S2-transcripts-boucheti
Reported
24176 CDS
Reproduced
20739 (-14.2%)
partial
S2-median-boucheti
Reported
513 bp
Reproduced
600 bp (+17%)
partial
S2-size-kojimai
Reported
28.68 Mb
Reproduced
26.81 Mb (-6.5%)
within tolerance
S2-gc-kojimai
Reported
53.03%
Reproduced
53.73% (+0.7)
within tolerance
S2-busco-kojimai
Reported
55.4% complete
Reproduced
54.2% (-1.2)
within tolerance
S2-transcripts-kojimai
Reported
33071 CDS
Reproduced
27270 (-17.5%)
partial
S2-median-kojimai
Reported
507 bp
Reproduced
615 bp (+21.3%)
did not match
S2-size-strummeri
Reported
26.07 Mb
Reproduced
24.44 Mb (-6.3%)
within tolerance
S2-gc-strummeri
Reported
53.00%
Reproduced
53.72% (+0.72)
within tolerance
S2-busco-strummeri
Reported
50.5% complete
Reproduced
48.8% (-1.7)
within tolerance
S2-transcripts-strummeri
Reported
30488 CDS
Reproduced
25362 (-16.8%)
partial
S2-median-strummeri
Reported
504 bp
Reproduced
597 bp (+18.5%)
partial
S2-nautilei-all
Reported
35654 CDS / 22.64 Mb / 51.21% / 447 bp / 30.3% BUSCO
Reproduced
not attempted — Ifremeria TSA (PRJNA741492) not retrievable from ENA (DXJZ01/DXJZ02/GJGK01 empty)
partial
meta-N-retained
Reported
126 retained (192 prepared, 66 excluded)
Reproduced
130 Illumina metagenomic runs present in ENA PRJNA523619
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 66/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

149.3 k
tokens (I/O) · 6.4 M incl. cache
16 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.