Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Symbiosis genes show a unique pattern of introgression and selection within a Rhizobium leguminosarum species complex.

Microb Genom · 2020
L1 79/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
79/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 55% of all assessed papers rank 514 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> 1:1 on the core result. Reproduced the paper's pangenome counts directly from the authors' shipped figshare intermediates (Data.zip v5: 22115 codon-aware ortholog alignments + 17111 pop-structure-corrected SNP matrices on «infra»). EXACT 1:1 on the three headline pangenome numbers: 22115 ortholog groups, 4204 core gene groups (=present in all 196 strains), 17911 accessory. Genes-in->=100-strains = 6542 vs reported 6529 (within-tol, Delta 13 = 0.2%; remainder is the extra bi-allelic/codon SNP-level filter). Total filtered SNP count NOT matched (raw npz columns 1,134,631 incl. multi-allelic-collapsed sites vs reported 441,287 after codon/bi-allelic filtering) - shipped npz cannot expose the filtered count without re-running the SNP filter on alignments (the optional 20%, not attempted). NOT attempted: SPAdes/Prokka/ProteinOrtho upstream assembly (very heavy); introgression scores Table1 (no runnable RapidNJ/traversal script shipped, no single genospecies-label table); Tajima's D Table2; genospecies=5 PCA; repA plasmid grouping. No fabrication flags: C1-C4 derive cleanly and consistently from shipped data.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 79
    assessed: 2026-06-15 ⛓ da527d6a68ff
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Is the frequent introgression of symbiosis genes among sympatric rhizobia a special case, or is horizontal gene transfer common for a wide range of genes across sibling Rhizobium leguminosarum species? The paper tests how many genes besides symbiosis genes cross species boundaries.

Core claims
  • The 196 R. leguminosarum sv. trifolii strains constitute a five-species complex (genospecies gsA-gsE) that occur in sympatry but show little recent between-species gene transfer in core or accessory genomes, except for a few highly mobile regions. finding
  • 171 genes frequently cross species boundaries and cluster into four linkage blocks; two largest blocks (125 genes) include the symbiosis genes, one block has 43 mainly chromosomal genes, and the last has three core symbiosis-essential genes of variable genomic location. finding
  • All introgression events were likely mediated by conjugation, but only the symbiosis linkage blocks displayed overrepresentation of distinct, high-frequency haplotypes (distinct selection signatures). mechanism
  • Inter-species introgression is not limited to symbiosis genes and plasmids, but non-symbiosis cases are infrequent. finding
  • A novel introgression-scoring method based on gene-tree traversal (counting genospecies shifts) detects introgression events. method
  • A method groups introgressed genes by intergenic linkage disequilibrium corrected for population structure rather than relying on a single reference strain's gene order. method
  • 196 newly sequenced, de novo assembled R. leguminosarum genomes provide a broad diversity resource for northern European populations. resource
  • The species complex has an open pan-genome, with accessory gene count increasing indefinitely as genomes are added. finding
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome shotgun sequencing (de novo assembly) 196 Rhizobium leguminosarum sv. trifolii strains from white-clover (Trifolium repens) nodules none genome assemblies / annotated gene content Illumina 2x250 bp paired-end (MicrobesNG); SPAdes v3.6.2 assembly
Long-read whole-genome sequencing 8 of the 196 R. leguminosarum strains none reference chromosomal/plasmid assemblies for scaffolding PacBio (Pacific Biosciences)
16S rDNA / rpoB phylogenetic analysis 196 R. leguminosarum strains plus known genospecies representatives none species confirmation and genospecies assignment
Orthologous gene prediction and pan-genome analysis 1,468,264 predicted CDS across 196 strains none orthologous gene groups, core/pan genome size Proteinortho v5.16b, Syntenizer3000, clustalo v1.2.0
SNP / variant calling and ANI/PCA analysis 196 strains, genes present in >=100 strains none bi-allelic SNPs, average nucleotide identity, shared SNP proportion
Gene-tree introgression scoring orthologous gene groups across 196 strains / 5 genospecies none introgression score (number of genospecies shifts) RapidNJ v2.3.2 (neighbour-joining)
Intergenic linkage disequilibrium analysis (population-structure-corrected, Mantel test) genotype matrix of 196 strains none intergenic LD / gene linkage blocks scipy linalg
Plasmid replicon (repABC) identification and nodulation test 196 genome assemblies; white clover inoculation (strain SM168B) none / inoculation plasmid replication groups; pink nodule formation tblastn (RepA/RepB/RepC queries)
Key results
  • 171 genes were identified that frequently cross species boundaries (introgressing genes) 171 genes
  • Introgressing genes clustered into four linkage blocks: two largest comprised 125 genes (including symbiosis genes), one had 43 mainly chromosomal genes, and one had three genes of variable location 125 / 43 / 3 genes
  • 196 strains clustered into five genospecies recognized at SNP identity above 96% >96% SNP identity
  • Only symbiosis linkage blocks showed overrepresentation of distinct high-frequency haplotypes, indicating distinct selection signatures
  • 196 strains shared 4204 core gene groups versus 17,911 accessory gene groups; pan-genome is open 4204 core / 17,911 accessory
  • Core gene groups had higher median GC content than accessory gene groups
  • One strain (SM168B) carried no symbiosis genes yet still nodulated clover (genes likely lost during processing); SM165B and SM95 had duplicated symbiosis regions
  • Genomes assembled into 10-96 scaffolds, total lengths 6,966,649-8,355,366 bp with 6,642-8,074 annotated genes 6.97-8.36 Mb; 6642-8074 genes
Key statistics
  • count 196 strains de novo assembled (out of 249 isolated) (R. leguminosarum sv. trifolii genomes from white-clover nodules)
  • count 171 introgressing genes (genes that frequently cross species boundaries)
  • count 6529 of 22,115 genes; 441,287 SNPs (genes passing filtering used for ANI/PCA analyses)
  • count 4204 core gene groups; 17,911 accessory gene groups; 22,115 total orthologous genes (pan-genome composition)
  • count 1,468,264 predicted coding sequences (input to orthologue identification)
  • other SNP identity >96% (threshold for recognizing the five genospecies)
  • count 3215 conserved chromosomal-backbone genes; 305 housekeeping genes for ANI (scaffolding reference set and ANI gene set)
  • other filtering: sequences >50, segregating sites >10, avg pairwise differences >10, ANI>0.7, introgression score >=10 (criteria for trustable top introgressed genes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This comparative genomics study assembled 196 Rhizobium leguminosarum genomes and used phylogenetic gene-tree traversal to compute a custom introgression score for each orthologous gene group, identifying genes that frequently cross species boundaries. Introgressed genes were then clustered into linkage blocks using population structure-corrected intergenic linkage disequilibrium (LD) estimated via a Mantel test on genomic relationship matrices. Population genetic parameters (Tajima's D, nucleotide diversity, pairwise differences, segregating sites) were estimated across the gene set, and species were delineated by SNP-identity thresholding (>96% shared SNPs) and average nucleotide identity of housekeeping genes. Results were reported descriptively, with filtering thresholds defining the set of high-confidence introgressing genes rather than formal hypothesis tests with stated p-values.

Replicationbiological Sample size196 strains isolated from white-clover root nodules across Denmark, France, and UK sampling sites; no formal power analysis reported GroupsFive genospecies (gsA–gsE) of R. leguminosarum sv. trifolii; no treatment/control groups — observational comparative genomics Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Custom introgression score (gene-tree traversal depth-first, counting genospecies shifts) All 22,115 orthologous gene groups; genes with score ≥10 and passing coverage/diversity filters retained as top introgressed 196 strains; score based on genes present in >50 strains not stated
Mantel test (Pearson correlation between pairwise distance matrices of pairs of genes) Intergenic linkage disequilibrium estimation between top-introgressed gene pairs, after population structure correction via GRM decorrelation 196 strains not stated
Neighbor-joining phylogenetic clustering (RapidNJ) Individual gene trees for each orthologous gene group, used as input for introgression scoring per-gene variable (presence in ≥50 strains required for introgression analysis) not stated
Tajima's D Population genetic characterization of gene groups across the 196 strains 196 strains not stated
PCA on SNP matrix Visual population structure assessment (Fig. S7b, c); restricted to genes present in ≥100 strains 6,529 genes, 441,287 SNPs across 196 strains na
Threshold-based SNP identity clustering (>96% shared SNPs) and ANI of 305 housekeeping genes Genospecies delineation (Fig. 1a, b) 196 strains, 6,529 genes for SNP matrix not stated
Approaches that could also have been used
  • Gene trees were constructed with neighbor-joining (RapidNJ) and used as the basis for the custom introgression score
    Could also: Maximum-likelihood or Bayesian phylogenetic methods (e.g., IQ-TREE, RAxML, FastTree) could also have been used to build gene trees — ML/Bayesian methods incorporate substitution-model selection and provide branch-support values (bootstrap, posterior probability), which can make topological comparisons and introgression inferences more statistically grounded; NJ is faster and may be preferred at this scale (>22,000 trees)
  • Introgression was quantified with a custom score counting genospecies-shift events in depth-first gene-tree traversal
    Could also: Formal D-statistics (ABBA-BABA test) or Patterson's D could also have been applied to test introgression between specific genospecies triplets — D-statistics yield a test statistic with a standard error estimable by jackknife, providing formal significance assessments for specific introgression hypotheses; the custom score is computationally efficient but does not yield a p-value or confidence bound
  • Intergenic LD was estimated using a Mantel test on pairwise GRM-based distance matrices after GRM decorrelation for population structure
    Could also: Standard r² or D' linkage disequilibrium statistics computed directly on the population structure-corrected pseudo-SNP matrix could also have been used — Direct r²/D' estimates on corrected genotype vectors are widely understood and software-supported (e.g., PLINK); the Mantel approach on distance matrices is a natural extension to whole-gene comparisons but introduces an additional layer of transformation whose properties may be less familiar to readers
  • Genospecies were delineated using a fixed SNP-identity threshold of >96% shared SNPs
    Could also: Model-based clustering (e.g., STRUCTURE, fastSTRUCTURE, or ADMIXTURE adapted for haploid bacteria) or hierarchical clustering with a statistically chosen cut-height could also have been used — Model-based approaches provide admixture proportions and can formally estimate the number of clusters K; threshold-based methods are transparent and fast but the choice of threshold (96%) is not derived from a statistical criterion
  • Population genetic parameters including Tajima's D were estimated but no formal neutrality test outcomes (p-values) are reported in the methods
    Could also: Tajima's D significance could also have been assessed against coalescent-simulated null distributions, or composite likelihood ratio tests (e.g., SweeD, OmegaPlus) could also have been applied to detect selective sweeps — Formal significance assessment against simulated nulls accounts for demographic history and unequal sample sizes across genospecies; reporting only the D value without a reference distribution leaves selective interpretation to visual inspection
  • Pan-genome openness was assessed by randomly adding genomes 50 times and plotting core/pan curves
    Could also: Parametric models (Heaps' law / power-law fit for pan-genome; exponential decay fit for core genome) could also have been fitted to the accumulation curves — Fitting parametric models and reporting the exponent α (Heaps' law) allows quantitative comparison of pan-genome openness across studies and provides a single summary statistic with an uncertainty estimate, complementing the visual curve approach
Software: SPAdes 3.6.2 · QUAST 4.6.3 · Prokka 1.12 · Proteinortho 5.16b · clustalo 1.2.0 · RapidNJ 2.3.2 · Python/dendropy · Python/scipy (linalg) · Custom Python scripts (Jigome, Syntenizer3000)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
8
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

10.6084/m9.figshare.11568894.v5 DOI in Article (http://semanticscience.org/resource/SIO_001029)
also used by 1 paper:
PRJNA510726 BioProject in Acknowledgments (http://purl.org/orb/Acknowledgments)
also used by 1 paper:
SAMN10617942 BioSamples in Acknowledgments (http://purl.org/orb/Acknowledgments)
also used by 1 paper:
SAMN10618137 BioSamples in Acknowledgments (http://purl.org/orb/Acknowledgments)
also used by 1 paper:

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-32176601 (Cavassim et al. 2020, Microb Genom)

"Symbiosis genes show a unique pattern of introgression and selection within a Rhizobium leguminosarum species complex." DOI 10.1099/mgen.0.000351. Repo: github.com/izabelcavassim/Rhizobium_analysis (authors' own code). Data: SRA PRJNA510726 (raw reads) + figshare 11568894.v5 (shipped intermediates).

In scope (pipeline-derived, attempted)

Counts derivable from the authors' shipped intermediate data (figshare Data.zip: codon-aware ortholog alignments + population-structure-corrected SNP matrices):

  • Number of orthologous gene groups (22115) -> pipeline: ProteinOrtho + clustalo (output shipped)
  • Core / accessory gene-group split (4204 / 17911) -> presence across 196 strains
  • SNP-filtering gene count (6529, genes in >=100 str.) -> custom SNP matrix builder
  • Total filtered SNPs (441287) -> custom SNP filter

In scope but NOT attempted (the optional ~20%; harder / under-specified)

  • Introgression scores (Table 1; 171 genes; 55% / 2.43%): described as RapidNJ gene trees + depth-first genospecies-transition traversal, but NO runnable script for it is shipped in the repo, and a single per-strain genospecies label table is not shipped.
  • Tajima's D selection (Table 2), via dendropy on per-block gene sets.
  • Genospecies count = 5 via PCA on the combined SNP matrix.
  • repA plasmid replicon grouping (24 groups / 8 major types).

Out of scope (heavy upstream / wet-lab)

  • Genome assembly (SPAdes 3.6.2, 196 strains), QUAST, Prokka annotation, ProteinOrtho ortholog inference -> upstream of shipped intermediates, very heavy; reproducing downstream counts from the shipped alignments+SNPs is equally valid (P16).
  • Wet-lab: strain isolation, sequencing, PacBio re-sequencing.
C1
Reported
22115 ortholog gene groups
Reproduced
22115
exact
C2
Reported
4204 core gene groups
Reproduced
4204
exact
C3
Reported
17911 accessory gene groups
Reproduced
17911
exact
C4
Reported
6529 genes passing SNP filter (>=100 strains)
Reproduced
6542
within tolerance
C5
Reported
441287 SNPs passing filter
Reproduced
1134631 (raw, unfiltered for codon/bi-allelic)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 79/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Reproduction used the authors' own shipped figshare intermediates (Data.zip v5), giving exact 1:1 matches on the three headline pangenome counts (22115 orthologs, 4204 core, 17911 accessory) and a within-tolerance C4 (6542 vs 6529, 0.2%). The single mismatch, C5 (1134631 vs reported 441287), is on our side / a preprocessing artifact - the shipped npz collapse multi-allelic sites and lack codon context, so a raw count necessarily exceeds the filtered value; it is derivable from the shipped alignments but was not chased, and there is no fabrication signal. Overall solid and explainable, but partial: the paper's central introgression/selection claims (Tables 1-2, PCA) were not reproduced, so q7 is limited and q8 is yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

97.5 k
tokens (I/O) · 6.1 M incl. cache
12 min
runtime · 0.01 CPU-h
0.1 GB
peak RAM
1
HPC jobs
hummel
machine