Symbiosis genes show a unique pattern of introgression and selection within a Rhizobium leguminosarum species complex.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> 1:1 on the core result. Reproduced the paper's pangenome counts directly from the authors' shipped figshare intermediates (Data.zip v5: 22115 codon-aware ortholog alignments + 17111 pop-structure-corrected SNP matrices on «infra»). EXACT 1:1 on the three headline pangenome numbers: 22115 ortholog groups, 4204 core gene groups (=present in all 196 strains), 17911 accessory. Genes-in->=100-strains = 6542 vs reported 6529 (within-tol, Delta 13 = 0.2%; remainder is the extra bi-allelic/codon SNP-level filter). Total filtered SNP count NOT matched (raw npz columns 1,134,631 incl. multi-allelic-collapsed sites vs reported 441,287 after codon/bi-allelic filtering) - shipped npz cannot expose the filtered count without re-running the SNP filter on alignments (the optional 20%, not attempted). NOT attempted: SPAdes/Prokka/ProteinOrtho upstream assembly (very heavy); introgression scores Table1 (no runnable RapidNJ/traversal script shipped, no single genospecies-label table); Tajima's D Table2; genospecies=5 PCA; repA plasmid grouping. No fabrication flags: C1-C4 derive cleanly and consistently from shipped data.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 79assessed: 2026-06-15 ⛓ da527d6a68ff
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusIs the frequent introgression of symbiosis genes among sympatric rhizobia a special case, or is horizontal gene transfer common for a wide range of genes across sibling Rhizobium leguminosarum species? The paper tests how many genes besides symbiosis genes cross species boundaries.
- ★ The 196 R. leguminosarum sv. trifolii strains constitute a five-species complex (genospecies gsA-gsE) that occur in sympatry but show little recent between-species gene transfer in core or accessory genomes, except for a few highly mobile regions. finding
- ★ 171 genes frequently cross species boundaries and cluster into four linkage blocks; two largest blocks (125 genes) include the symbiosis genes, one block has 43 mainly chromosomal genes, and the last has three core symbiosis-essential genes of variable genomic location. finding
- ★ All introgression events were likely mediated by conjugation, but only the symbiosis linkage blocks displayed overrepresentation of distinct, high-frequency haplotypes (distinct selection signatures). mechanism
- ★ Inter-species introgression is not limited to symbiosis genes and plasmids, but non-symbiosis cases are infrequent. finding
- ★ A novel introgression-scoring method based on gene-tree traversal (counting genospecies shifts) detects introgression events. method
- ★ A method groups introgressed genes by intergenic linkage disequilibrium corrected for population structure rather than relying on a single reference strain's gene order. method
- 196 newly sequenced, de novo assembled R. leguminosarum genomes provide a broad diversity resource for northern European populations. resource
- The species complex has an open pan-genome, with accessory gene count increasing indefinitely as genomes are added. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome shotgun sequencing (de novo assembly) | 196 Rhizobium leguminosarum sv. trifolii strains from white-clover (Trifolium repens) nodules | none | genome assemblies / annotated gene content | Illumina 2x250 bp paired-end (MicrobesNG); SPAdes v3.6.2 assembly |
| Long-read whole-genome sequencing | 8 of the 196 R. leguminosarum strains | none | reference chromosomal/plasmid assemblies for scaffolding | PacBio (Pacific Biosciences) |
| 16S rDNA / rpoB phylogenetic analysis | 196 R. leguminosarum strains plus known genospecies representatives | none | species confirmation and genospecies assignment | — |
| Orthologous gene prediction and pan-genome analysis | 1,468,264 predicted CDS across 196 strains | none | orthologous gene groups, core/pan genome size | Proteinortho v5.16b, Syntenizer3000, clustalo v1.2.0 |
| SNP / variant calling and ANI/PCA analysis | 196 strains, genes present in >=100 strains | none | bi-allelic SNPs, average nucleotide identity, shared SNP proportion | — |
| Gene-tree introgression scoring | orthologous gene groups across 196 strains / 5 genospecies | none | introgression score (number of genospecies shifts) | RapidNJ v2.3.2 (neighbour-joining) |
| Intergenic linkage disequilibrium analysis (population-structure-corrected, Mantel test) | genotype matrix of 196 strains | none | intergenic LD / gene linkage blocks | scipy linalg |
| Plasmid replicon (repABC) identification and nodulation test | 196 genome assemblies; white clover inoculation (strain SM168B) | none / inoculation | plasmid replication groups; pink nodule formation | tblastn (RepA/RepB/RepC queries) |
- – 171 genes were identified that frequently cross species boundaries (introgressing genes) 171 genes
- – Introgressing genes clustered into four linkage blocks: two largest comprised 125 genes (including symbiosis genes), one had 43 mainly chromosomal genes, and one had three genes of variable location 125 / 43 / 3 genes
- – 196 strains clustered into five genospecies recognized at SNP identity above 96% >96% SNP identity
- – Only symbiosis linkage blocks showed overrepresentation of distinct high-frequency haplotypes, indicating distinct selection signatures
- – 196 strains shared 4204 core gene groups versus 17,911 accessory gene groups; pan-genome is open 4204 core / 17,911 accessory
- ▲ Core gene groups had higher median GC content than accessory gene groups
- – One strain (SM168B) carried no symbiosis genes yet still nodulated clover (genes likely lost during processing); SM165B and SM95 had duplicated symbiosis regions
- – Genomes assembled into 10-96 scaffolds, total lengths 6,966,649-8,355,366 bp with 6,642-8,074 annotated genes 6.97-8.36 Mb; 6642-8074 genes
- count 196 strains de novo assembled (out of 249 isolated) (R. leguminosarum sv. trifolii genomes from white-clover nodules)
- count 171 introgressing genes (genes that frequently cross species boundaries)
- count 6529 of 22,115 genes; 441,287 SNPs (genes passing filtering used for ANI/PCA analyses)
- count 4204 core gene groups; 17,911 accessory gene groups; 22,115 total orthologous genes (pan-genome composition)
- count 1,468,264 predicted coding sequences (input to orthologue identification)
- other SNP identity >96% (threshold for recognizing the five genospecies)
- count 3215 conserved chromosomal-backbone genes; 305 housekeeping genes for ANI (scaffolding reference set and ANI gene set)
- other filtering: sequences >50, segregating sites >10, avg pairwise differences >10, ANI>0.7, introgression score >=10 (criteria for trustable top introgressed genes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This comparative genomics study assembled 196 Rhizobium leguminosarum genomes and used phylogenetic gene-tree traversal to compute a custom introgression score for each orthologous gene group, identifying genes that frequently cross species boundaries. Introgressed genes were then clustered into linkage blocks using population structure-corrected intergenic linkage disequilibrium (LD) estimated via a Mantel test on genomic relationship matrices. Population genetic parameters (Tajima's D, nucleotide diversity, pairwise differences, segregating sites) were estimated across the gene set, and species were delineated by SNP-identity thresholding (>96% shared SNPs) and average nucleotide identity of housekeeping genes. Results were reported descriptively, with filtering thresholds defining the set of high-confidence introgressing genes rather than formal hypothesis tests with stated p-values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Custom introgression score (gene-tree traversal depth-first, counting genospecies shifts) | All 22,115 orthologous gene groups; genes with score ≥10 and passing coverage/diversity filters retained as top introgressed | 196 strains; score based on genes present in >50 strains | not stated |
| Mantel test (Pearson correlation between pairwise distance matrices of pairs of genes) | Intergenic linkage disequilibrium estimation between top-introgressed gene pairs, after population structure correction via GRM decorrelation | 196 strains | not stated |
| Neighbor-joining phylogenetic clustering (RapidNJ) | Individual gene trees for each orthologous gene group, used as input for introgression scoring | per-gene variable (presence in ≥50 strains required for introgression analysis) | not stated |
| Tajima's D | Population genetic characterization of gene groups across the 196 strains | 196 strains | not stated |
| PCA on SNP matrix | Visual population structure assessment (Fig. S7b, c); restricted to genes present in ≥100 strains | 6,529 genes, 441,287 SNPs across 196 strains | na |
| Threshold-based SNP identity clustering (>96% shared SNPs) and ANI of 305 housekeeping genes | Genospecies delineation (Fig. 1a, b) | 196 strains, 6,529 genes for SNP matrix | not stated |
-
Gene trees were constructed with neighbor-joining (RapidNJ) and used as the basis for the custom introgression score↳ Could also: Maximum-likelihood or Bayesian phylogenetic methods (e.g., IQ-TREE, RAxML, FastTree) could also have been used to build gene trees — ML/Bayesian methods incorporate substitution-model selection and provide branch-support values (bootstrap, posterior probability), which can make topological comparisons and introgression inferences more statistically grounded; NJ is faster and may be preferred at this scale (>22,000 trees)
-
Introgression was quantified with a custom score counting genospecies-shift events in depth-first gene-tree traversal↳ Could also: Formal D-statistics (ABBA-BABA test) or Patterson's D could also have been applied to test introgression between specific genospecies triplets — D-statistics yield a test statistic with a standard error estimable by jackknife, providing formal significance assessments for specific introgression hypotheses; the custom score is computationally efficient but does not yield a p-value or confidence bound
-
Intergenic LD was estimated using a Mantel test on pairwise GRM-based distance matrices after GRM decorrelation for population structure↳ Could also: Standard r² or D' linkage disequilibrium statistics computed directly on the population structure-corrected pseudo-SNP matrix could also have been used — Direct r²/D' estimates on corrected genotype vectors are widely understood and software-supported (e.g., PLINK); the Mantel approach on distance matrices is a natural extension to whole-gene comparisons but introduces an additional layer of transformation whose properties may be less familiar to readers
-
Genospecies were delineated using a fixed SNP-identity threshold of >96% shared SNPs↳ Could also: Model-based clustering (e.g., STRUCTURE, fastSTRUCTURE, or ADMIXTURE adapted for haploid bacteria) or hierarchical clustering with a statistically chosen cut-height could also have been used — Model-based approaches provide admixture proportions and can formally estimate the number of clusters K; threshold-based methods are transparent and fast but the choice of threshold (96%) is not derived from a statistical criterion
-
Population genetic parameters including Tajima's D were estimated but no formal neutrality test outcomes (p-values) are reported in the methods↳ Could also: Tajima's D significance could also have been assessed against coalescent-simulated null distributions, or composite likelihood ratio tests (e.g., SweeD, OmegaPlus) could also have been applied to detect selective sweeps — Formal significance assessment against simulated nulls accounts for demographic history and unequal sample sizes across genospecies; reporting only the D value without a reference distribution leaves selective interpretation to visual inspection
-
Pan-genome openness was assessed by randomly adding genomes 50 times and plotting core/pan curves↳ Could also: Parametric models (Heaps' law / power-law fit for pan-genome; exponential decay fit for core genome) could also have been fitted to the accumulation curves — Fitting parametric models and reporting the exponent α (Heaps' law) allows quantitative comparison of pan-genome openness across studies and provides a single summary statistic with an uncertainty estimate, complementing the visual curve approach
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32176601 (Cavassim et al. 2020, Microb Genom)
"Symbiosis genes show a unique pattern of introgression and selection within a Rhizobium leguminosarum species complex." DOI 10.1099/mgen.0.000351. Repo: github.com/izabelcavassim/Rhizobium_analysis (authors' own code). Data: SRA PRJNA510726 (raw reads) + figshare 11568894.v5 (shipped intermediates).
In scope (pipeline-derived, attempted)
Counts derivable from the authors' shipped intermediate data (figshare Data.zip: codon-aware ortholog alignments + population-structure-corrected SNP matrices):
- Number of orthologous gene groups (22115) -> pipeline: ProteinOrtho + clustalo (output shipped)
- Core / accessory gene-group split (4204 / 17911) -> presence across 196 strains
- SNP-filtering gene count (6529, genes in >=100 str.) -> custom SNP matrix builder
- Total filtered SNPs (441287) -> custom SNP filter
In scope but NOT attempted (the optional ~20%; harder / under-specified)
- Introgression scores (Table 1; 171 genes; 55% / 2.43%): described as RapidNJ gene trees + depth-first genospecies-transition traversal, but NO runnable script for it is shipped in the repo, and a single per-strain genospecies label table is not shipped.
- Tajima's D selection (Table 2), via dendropy on per-block gene sets.
- Genospecies count = 5 via PCA on the combined SNP matrix.
- repA plasmid replicon grouping (24 groups / 8 major types).
Out of scope (heavy upstream / wet-lab)
- Genome assembly (SPAdes 3.6.2, 196 strains), QUAST, Prokka annotation, ProteinOrtho ortholog inference -> upstream of shipped intermediates, very heavy; reproducing downstream counts from the shipped alignments+SNPs is equally valid (P16).
- Wet-lab: strain isolation, sequencing, PacBio re-sequencing.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Reproduction used the authors' own shipped figshare intermediates (Data.zip v5), giving exact 1:1 matches on the three headline pangenome counts (22115 orthologs, 4204 core, 17911 accessory) and a within-tolerance C4 (6542 vs 6529, 0.2%). The single mismatch, C5 (1134631 vs reported 441287), is on our side / a preprocessing artifact - the shipped npz collapse multi-allelic sites and lack codon context, so a raw count necessarily exceeds the filtered value; it is derivable from the shipped alignments but was not chased, and there is no fabrication signal. Overall solid and explainable, but partial: the paper's central introgression/selection claims (Tables 1-2, PCA) were not reproduced, so q7 is limited and q8 is yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.