Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genome-wide signatures of convergent evolution in echolocating mammals.

Nature · 2013
not yet assessed 2/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

This is a partial reproduction. The paper's central scientific claim -- a genome-wide signature of convergent molecular evolution (ΔSSLS/Uc statistic finding ~200 of 2,326 orthologous loci with significant convergence support between echolocating bats and dolphin) -- cannot be reproduced from public materials and is scoped as a well-founded drop: the custom statistical pipeline is explicitly 'available on request' in the paper (no public code), and none of the underlying alignments, orthology calls, ancestral reconstructions, or constrained tree topologies were ever deposited (checked Nature SI, NCBI, and Dryad; the one related Dryad record is an independent 2015 reanalysis with its own different alignment set, not the original data). What IS genuinely in-scope, pipeline-derived, and backed by public data is the paper's other real contribution: four newly sequenced bat draft genomes (Eidolon helvum, Pteronotus parnellii, Rhinolophus ferrumequinum, Megaderma lyra), whose raw Illumina WGS reads are on SRA (SRR924356/359/361/427, ~185-195M read pairs / ~33-35 Gbases each) and whose resulting draft assemblies the same authors deposited at NCBI (GCA_000465285.1/405.1/495.1/345.1, scaffold N50s 17-28kb, 130k-190k scaffolds each -- fragmented, shallow-coverage drafts, consistent with the described SOAPdenovo/GapCloser short-read pipeline). We set up a from-scratch reassembly pipeline (sra-tools -> SOAPdenovo2 k=31, matching the paper's assembler) on «our HPC» for the RU's designated accession, SRR924356 (E. helvum), to produce a genuinely gradable comparison against the deposited assembly stats. The pipeline hit and recovered from one real infrastructure snag (fasterq-dump stalled writing many small temp files to «infra» network storage; fixed by redirecting temp I/O to node-local /tmp), and was still in the FASTQ-extraction stage (~80 minutes of an expected multi-hour run) when this result was finalized on explicit instruction to publish an honest partial result now rather than keep waiting/perfecting. No numeric claim is graded in this result because none has a completed comparison yet: the core claim's comparison is impossible (missing code/data, a genuine drop, not a timing issue), and the assembly comparison's reproduced side (our own assembly stats) does not exist yet (a timing issue, not a data/code gap) -- «job» keeps running on «our HPC» independently of this session and a follow-up worker can pick up its output («path» once it appears) to complete the grade. Not attempted at all in this pass: assembling the other 3 bat genomes, and any RAxML run (noted as not meaningful without the authors' actual alignments/topologies).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Using the independent evolution of echolocation in bats and cetaceans as a model of phenotypic convergence, the authors test whether adaptive convergent amino-acid sequence evolution is restricted to a handful of candidate loci or is instead widespread across the genome, and whether such convergence is driven by natural selection rather than neutral homoplasy.

Core claims
  • Genome-wide convergent sequence evolution between echolocating lineages is not rare but widespread and continuously distributed, with signatures consistent with convergence in nearly 200 loci out of 2,326 examined. finding
  • Convergence is commonly driven by natural selection acting on a small number of sites per locus, rather than by neutral processes. mechanism
  • Genes linked to hearing or deafness show strong and significant support for bat-dolphin convergence, consistent with an involvement in echolocation. finding
  • Genes linked to vision/blindness also show convergent signal (though weaker), suggesting convergence associated with a shift in primary sensory modality. finding
  • Sitewise support for convergence (|ΔSSLS|) is robustly correlated with the sitewise strength of natural selection (ω = dN/dS), indicating adaptive rather than neutral convergence. mechanism
  • A maximum-likelihood pipeline comparing sitewise log-likelihood support (SSLS) between the accepted species tree (H0) and forced-monophyly convergence topologies (H1 bat-bat, H2 bat-dolphin) can detect genome-wide sequence convergence. method
  • Four new draft bat genomes (Rhinolophus ferrumequinum, Megaderma lyra, Eidolon helvum, Pteronotus parnellii) and a 22-mammal alignment of 2,326 orthologous CDSs constitute a new resource. resource
  • 17 of the top 5% H2 loci form a single protein-protein interaction network centred on TP63 (p63) and CDK1 (p34), both implicated in cochlea/hair-cell biology. finding
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome shotgun sequencing (paired-end short reads, 500 bp insert libraries) Four bat species: Rhinolophus ferrumequinum, Megaderma lyra, Eidolon helvum, Pteronotus parnellii none Short read sequence data (~33-41 Gb per species); SRA accessions SRR924356, SRR924359, SRR924361, SRR924427 Illumina Hi-Seq 2000 / Illumina Genome Analyzer (Illumina Inc.); sequencing by BGI
De novo genome assembly and scaffolding Raw reads from the four newly sequenced bat genomes none Contigs and scaffolds CLC bio de novo assembler (k-mer 32-50); SOAPdenovo
Homology-based gene prediction and 1-to-1 orthology screening Four draft bat genomes vs. eutherian mammals / Homo sequences none Number of predicted genes per species and count of single-copy orthologous protein-coding nuclear genes shared across all four genomes
Multiple sequence alignment of coding sequences (codon-aware, in frame, ambiguous sites/codons removed) 2,326 orthologous CDSs across 22 mammals (6 bats, bottlenose dolphin Tursiops truncatus, dog, horse, cow, mouse, human and others from Ensembl Release 63) none Alignments of >=450 bp per locus, 805,053 amino acids total MAFFT; Ensembl Release 63
Maximum-likelihood phylogenetic reconstruction and sitewise log-likelihood support (SSLS / ΔSSLS) convergence pipeline 2,326 mammalian orthologous CDS alignments other (topological constraint: echolocating taxa forced into spurious monophyletic clades H1 bat-bat and H2 bat-dolphin vs. species tree H0) Per-site and per-locus ΔSSLS = SSLS(H0) - SSLS(H1/H2); number of loci with mean support for each convergence hypothesis; preferred topology per locus RAxML (soft-polytomy resolution); TreeAnnotator v1.7.4 / BEAST v1.7.4 for majority clade-consensus tree; ancestral states inferred by parsimony
Sitewise dN/dS (omega) selection estimation and correlation with convergence signal Echolocating lineages within the 2,326-locus mammalian CDS alignment none (lineage-specific selection comparison: echolocators vs. other taxa) Sitewise ω; correlation of absolute sitewise ΔSSLS with sitewise ω after correcting for locus identity published software (unspecified)
Least-squares regression of ΔSSLS on sitewise ω per locus, with extrapolation to ω=2 Individual loci from the 2,326-CDS dataset (sites with differing selection pressure between echolocators and other taxa) none (in silico modelling of prolonged diversifying selection, ω=2) Slope and 95% confidence intervals; classification of loci as adaptive convergent (c.i. of slope < 0) or adaptive divergent (c.i. > 0); predicted convergence under strong selection
Gene-ontology annotation plus randomisation/simulation testing (and protein-protein interaction network and human tissue-specific RNA expression database scans) 21 hearing/deafness-linked genes, 75 vision/blindness-linked genes, previously reported echolocation loci, and the top 5% (117 genes) for H2; human orthologues for expression none Mean ΔSSLS of gene sets vs. null distributions (z scores, P values); network membership; tissue-enriched expression DAVID gene ontology database (http://david.abcc.ncifcrf.gov); published protein-protein interaction and human tissue expression databases; randomisation n=1,000; MCMC mixture-model neutral simulations n>=1,000; random alternative topologies n=100
Key results
  • Of 2,326 loci, 824 showed mean ΔSSLS support for H1 (bat-bat convergence) and 392 for H2 (bat-dolphin convergence); signatures consistent with convergence were found in nearly 200 loci. 824 loci (H1); 392 loci (H2); ~200 loci with convergence signatures
  • Hearing/deafness genes had significantly more negative average ΔSSLS than expected by chance for bat-dolphin convergence. z = -0.0194, P < 0.05
  • Vision/blindness genes showed weaker support for convergence under both hypotheses. z = -0.0020, P <= 0.055 and z = -0.0097, P <= 0.09
  • Absolute sitewise ΔSSLS correlated strongly with sitewise ω after correcting for locus identity, under both convergence hypotheses. H1: P = 0.0336; H2: P < 0.001
  • Locus-wise regression classified loci as putatively adaptive convergent versus adaptive divergent based on the 95% c.i. of the slope. H1: 92 adaptive convergent vs 111 adaptive divergent; H2: 59 adaptive convergent vs 212 adaptive divergent
  • Prestin (SLC26A5), a known convergent hearing gene, ranked highly in the genome-wide convergence distribution, validating the pipeline; other hearing genes (ITM2B, SLC4A11 for H1; COCH, ITM2B, ERCC3, OPA1 for H2) and vision genes (LCAT, SLC45A2, RABGGTB, RP1 for H1; JMJD6, SIX, RHO for H2) fell in the top 5%. Prestin rank #43 (H1) and #22 (H2)
  • Across all topologies compared, the species phylogeny H0 was preferred at most loci, with the bat-bat topology H1 preferred next most often. H0 preferred at 1,170 loci (55%); H1 at 548 loci (26%)
  • Hearing and vision genes (including SLC44A2, SLC4A11, ITGAL) showed significantly stronger predicted convergence than background loci when regression models were extrapolated to prolonged diversifying selection (ω=2). ω = 2 extrapolation
Key statistics
  • pvalue z = -0.0194, P < 0.05 (Randomisation test of mean ΔSSLS for hearing/deafness genes under H2 (bat-dolphin convergence))
  • pvalue z = -0.0020, P <= 0.055 and z = -0.0097, P <= 0.09 (Randomisation test of mean ΔSSLS for vision/blindness genes under the two convergence hypotheses)
  • pvalue H1: P = 0.0336; H2: P < 0.001 (Correlation between absolute sitewise ΔSSLS and sitewise ω, corrected for locus identity)
  • pvalue P <= 0.01 in both cases (Randomisation support for loci previously associated with echolocation, under both H1 and H2)
  • count 824 loci (H1); 392 loci (H2) (Loci with mean ΔSSLS supporting each convergence hypothesis)
  • count 92 adaptive convergent / 111 adaptive divergent (H1); 59 adaptive convergent / 212 adaptive divergent (H2) (Locus classification from 95% c.i. of the ΔSSLS-vs-ω regression slope)
  • count 805,053 amino acids in 2,326 orthologous CDSs across 22 mammals; 7,612 1-to-1 orthologues present in all four new bat genomes (Scale of the systematic convergence analysis)
  • count 20,424 (R. ferrumequinum), 20,043 (M. lyra), 20,455 (E. helvum), 20,357 (P. parnellii) (Genes identified by homology-based prediction in the four new bat genomes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper used a maximum-likelihood phylogenetic approach, comparing sitewise log-likelihood support (ΔSSLS) for a fixed species tree (H0) against two alternative convergent topologies (H1, H2) across 2,326 orthologous coding gene alignments spanning 22 mammals. Support for convergence within defined gene sets (hearing-linked, vision-linked, echolocation-associated genes) was tested with a randomization procedure comparing observed mean ΔSSLS values to null distributions built from 1,000 randomizations, and per-locus classification as 'adaptive convergent' or 'adaptive divergent' used least-squares regression of ΔSSLS on sitewise dN/dS (ω) with 95% confidence intervals on the fitted slope. Results were reported as z-scores, p-values (both exact and thresholded), and regression-based confidence intervals, with additional simulation-based null distributions used to check robustness to substitution model choice and neutral evolutionary processes.

Replicationunclear Sample sizeSample sizes given as counts of loci/genes analyzed (2,326 CDS alignments; 21 hearing-linked genes; 75 vision-linked genes) and number of randomizations/simulations (n=1,000; n=100; n≥1,000); no formal statistical power analysis described GroupsGenes/loci linked to hearing or vision (and echolocation) vs. background loci; species tree (H0) vs. alternative convergence topologies (H1: bat-bat, H2: bat-dolphin) Pairingna Randomization/blindingna DispersionCI Exact p-valuesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Randomization/permutation test (observed mean ΔSSLS vs. null distribution, 1,000 randomizations) Comparison of hearing-linked (n=21) and vision-linked (n=75) gene sets, and echolocation-associated loci, for H1 and H2 convergence hypotheses 21 hearing genes; 75 vision genes; 1,000 randomizations not stated
Correlation between sitewise ΔSSLS and sitewise ω, 'correcting for locus identity' Across all sites/loci, for H1 and H2 not explicitly stated not stated
Least-squares linear regression of ΔSSLS on ω per locus, with 95% CI on slope used for classification Classification of loci as adaptive convergent vs. adaptive divergent under H1 and H2 2,326 loci not stated
Simulation-based null distribution (Markov-chain Monte Carlo mixture models) Testing whether convergent signal (ΔSSLS) could arise from neutral processes n≥1,000 simulated neutrally-evolving amino acids not stated
Comparison against random alternative topologies Robustness check for convergence hypotheses H1/H2 n=100 random topologies not stated
Approaches that could also have been used
  • Statistical significance for gene-set convergence (hearing, vision, echolocation genes) was assessed via a custom randomization test comparing observed mean ΔSSLS to a null distribution from 1,000 randomizations, yielding a separate z-score/p-value for each gene set and hypothesis.
    Could also: A formal multiple-testing correction (e.g., Benjamini-Hochberg FDR) applied jointly across the gene-set and hypothesis (H1/H2) comparisons — would provide a family-wise or false-discovery-rate-adjusted view of the several p-values reported across different gene categories and topologies, complementing the individual randomization p-values
  • Loci were classified as putatively 'adaptive convergent' or 'adaptive divergent' based on whether the 95% CI of a per-locus regression slope excluded zero, applied across 2,326 loci genome-wide.
    Could also: A false-discovery-rate procedure (e.g., Benjamini-Hochberg) applied to the per-locus slope test statistics — would give an estimated false-discovery rate for the genome-wide set of loci classified as adaptive convergent/divergent from a single large-scale screen
  • The relationship between sitewise ΔSSLS and sitewise ω was evaluated by correlation 'after correcting for locus identity'.
    Could also: A linear mixed-effects model with locus fitted as a random effect — would explicitly model non-independence of sites nested within the same locus and yield standard errors/CIs that account for this clustering structure
  • Convergence signal (ΔSSLS) was assessed by comparing sitewise maximum-likelihood log-likelihood support between the species tree and alternative topologies.
    Could also: A Bayesian phylogenetic framework (e.g., Bayes factor comparison of topologies) — would produce posterior probabilities for each topology and propagate branch-length and tree uncertainty directly into the convergence measure
  • Selection strength was measured as a sitewise dN/dS (ω) ratio using an unspecified 'published software'.
    Could also: Codon-based branch-site models (e.g., PAML branch-site test or aBSREL) testing explicitly for episodic diversifying selection on pre-specified foreground (echolocating) branches — would offer a complementary, lineage-specific test for positive selection alongside the sitewise ω values already used to link convergence to selection
  • Some p-values for gene-set convergence tests were reported as thresholds (e.g., P≤0.055, P≤0.09) rather than exact values.
    Could also: Reporting exact p-values and/or standardized effect sizes with confidence intervals consistently across all gene-set comparisons — would let readers assess the precision of borderline results and compare effect magnitudes directly across gene sets
Software: CLC bio (de novo assembly) · SOAPdenovo · MAFFT · RAxML · BEAST/TreeAnnotator 1.7.4 · DAVID (gene ontology)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

No individual results have been recorded for this entry yet.

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 31/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴

Nothing was actually compared. ROOM_RESULT.json has claims: [] and claims_graded: false — the headline result (~200 of 2,326 orthologous loci with significant ΔSSLS/Uc convergence support) is unreachable because the implementing code is 'available on request' in the Methods and no alignments, orthology calls, ancestral states or constrained topologies were ever deposited (SI, NCBI and Dryad all checked; doi:10.5061/dryad.16qc5 is Thomas & Hahn's independent 2015 reanalysis on their own 9-mammal set). The blocker is therefore on the authors' side as a reproducibility/materials failure (q4/q5 red), while the only in-scope gradable item — reassembling SRR924356 and comparing to the deposited GCA_000465285.1 (1.84 Gb, scaffold N50 27,684) — is our incompleteness: «job» was still in fasterq-dump when the result was force-finalized. I keep q7/q8 at yellow rather than red deliberately: we observed no numeric contradiction at all, and the rubric's red reserves itself for demonstrated substantive discrepancy, not for absence of evidence. Everything that could be checked (19GB run size, 195,215,120 spots, ~17-19x coverage) matched metadata exactly.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.