Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

The evolution of sexual signaling is linked to odorant receptor tuning in perfume-collecting orchid bees.

Nat Commun · 2020
L1 93/100 PQI 98
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 85% of all assessed papers rank 154 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1. The authors' own repo (pbrec/popgen-popchem @ 2ca05c9) ships the derived data matrices, so the downstream pipeline-derived results run directly without re-deriving from SRA reads. All 4 in-scope claims reproduced on «our HPC» (conda R 4.5.3 + SNPRelate/vegan/ecodist): moments AIC weight=1 (exact), GBS PCA PC1+PC2=3.81% which is <4% (exact bound), perfume ANOSIM R=0.7751~0.8 & p=0.001 (within-tol), and the 306-individual perfume filter (exact). NOT attempted: (1) genome-wide FST/pi/Tajima's D sliding-window scan (genome_analysis.R) because its VCFs are not in the repo and the script says to request them by email -> data_restricted; (2) Or41 functional electrophysiology (or41_functional_plotting.R, Fig 4) which reads Fig-4 Source Data not in the repo and is wet-lab -> non_pipeline. Minor honesty flags: a shipped-code typo in gbs_analysis.R's filename (double .txt) had to be corrected to run; the paper cites 16,369 SNPs while the shipped PCA matrix has 5428 (MAF>=0.05 set); the shipped ANOSIM code filters to 270 individuals (top-50 compounds) vs the paper's headline 306 (looser filter) but both give R=0.8,p=0.001. No fabrication detected -- all reproduced values derive from shipped data+code.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 93
    assessed: 2026-06-16 ⛓ f2f3347434dc
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Do specific odorant receptor (OR) genes facilitate the divergence of sexual chemical signaling (perfume communication) and reproductive isolation during the early speciation of the orchid bee lineages Euglossa dilemma and E. viridissima?

Core claims
  • E. dilemma and E. viridissima are reproductively isolated genetically distinct lineages despite low genome-wide differentiation. finding
  • Perfume chemistry is species-specific, driven mainly by lineage-specific major compounds HNDB (E. dilemma) and L97 (E. viridissima). finding
  • Perfume differentiation coincides with two species-specific selective sweeps harboring tandem arrays of odorant receptor genes (43 ORs total). finding
  • The odorant receptor Or41 evolved under positive selection in E. dilemma (selective sweep) but purifying selection in E. viridissima. mechanism
  • The derived Or41 variant in E. dilemma is specifically tuned to its major perfume compound, while the ancestral E. viridissima variant is broadly tuned to multiple odorants. finding
  • OR evolution likely contributed to the divergence of sexual communication and pre-mating reproductive barriers in natural populations. mechanism
  • Functional in vitro assays of OR variants link OR genotype to odorant tuning phenotype. method
  • Net interspecific differentiation (∆FST') leveraging intraspecific population structure identifies islands of divergence. method
Experimental setups
Assay System Perturbation Readout Platform
SNP genotyping / population genetics (PCA, ADMIXTURE, f4-test, demographic modeling) 232 male orchid bees (E. dilemma and E. viridissima) across Central America none genetic differentiation, population structure, admixture (16,369 SNPs)
Whole-genome resequencing / genome-wide divergence scan 30 males from three genetic lineages (n=10 each Ed north, Ed south, Ev) none FST, ∆FST', Dxy, nucleotide diversity (π), LD, selective sweeps (CLR) E. dilemma reference genome
Gas chromatography–mass spectrometry (GC–MS) of perfume chemistry 384 male orchid bees (hind tibial perfume extracts) none perfume chemical composition / relative compound abundance (nMDS, ANOSIM, SIMPER) GC–MS
Morphometric analysis of mandible dentation E. dilemma and E. viridissima males (sympatric and allopatric) none number of mandibular teeth
Maximum likelihood phylogeny and dN/dS molecular evolution analysis Or41 sequences from 47 individuals plus five outgroup species none genotype species-specificity, dN/dS, selection signature
Functional odorant receptor assay (heterologous expression) Or41 derived (E. dilemma) and ancestral (E. viridissima) variants OR variant comparison odorant tuning / response specificity to perfume compounds
Key results
  • PCA of 16,369 SNPs separated E. dilemma and E. viridissima in allopatry and sympatry, with first two PCs explaining <4% of variation. <4% variance on PC1+PC2
  • Perfume composition differentiated into two distinct lineage-specific chemical phenotypes independent of geography. ANOSIM R=0.8, p=0.001
  • HNDB (E. dilemma) and L97 (E. viridissima) are diagnostic compounds accounting for the largest perfume proportions and a large share of chemical differentiation. HNDB 55%, L97 37%; together 46.3% of differentiation
  • Two species-specific selective sweeps identified, harboring tandem OR arrays totaling 43 OR genes. 39 ORs (Ed sweep) + 4 ORs (Ev sweep)
  • Or41 shows positive selection on the E. dilemma branch but purifying selection in E. viridissima. dN/dS=3.6 (E. dilemma) vs 0.3 (E. viridissima)
  • Nucleotide diversity (π) at Or41 was much lower in E. dilemma than E. viridissima, consistent with a sweep. 5-fold lower π in E. dilemma
  • Of 19 substitutions mapped on Or41 membrane topology, 17 were non-synonymous and 2 synonymous. 17/19 non-synonymous
  • In sympatry 26% of E. viridissima males had three mandibular teeth vs 3% in allopatry, consistent with introgression. 26% vs 3%; Fisher p=0.0009
Key statistics
  • correlation ANOSIM R=0.8 (perfume composition differentiation between lineages, p=0.001)
  • fold_change dN/dS=3.6 (positive selection on Or41 E. dilemma branch)
  • fold_change dN/dS=0.3 (purifying selection on Or41 E. viridissima branch)
  • pvalue p=0.0009 (Fisher's exact test, three-tooth frequency sympatry vs allopatry)
  • other f4=0.001, z=2.7, p=0.007 (f4-test rejecting simple bifurcating phylogeny)
  • correlation r=−0.13, p=0 (Dxy negatively correlated with ∆FST')
  • other pairwise FST: 0.04–0.18 (low differentiation among three genetic lineages)
  • count 43 OR genes (39 + 4) (ORs in the two selective sweep tandem arrays)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combines population-genomic, chemical-ecology, and molecular-evolution analyses across distribution-wide samples (232 genotyped males, 384 perfume samples, 30 re-sequenced genomes). Genetic structure was assessed with PCA, ADMIXTURE clustering, F_ST, an f4-test, and AIC-based demographic model selection; perfume chemistry was compared with nMDS plus ANOSIM and SIMPER; genome divergence used a net-F_ST (ΔF_ST') window scan with correlation analyses, a Mann–Whitney U comparison of D_xy, selective-sweep (CLR) detection, and dN/dS and maximum-likelihood phylogenetics on Or41. A categorical trait was tested with Fisher's exact test. Results were reported with test statistics and p-values, with dispersion shown as SEM and box-plot quartiles/IQR.

Replicationbiological Sample sizeSample sizes stated per analysis: 232 genotyped males, 384 GC–MS perfume samples, 30 re-sequenced genomes (n = 10 each for Ed-north, Ed-south, Ev), 47 individuals for Or41 phylogeny; no formal power analysis described GroupsE. dilemma (Ed-north, Ed-south) vs E. viridissima; outlier vs non-outlier genomic windows; sympatric vs allopatric populations Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Principal components analysis (PCA) of genetic variance Genetic differentiation of E. dilemma vs E. viridissima, Fig. 1b 232 males / 16,369 SNPs not stated
ADMIXTURE genetic clustering analysis Population/species structure, Fig. 1c 232 males not stated
Pairwise F_ST estimation Differentiation among Ev, Ed-south, Ed-north (0.04–0.18) na
f4-test (four-population test) Test of bifurcating phylogeny among lineages (z = 2.7, p = 0.007) not stated
Demographic model selection via Akaike Information Criterion (AIC) Species differentiation model (AIC weight = 1), Supplementary Table 5 30 re-sequenced genomes not stated
Fisher's exact test Three-tooth frequency in sympatric vs allopatric E. viridissima males (p = 0.0009) na
Non-metric multidimensional scaling (nMDS) with ANOSIM Perfume composition differentiation (R = 0.8, p = 0.001), Fig. 1d 384 individuals not stated
Similarity Percentage (SIMPER) analysis Contribution of HNDB and L97 to chemical differentiation (46.3%) na
Net interspecific F_ST (ΔF_ST') window scan Genome-wide divergence in 50 kb windows (>99th percentile outliers), Fig. 2 30 genomes (n = 10 per lineage) na
Pearson's correlation ΔF_ST' vs gene density (r = 0.17), π (r < −0.24), D_xy (r = −0.13), LD (r ≥ 0.1), all p = 0 not stated
Mann–Whitney U-test D_xy in outlier vs non-outlier regions (p = 0.001), Fig. 2b na
Composite likelihood ratio (CLR) selective-sweep test Sweep detection in outlier windows / Or41, Fig. 3a n = 10 per lineage na
Maximum likelihood phylogeny with bootstrap support Or41 genotype phylogeny, Fig. 3b 47 individuals na
dN/dS analysis Selection on Or41 branches (E. dilemma dN/dS = 3.6; E. viridissima = 0.3), Fig. 3b 5 outgroup species na
Approaches that could also have been used
  • Genome-wide divergence outliers were identified using a fixed 99th-percentile threshold on ΔF_ST' windows.
    Could also: A model-based or empirical-null approach (e.g., simulation under demography, or a formal false-discovery-rate control across windows) could ALSO be used to flag outliers. — An explicit FDR or null-model calibration would attach a quantified error rate to each candidate region, complementing the percentile cutoff.
  • Several p-values from correlation tests are reported as 'p = 0'.
    Could also: Reporting exact small p-values (e.g., p < 1e-16) or the test statistic with degrees of freedom would ALSO convey the result. — Precise small-value reporting communicates the strength of evidence and aids reproducibility and meta-analysis.
  • Dispersion for D_xy was shown as 1 SEM, alongside IQR-based box plots elsewhere.
    Could also: Standard deviation or a 95% confidence interval could ALSO be reported. — SD or a CI conveys the spread or estimation uncertainty directly and is often preferred for describing variability, providing consistent dispersion reporting across panels.
  • Perfume composition was compared using ANOSIM following nMDS ordination.
    Could also: PERMANOVA (e.g., adonis on the dissimilarity matrix) could ALSO be used to test group differences. — PERMANOVA partitions variance, accommodates covariates such as geography, and is less sensitive to within-group dispersion differences than ANOSIM.
  • A categorical tooth-count difference between population types was tested with Fisher's exact test.
    Could also: A logistic regression or generalized linear model could ALSO model the trait while incorporating site or population as a covariate. — A regression framework would allow adjustment for additional factors and yield an effect size (e.g., odds ratio) with a confidence interval.
  • Branch-wise selection on Or41 was summarized with a single dN/dS ratio per branch.
    Could also: A branch-site likelihood test with a formal null comparison (e.g., LRT) could ALSO be applied. — A branch-site test provides a statistical significance assessment for positive selection at specific codons, complementing the descriptive dN/dS values.
Software: ADMIXTURE (clustering)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
57
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJNA388474 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
PRJNA529235 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31932598

Paper: Brand et al. 2020, Nat Commun 11:244. "The evolution of sexual signaling is linked to odorant receptor tuning in perfume-collecting orchid bees." DOI 10.1038/s41467-019-14162-6.

Code: https://github.com/pbrec/popgen-popchem @ commit 2ca05c9c32ddc55b022ec44ae92cd4a01ad64d1d (pushed 2019-12-15). Authors' own code (R scripts). README is a one-line title only — no run instructions; scripts are self-documenting and read shipped data files.

Data: SRA PRJNA529235 (WGS + GBS reads). Crucially, the repo ships the derived data matrices used by most downstream analyses, so those analyses can be reproduced 1:1 without re-running the upstream read→alignment→SNP-call pipeline.

In scope (pipeline-derived, reproducible from SHIPPED data)

# Result Script Shipped input Pipeline / tool
C1 Demographic model selection by AIC: best moments model has AIC weight = 1 moments_AIC.R moments_results.txt (16 models, log-likelihoods) base R AIC + Akaike weights over moments (Jouganous et al.) fit log-likelihoods
C2 GBS SNP PCA: PC1+PC2 jointly explain <4% of genetic variation gbs_analysis.R GBS_SNPs_MAF0.05_SNPrelate_format.txt (5428 SNPs), pop_ids_gbs.txt SNPRelate snpgdsPCA
C3 Perfume chemistry ANOSIM E. dilemma vs E. viridissima: R = 0.8, p = 0.001 perfume_analysis_and_plotting.R (ANOSIM block) peak_area_matrix_Brandetal.txt (384 indiv × 653 compounds), pop_ids_gcms.txt vegan/ecodist Bray-Curtis + anosim (999 perms)
C4 Perfume filtering: 384 → 306 individuals after ≥10-peak filter same same data filtering step (deterministic)

Out of scope (NOT attempted, with reason)

Result Script Why out of scope
Genome-wide FST / π / Tajima's D sliding-window scan (Fig 2/3) genome_analysis.R Header states: "request VCF files (very large) by emailing Philipp Brand … or Santiago Ramirez." The VCFs are not in the repo and are on-request onlydata_restricted. Re-deriving them would require the full WGS read→ref-map→SNP-call pipeline plus the reference genome + GFF (not pointed to), a separate large effort outside the shipped-data reproduction.
Or41 functional electrophysiology (Fig 4) or41_functional_plotting.R Reads data.txt / hndb_dilutions.txt = "Source Data file for Figure 4" — not in repo; these are single-sensillum-recording (wet-lab) measurements, non-pipeline. Plotting only.

Notes

  • The shipped GBS matrix is the MAF≥0.05 SNPRelate-format set with 5428 SNPs; the paper text mentions 16,369 SNPs (likely a pre-filter count). The PCA script re-applies maf=0.05 + missing.rate=0.5 on this matrix, so the reproduced variance % is computed on the shipped set. We compare the PC1+PC2 magnitude claim.
  • Compute: conda R env built inside a «our HPC» SLURM job (compute nodes have internet); shipped data already on «infra».
Figures / tables: Fig.1Fig.1d
C1
Reported
AIC weight = 1
Reproduced
1.000000 (best model ESm222geno; delta-AIC to 2nd = 46.4)
exact
C2
Reported
PC1+PC2 jointly explained <4% of genetic variation
Reproduced
PC1+PC2 = 3.81% (2.62 + 1.19); N=232 indiv, 5428 SNPs
within tolerance
C3
Reported
perfume ANOSIM R = 0.8, p = 0.001
Reproduced
R = 0.7751 (rounds to 0.8), p = 0.001; N=270 (163 dilemma / 107 viridissima)
within tolerance
C4
Reported
306 individuals after <10-compound filter
Reproduced
306 (>=3-prevalence subset, >=10 compounds)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

All four in-scope, pipeline-derived claims reproduce 1:1 from the authors' own deposited matrices and R scripts: moments AIC weight=1 (exact), GBS PCA PC1+PC2=3.81% (<4% as stated), perfume ANOSIM R=0.7751→0.8 with p=0.001, and the 306-individual filter (exact). The only gaps are on our/technical side or harmless packaging — a 1-decimal rounding on R, a filename typo, a 16,369-vs-5428 SNP labelling nuance, and a looser-vs-stricter filter that gives the same ANOSIM result. No fabrication; the genome-wide diversity scan (on-request VCFs) and Or41 electrophysiology (wet-lab) were legitimately out of scope, which is a data-availability/non-pipeline matter, not an authors' defect on the reproduced claims.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

147.7 k
tokens (I/O) · 12.4 M incl. cache
28 min
runtime · 0.01 CPU-h
1.7 GB
peak RAM
2
HPC jobs
hummel
machine