Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genome and cuticular hydrocarbon-based species delimitation shed light on potential drivers of speciation in a Neotropical ant species complex.

Ecol Evol · 2022
L1 87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for a 1:1 on the metadata claims and a genuine partial on the de-novo pipeline. C1-C4 verified EXACTLY/within-tol directly from deposited data: 3RAD read range 75,193-1,007,069 (SRA) exact; 35-run composition (34 E. ruidum + 1 E. tuberculatum outgroup) exact; UCE 100% matrix length 508,859 chars exact (Figshare partition); UCE 100% loci 640 vs 642 within-tol. C5 is the REAL pipeline reproduction and was run on «our HPC» compute: ipyrad 0.9.108 + VSEARCH 2.31.0 de-novo assembly (clust_threshold 0.98) on the 35 deposited 3RAD fastqs, branched to the four reported min-taxa thresholds (25/28/30/33). Reproduced matrix loci 447-6722 vs reported 986-7094: the inclusive matrix (min25) 6722 reproduces the reported upper bound 7094 to within ~5% and the ranges overlap (same order of magnitude); the strict-occupancy matrices fall lower because SRR17568700 failed clustering (34 not 35 samples assembled -> min33 = 33/34 = 97% occupancy, relatively stricter than the paper's 33/35 = 94%), compounded by the ipyrad 0.6.19->0.9.108 / vsearch 2.0.3->2.31.0 version delta (the paper's exact 0.6.19/py2.7 stack is unsolvable on current bioconda) and unspecified params (restriction_overhang/filters; ipyrad defaults used). NOT attempted (out of scope / stochastic / wet-lab): downstream RAxML/ASTRAL/STRUCTURE/BFD*/SVDquartets phylogenies + species-delimitation conclusions, CHC chemistry, cox1 Sanger. DATA DEFECT flagged: the paper's UCE accession 'PRJNA79660' is malformed; nearest valid BioProjects (PRJNA796600/796601) resolve to Saccharomyces cerevisiae RNA-Seq, so the UCE raw reads are not locatable from the stated accession (downstream UCE results rely on the Figshare matrices instead).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 87
    assessed: 2026-06-19 ⛓ 8a95cfe382b2
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether species boundaries within the Neotropical Ectatomma ruidum ant species complex coincide with geographic barriers, using integrative genomic (3RAD, UCE), mitochondrial DNA, and cuticular hydrocarbon (CHC) evidence to delimit species and identify alternative (non-geographic) drivers of speciation such as mitonuclear incompatibility, CHC differentiation, and colony structure.

Core claims
  • At least five distinct species exist within the E. ruidum species complex finding
  • Species boundaries in the complex do not coincide with known geographic barriers finding
  • Three of the five species are restricted to southeast Mexico and show high levels of mitochondrial heteroplasmy finding
  • Integrative use of 3RAD, UCEs, cox1 mtDNA, and CHC profiling improves rigor of species delimitation in taxonomically complex groups method
  • Mitonuclear incompatibility, CHC differentiation, and colony structure are proposed as alternative (non-geographic) drivers of speciation in this complex mechanism
  • CHC profiles vary considerably among populations, supporting existence of distinct evolutionary lineages finding
  • A population from Huaxpaltepec, Oaxaca differs significantly in distress call, suggesting it is a distinct undescribed species finding
Experimental setups
Assay System Perturbation Readout Platform
cox1 mtDNA Sanger sequencing E. ruidum workers (250 specimens) + E. gibbum outgroup none cox1 sequence variation, haplotypes, phylogenetic relationships Sanger sequencing, edited/aligned in Geneious v10.1
3RAD genome-wide sequencing E. ruidum (34 specimens) + E. tuberculatum outgroup none genome-wide SNP/RAD loci for phylogenetic and species-tree analysis Illumina HiSeq2500 Rapid PE150
Ultraconserved element (UCE) sequencing E. ruidum (13 specimens, plus 4 archived males) + E. gibbum outgroup none UCE loci alignments, phased SNPs, phylogenetic relationships Illumina HiSeq 2500 (PE150) and HiSeq X Ten (PE150); ant-specific hym-v2 bait set
Cuticular hydrocarbon (CHC) profiling E. ruidum workers (24 new workers from Puerto Morelos and Cuatode, pooled with prior data to 132 workers total) none CHC profile composition/differentiation among populations
Maximum likelihood/Bayesian phylogenetic analysis cox1, 3RAD, and UCE data matrices none tree topology, branch support (bootstrap/posterior) RAxML v8, ExaBayes v1.5, SVDquartets/PAUP v4.0a
Key results
  • Analyses consistently identified at least five distinct species in the E. ruidum complex
  • Two species are widely distributed across the Neotropics; three are restricted to southeast Mexico
  • Species boundaries did not coincide with geographic barriers
  • Southeast Mexico-restricted species (spp. 3, 4, 2×3) show apparently high levels of heteroplasmy
  • UCE bait set enriched 2524 targeted UCE loci plus 452 exon baits using 9446 custom probes
  • cox1 fragment of 626 bp examined across 250 E. ruidum specimens plus 1 E. gibbum outgroup
Key statistics
  • count 250 specimens (cox1) + 1 outgroup (cox1 mtDNA data set sample size)
  • count 626 bp (length of cox1 gene fragment analyzed)
  • count 35 specimens (34 E. ruidum + 1 outgroup) (3RAD data set sample size)
  • count 14 specimens (13 E. ruidum + 1 outgroup) (UCE data set sample size)
  • count 9446 custom probes targeting 2524 UCE loci and 452 baits targeting 16 exons (hym-v2 bait set composition for UCE enrichment)
  • count 24 new CHC workers pooled into total of 132 workers (CHC data set sample size)
  • other 5 independent MCMC chains of 1,000,000 generations each, 10% burn-in (ExaBayes Bayesian phylogenetic analysis parameters)
  • other 500 bootstrap replicates (SVDquartets nonparametric bootstrapping for 3RAD species tree)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper employed an integrative taxonomic approach combining four data types—3RAD RADseq, ultraconserved elements (UCEs), mitochondrial cox1 Sanger sequences, and cuticular hydrocarbon (CHC) profiles—to delimit species within the Ectatomma ruidum ant complex. Phylogenetic inference was performed using maximum likelihood (RAxML, GTR+Γ) and Bayesian MCMC methods (ExaBayes, GTR+G), supplemented by a coalescent-based species tree approach (SVDquartets). Node support was assessed via nonparametric bootstrapping, with ≥70% bootstrap support defined as well-supported; CHC and formal species-delimitation statistical details are not captured in the provided text excerpt, which ends mid-Methods.

Replicationbiological Sample sizecox1: 251 specimens; 3RAD: 35 specimens (34 E. ruidum + 1 outgroup); UCE: 14 specimens (13 E. ruidum + 1 outgroup); CHC: 132 workers, 2–5 per colony from multiple nests GroupsFive putative species/lineages within the E. ruidum complex, plus outgroup taxa Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Maximum likelihood phylogenetic inference (RAxML v8, GTR+Γ, automatic bootstrap stopping rule or 200 replicates depending on dataset) cox1 dataset; 3RAD (4 matrices, min_sample_locus 25/28/30/33); UCE (3 matrices, 90/95/100% occupancy) cox1: 251 sequences; 3RAD: 35 specimens; UCE: 14 specimens not stated
Bayesian MCMC phylogenetic inference (ExaBayes v1.5, GTR+G, 5 independent chains, 1,000,000 generations each, 10% burn-in, convergence criterion: max discrepancy among chains <0.1) 3RAD datasets (4 matrices) 35 specimens not stated
Coalescent-based species tree estimation (SVDquartets v1.0 in PAUP v4.0a, nonparametric bootstrap, 500 replicates) 3RAD datasets 35 specimens not stated
Haplotype network construction (TCS algorithm implemented in POPART) cox1 haplotype relationships 251 sequences na
Clustering threshold sensitivity series (0.80–0.98) evaluated by locus-count inflection point (Ilut et al. 2014 approach) and ML nodal support 3RAD ipyrad assembly parameter optimisation 35 specimens not stated
Approaches that could also have been used
  • Phylogenetic node support was summarised as a bootstrap percentage with a fixed threshold of ≥70% defining 'well-supported' across all analyses
    Could also: Bayesian posterior probabilities (e.g., PP ≥0.95) could also serve as a complementary support metric reported alongside bootstrap values — Bootstrap values and posterior probabilities measure different aspects of nodal support; presenting both simultaneously allows readers to assess concordance between frequentist and Bayesian frameworks and to detect nodes where the two measures diverge
  • The GTR+Γ substitution model was applied uniformly across all partitions and data types without explicit model-selection testing described in the provided text
    Could also: Explicit model selection via information criteria (AIC, BIC, or Bayes factors) using tools such as ModelTest-NG or PartitionFinder could also identify the best-fitting substitution model per data partition — Formal model selection can reveal whether simpler models fit equally well, potentially reducing over-parameterisation and improving branch-length estimation, particularly for datasets with few taxa such as the UCE matrix (n = 14)
  • MCMC convergence in ExaBayes was assessed primarily by the maximum discrepancy statistic among chains (<0.1) inspected in TRACER
    Could also: Effective sample size (ESS ≥200 per parameter) and potential scale reduction factors (PSRF ≈1.0) could also be reported as complementary convergence diagnostics — ESS quantifies mixing efficiency at the parameter level and PSRF detects between-chain disagreement; together they provide finer-grained evidence of stationarity beyond the single discrepancy summary
  • Concordance across multiple inference methods (ML, Bayesian, SVDquartets) applied to the same 3RAD data was assessed qualitatively
    Could also: Site concordance factors (sCF, e.g., implemented in IQ-TREE 2) could also quantify the proportion of sites independently supporting each node across the genome — Concordance factors capture site- or gene-level discordance that bootstrap support alone does not reveal, providing additional resolution for nodes affected by incomplete lineage sorting or reticulate evolution — both relevant here given the heteroplasmy context
  • Species delimitation conclusions were drawn by integrating qualitative concordance across four independent data types (3RAD, UCE, cox1, CHC)
    Could also: Probabilistic species delimitation methods such as BPP (Bayesian multispecies coalescent), STACEY, or mPTP could also formally test competing species-boundary hypotheses — Model-based delimitation methods assign posterior probabilities or likelihood ratios to alternative delimitation schemes, complementing the convergence-of-evidence approach with quantified statistical uncertainty around species boundaries
  • Matrix completeness was explored via a sensitivity series of taxon-occupancy thresholds (75–100% for UCEs; min_sample_locus 25–33 for 3RAD) and results compared qualitatively
    Could also: Explicit missing-data simulations or SVDquartets weighting schemes that account for unequal locus sampling could also be used to quantify the directional effect of matrix incompleteness on topology and support values — Threshold sensitivity analyses as performed here are standard practice; adding simulations would further characterise whether missing data systematically bias inferred relationships, which can be relevant in reduced-representation datasets with uneven per-individual coverage
Software: RAxML 8.0 · ExaBayes 1.5 · SVDquartets / PAUP SVDquartets 1.0 / PAUP 4.0a · TRACER 1.7.0 · ipyrad 0.6.19 · Stacks 2.0 · PHYLUCE 1.5.0 · Geneious 10.1 · MAFFT 7.130b · Gblocks 0.91b · VSEARCH 2.0.3 · MUSCLE 3.8.31 · ABySS 1.3.6 · POPART

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 35342602 (Ectatomma ruidum species complex; 3RAD + UCE + CHC + cox1)

Title: Genome and cuticular hydrocarbon-based species delimitation shed light on potential drivers of speciation in a Neotropical ant species complex. Ecol Evol 2022;12:e8704. DOI 10.1002/ece3.8704 · PMCID PMC8928884. Listed code: https://github.com/torognes/vsearch (VSEARCH 2.0.3 — the clustering engine used inside ipyrad). Data: SRA PRJNA796376 (3RAD) + "PRJNA79660" (UCE, malformed accn). Analyzed matrices + topologies: Figshare collection 10.6084/m9.figshare.c.5821973.v1.

Pipelines used by the paper

  • 3RAD: Stacks process_radtags (demux/clean/trim, drop -c uncalled / -q low-qual) → ipyrad v0.6.19 de-novo assembly (clusters reads with VSEARCH 2.0.3, aligns with MUSCLE 3.8.31; clustering-threshold series à la Ilut 2014, threshold 0.98 retained) → concatenated NEXUS matrices at min-taxa 25/28/30/33 of 35.
  • UCE: ILLUMIPROCESSOR (trimmomatic) → ABySS 1.3.6 de-novo → PHYLUCE map to hym-v2 baits → MAFFT align → Gblocks trim → matrices at 75/80/90/95/100% taxon occupancy; SNP phasing (PHYLUCE Tutorial II).
  • Downstream (phylogenetics / delimitation): RAxML 8, ASTRAL 4.10.8, SplitsTree, STRUCTURE, BFD*/SNAPP, SVDquartets — multispecies-coalescent + Bayesian.
  • CHC (cuticular hydrocarbons) + cox1 mtDNA: wet-lab / Sanger.

IN SCOPE (pipeline-derived, deterministic-ish)

  • C1 — 3RAD per-sample read counts: range 75,193–1,007,069 (Results §3.1). Directly checkable against SRA read_count of the 35 deposited runs. → reproduced EXACT.
  • C2 — 3RAD dataset N = 35 (34 ingroup E. ruidum + 1 outgroup E. tuberculatum). → reproduced EXACT from SRA scientific_name.
  • C3 — UCE 100%-occupancy matrix length = 508,859 characters (Results §3.1, lower bound of "508,859 to 1,817,455 characters"). → EXACT vs deposited Figshare partition.
  • C4 — UCE 100%-occupancy matrix = 642 loci (Results §3.1). → 640 in deposited partition (within-tol, off by 2).
  • C5 — 3RAD matrices contained 986–7094 loci (Table S2). Requires re-running ipyrad/VSEARCH on the 35 deposited fastqs on «our HPC». ATTEMPTED — see status.

OUT OF SCOPE (not pipeline / not reproducible here)

  • UCE loci range 642–2196 upper bound (2196 = 75%-occupancy matrix; NOT deposited on Figshare — only 90/95/100% matrices are there). Partial verification only.
  • 3RAD outgroup shared-loci 375–1322 (Table S2; needs ipyrad rerun per-pair).
  • All downstream phylogenies / STRUCTURE K / BFD* marginal likelihoods / ASTRAL trees — stochastic MCMC/ML, integrative species-delimitation conclusions; out of scope.
  • CHC chemistry, cox1 Sanger, morphology — wet-lab, out of scope.

Datasets the paper relies on

  • PRJNA796376 — 3RAD, 35 Illumina single-end runs, Ectatomma (open). Profiled.
  • "PRJNA79660" (UCE) — MALFORMED accession. Nearest valid PRJNA796600/PRJNA796601 both resolve to Saccharomyces cerevisiae RNA-Seq (18 runs) — NOT ant UCE. UCE raw reads effectively unlocatable from the stated accession. Profiled as delivers_promised=no.
  • Figshare c.5821973 — analyzed NEXUS matrices + partitions + topologies (open). Profiled.
Figures / tables: TableFigshare
C1
Reported
3RAD per-sample reads 75,193-1,007,069
Reproduced
min=75193 (SRR17568680), max=1007069 (SRR17568701)
exact
C2
Reported
35 individuals (34 E. ruidum + 1 outgroup E. tuberculatum)
Reproduced
35 SRA runs: 34 E. ruidum + 1 E. tuberculatum
exact
C3
Reported
UCE 100%-occupancy matrix = 508,859 characters
Reproduced
508859 (max coord, deposited Figshare partition)
exact
C4
Reported
UCE 100%-occupancy matrix = 642 loci
Reproduced
640 charsets in deposited partition
within tolerance
C5
Reported
3RAD matrices 986-7094 loci (Table S2; min-taxa 25/28/30/33)
Reproduced
447-6722 loci (min25=6722, min28=4820, min30=3010, min33=447); ipyrad 0.9.108 + vsearch 2.31.0 de-novo clust 0.98 on 35 deposited fastqs, 34 reached across-DB
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

93.3 k
tokens (I/O) · 4.2 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.