Ordinal-level phylogenomics of the arthropod class Diplopoda (millipedes) based on an analysis of 221 nuclear protein-coding loci generated using next-generatio
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Core pipeline (SRA download -> FASTX QC/trim -> Trinity de novo assembly -> BUSCO completeness) reproduced within tolerance for all 9 taxa: processed read-pair volumes match the paper's Table 1 to ~2%, contig counts are systematically 7-16% higher (explainable by Trinity version/read-retention differences), and BUSCO-vs-CEGMA completeness ranking across taxa is concordant (Spearman rho=0.967). One dataset-accounting discrepancy (raw-read counts ~76-80% of paper's column despite near-exact processed-read match) is flagged as an unreconciled definitional ambiguity, not asserted either way. Downstream HaMStR orthology assignment and everything depending on it (locus alignment, supermatrix, ML tree) was NOT attempted: the 2013-era HaMStR toolchain and the paper's curated reference proteome are unavailable, and the maintained successor uses incompatible reference databases; a self-designed substitute was rejected as producing false precision rather than genuine reproduction evidence. This is a documented scope boundary, not a silent drop.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a phylogenomic dataset of hundreds of nuclear protein-coding loci derived from next-generation transcriptome sequencing resolve ordinal-level relationships within the millipede class Diplopoda, and do the resulting relationships agree with the existing 16-order classification and traditional morphology-based groupings?
- ★ An ordinal-level phylogeny of Diplopoda reconstructed from 221 nuclear protein-coding loci (61,641 aligned amino acid columns) differs from existing classifications in fundamental ways. finding
- ★ The phylogeny recovers a previously undescribed grouping of Juliformia + Merocheta + Stemmiulida. finding
- ★ Ancestral state reconstruction of spinnerets suggests caution in using spinnerets as a unifying character for the Nematophora. finding
- ★ Likelihood-based topology tests show the new phylogeny has significantly stronger support than previous millipede ordinal phylogeny hypotheses given these data. finding
- ★ Illumina RNA-Seq transcriptomes combined with HaMStR orthology assignment and rigorous dataset optimization constitute the first phylogenomic dataset produced for a myriapod taxon. resource
- A dataset optimization pipeline (SCaFoS locus selection, Gblocks, Aliscore/ALICUT masking, paralog and short-sequence removal) reduces 1,005 candidate orthologs to 221 high-quality loci with substantially less missing data. method
- Divergence times for major millipede lineages were estimated using relaxed molecular clocks calibrated with two fossil constraints. method
- Gonopods and other morphological character systems are of limited use for high-level millipede phylogeny because they may not be homologous across orders. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Transcriptome sequencing (Illumina RNA-Seq) | 9 exemplar millipede species plus a Lithobius centipede outgroup; anterior head region plus body rings of larger animals, whole bodies of smaller animals, preserved in RNAlater | none | Paired-end 50 bp raw reads (29.7–63.9 million per taxon) | Illumina HiSeq, paired-end 50 bp chemistry; libraries prepared and sequenced at Hudson Alpha; RNA extracted with Qiagen RNEasy kit and Shredder columns |
| Read quality control/trimming | Paired-end FASTQ read files from each of 9 sequenced taxa | none | Number of reads retained after quality-score-20 trimming, <30 bp removal, and first-9-base clipping | FASTX Toolkit; FASTQC; sync_paired_end_reads.py |
| De novo transcriptome assembly | Cleaned paired-end reads from 9 sequenced arthropod taxa | none | Number of assembled contigs representing unique mRNA transcripts | Trinity pipeline (min_kmer_cov 2, run_butterfly, CPU 6, bflyHeapSpace 10G) |
| Transcriptome completeness estimation | Assembled transcriptomes of each exemplar taxon | none | Number and percentage of 248 core eukaryotic genes (CEGs) recovered at complete and partial coverage | CEGMA, using KOG/COG-derived core eukaryotic gene set |
| Orthology assignment and alignment/dataset optimization | Trinity assemblies plus GenBank EST contigs for Ixodes scapularis and Archispirostreptus gigas; arthropod core ortholog set with Daphnia pulex as primary reference | none | Number of orthologous loci per taxon and final concatenated supermatrix dimensions (loci, characters, gaps, missing data) | HaMStR (e-value 1e-20, representative option), MAFFT, SCaFoS, Gblocks, EMBOSS infoalign, Aliscore 1.0, ALICUT 2.0, FASconCAT |
| Maximum likelihood phylogenetic inference | Concatenated amino acid supermatrix (221 loci) of 12 arthropod exemplars, partitioned by ortholog; also individual locus alignments and an outgroup-removed dataset | outgroup removal (long-branch attraction/outgroup sensitivity test) | ML tree topology and bootstrap support values | RAxML v7.2.8, PROTGAMMAWAG model, 1,000 random addition sequence replicates, 1,000 bootstrap replicates |
| Bayesian phylogenetic inference and likelihood-based topology testing | Same amino acid supermatrix; previously published tree topologies trimmed to the taxa included here | alternative topology constraints from prior published hypotheses | Posterior probabilities/consensus tree; per-topology likelihood scores compared by AU, NP, BP, PP, KH, SH, WKH, and WSH tests | Phylobayes v3.3b (5 chains, 10,000 cycles, 20% burn-in, bpcomp); FastTree 2.1; CONSEL |
| Ancestral state reconstruction and molecular divergence dating | Bayesian consensus phylogeny of millipede ordinal exemplars | fossil calibration constraints (Pneumodesmus newmani 428 MYA; Sigmastria dilata 410 MYA; Diplopoda max ~480 MYA; root 560 MYA) | Reconstructed ancestral states for gonopods (0/1/2), ozopores (0/1), and spinnerets (0/1); divergence dates with 95% confidence intervals | Mesquite (parsimony model); Phylobayes 3.3 (UGAM and lognormal relaxed clock, 10,000 cycles, 2,000 burn-in) |
- – The recovered phylogeny includes a clade never previously described: Juliformia + Merocheta + Stemmiulida.
- – The new topology is significantly better supported than previously published millipede ordinal phylogeny hypotheses given the data.
- – Ancestral reconstruction indicates spinnerets are not a reliable unifying characteristic for the Nematophora.
- ▼ Dataset optimization reduced the supermatrix from 1,005 loci/532,002 characters to 221 loci/61,641 characters. 1005 to 221 loci; 532,002 to 61,641 aligned AA columns
- ▼ Optimization decreased gaps and missing data proportions in the supermatrix. gaps 28.02% to 7.20%; missing data 28.73% to 10.16%
- – Quality screening retained the large majority of raw reads across taxa. average 74.18% (range 70.77%-76.39%)
- – CEGMA complete protein coverage varied widely among taxa, lowest in Prostemmiulus sp. and highest in Brachycybe lecontii. 33.47% to 81.85%, mean 57.48%
- – HaMStR ortholog recovery for newly sequenced taxa averaged 810.11 loci. mean 810.11 (range 594-924)
- count 221 (Nuclear protein-coding loci in the final optimized supermatrix (1,005 pre-optimization))
- count 61,641 (Aligned amino acid columns post-optimization (532,002 pre-optimization))
- other gaps 7.20% post-optimization vs 28.02% pre-optimization; missing data 10.16% vs 28.73% (Supermatrix gap and missing-data proportions before and after optimization)
- mean 74.18% (70.77%-76.39%) (Proportion of raw Illumina reads retained after FASTX Toolkit quality filtering)
- mean 17,833.44 (10,524-33,692) (Trinity-assembled contigs per novel transcriptome; skewed upward by Lithobius)
- mean complete coverage 57.48% (33.47%-81.85%); partial coverage 73.93% (56.85%-93.95%) (CEGMA transcriptome completeness against 248 core eukaryotic genes)
- mean 810.11 (594-924) (HaMStR orthologs recovered per newly sequenced taxon)
- count 12 exemplar arthropod specimens (9 millipedes and 3 outgroups) (Taxon sampling for the supermatrix; alignments required >=11 of 12 taxa)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper reconstructed an ordinal-level millipede phylogeny from 221 nuclear protein-coding loci using maximum likelihood (RAxML, 1000 bootstrap replicates) and Bayesian inference (Phylobayes, 5 chains x 10,000 cycles) on a concatenated supermatrix. Statistical comparison of the resulting topology against four previously published hypotheses was performed with likelihood-based topology tests (AU, NP, BP, PP, KH, SH, weighted KH, weighted SH) implemented in CONSEL. Ancestral states of three morphological characters were reconstructed with a parsimony model, and divergence times were estimated under relaxed-clock Bayesian dating with 95% confidence intervals from two fossil calibrations.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Likelihood-based tree topology tests (AU, NP, BP, Bayesian PP, KH, SH, weighted KH, weighted SH) via CONSEL | comparing the study's ML/BI phylogeny to four previously published millipede ordinal-level topology hypotheses | — | not stated |
| Maximum likelihood bootstrap support (RAxML, PROTGAMMAWAG model) | node support on the ML supermatrix and individual-locus phylogenies | 1,000 bootstrap replicates; 1,000 random addition sequence replicates | not stated |
| Bayesian inference posterior probabilities (Phylobayes) | node support on the BI supermatrix phylogeny (Figure 1) | 5 independent chains, 10,000 cycles each, first 20% discarded as burn-in | not stated |
| Parsimony-based ancestral state reconstruction (Mesquite) | gonopod type, ozopore presence, and spinneret presence mapped onto the BI phylogeny | — | not stated |
| Relaxed-clock Bayesian divergence dating (Phylobayes, UGAM and lognormal models) with 95% confidence intervals | estimated timing of major millipede lineage divergences | two fossil calibration points; 10,000 cycles per clock model, first 2,000 discarded as burn-in | not stated |
-
Competing tree topologies were evaluated using eight CONSEL statistics (AU, NP, BP, PP, KH, SH, weighted KH, weighted SH) without a stated correction for evaluating multiple tests on the same comparisons.↳ Could also: Designating one test (e.g., the AU test, which corrects for selection bias) as the primary criterion, or applying a formal multiple-comparison adjustment across tests, could also be used. — The different CONSEL statistics have distinct bias properties (e.g., SH tends to be conservative, AU corrects for selection bias), so treating one as primary or explicitly accounting for running several together is a standard way to summarize this kind of multi-test topology comparison.
-
Node support on the ML phylogeny was assessed with nonparametric bootstrapping (1,000 replicates).↳ Could also: Transfer bootstrap expectation (TBE) or approximate likelihood ratio tests (aLRT) could also be used to summarize branch support. — TBE and aLRT were developed to give more informative support values on large or sparsely sampled phylogenomic matrices and are often reported alongside standard bootstrap values.
-
Ancestral states for gonopods, ozopores, and spinnerets were reconstructed using a parsimony model.↳ Could also: A likelihood- or Bayesian-based ancestral state reconstruction (e.g., a Markov model of character evolution on the same BI tree) could also be used. — Model-based reconstruction incorporates branch-length information and yields probabilities for each ancestral state rather than a single most-parsimonious assignment, and is commonly reported alongside parsimony reconstructions for comparison.
-
Divergence times were estimated from a single Bayesian relaxed-clock analysis (Phylobayes, UGAM and lognormal models) calibrated with two fossil points.↳ Could also: Cross-checking estimates with an independent dating program (e.g., BEAST or MCMCtree) or additional fossil calibrations could also be used. — Comparing divergence-time estimates across independent software and calibration schemes is a common way to gauge how sensitive date estimates are to modeling choices, complementing the reported credibility intervals.
-
Transcriptome completeness was summarized as percentages of complete/partial gene recovery using CEGMA, with ranges given across taxa.↳ Could also: BUSCO-based completeness scoring could also be used as an alternative to CEGMA. — BUSCO uses more taxon-specific, updated ortholog sets and became a widely adopted successor to CEGMA for assessing transcriptome or genome completeness.
-
A single amino-acid substitution model (PROTGAMMAWAG) was applied uniformly across all partitions in the ML analysis.↳ Could also: Partition-specific model selection (e.g., testing several candidate AA substitution models per locus with a tool such as ProtTest) could also be used. — Selecting models separately per partition can capture variation in substitution patterns among loci and is a commonly used complementary strategy in phylogenomic analyses with many loci.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The upstream half of this study reproduces well: the same nine SRA accessions yield processed read-pair counts within ~2% of Table 1 (e.g. 23,074,205 vs 23,473,057 for Glomeridesmus), contig counts 7–16% higher in one consistent direction (assembler-version drift), and a BUSCO completeness ranking concordant with the paper's CEGMA ranking at rho=0.967. Two things keep this out of green. First, the paper's 'Raw Reads' column cannot be reconstructed from the deposited runs — ours is 76–80% of it despite the near-exact processed-read match — an unresolved counting-definition issue on the authors' reporting side. Second, and more importantly, the paper's actual result — the 221-locus supermatrix ML phylogeny — was not attempted at all because HaMStR and its curated reference proteome are unarchived and no compatible successor exists; the curators correctly refused a BLAST-RBH substitute as false precision. Nothing here looks fabricated or 'too perfect'; this is a well-documented partial reproduction whose limit is toolchain decay plus an undeposited reference set, not a discrepancy in the reported science.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.