Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Ordinal-level phylogenomics of the arthropod class Diplopoda (millipedes) based on an analysis of 221 nuclear protein-coding loci generated using next-generatio

PLoS One · 2013
L1 82/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
82/100
Reproducibility score
0.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 60% of all assessed papers rank 459 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Core pipeline (SRA download -> FASTX QC/trim -> Trinity de novo assembly -> BUSCO completeness) reproduced within tolerance for all 9 taxa: processed read-pair volumes match the paper's Table 1 to ~2%, contig counts are systematically 7-16% higher (explainable by Trinity version/read-retention differences), and BUSCO-vs-CEGMA completeness ranking across taxa is concordant (Spearman rho=0.967). One dataset-accounting discrepancy (raw-read counts ~76-80% of paper's column despite near-exact processed-read match) is flagged as an unreconciled definitional ambiguity, not asserted either way. Downstream HaMStR orthology assignment and everything depending on it (locus alignment, supermatrix, ML tree) was NOT attempted: the 2013-era HaMStR toolchain and the paper's curated reference proteome are unavailable, and the maintained successor uses incompatible reference databases; a self-designed substitute was rejected as producing false precision rather than genuine reproduction evidence. This is a documented scope boundary, not a silent drop.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a phylogenomic dataset of hundreds of nuclear protein-coding loci derived from next-generation transcriptome sequencing resolve ordinal-level relationships within the millipede class Diplopoda, and do the resulting relationships agree with the existing 16-order classification and traditional morphology-based groupings?

Core claims
  • An ordinal-level phylogeny of Diplopoda reconstructed from 221 nuclear protein-coding loci (61,641 aligned amino acid columns) differs from existing classifications in fundamental ways. finding
  • The phylogeny recovers a previously undescribed grouping of Juliformia + Merocheta + Stemmiulida. finding
  • Ancestral state reconstruction of spinnerets suggests caution in using spinnerets as a unifying character for the Nematophora. finding
  • Likelihood-based topology tests show the new phylogeny has significantly stronger support than previous millipede ordinal phylogeny hypotheses given these data. finding
  • Illumina RNA-Seq transcriptomes combined with HaMStR orthology assignment and rigorous dataset optimization constitute the first phylogenomic dataset produced for a myriapod taxon. resource
  • A dataset optimization pipeline (SCaFoS locus selection, Gblocks, Aliscore/ALICUT masking, paralog and short-sequence removal) reduces 1,005 candidate orthologs to 221 high-quality loci with substantially less missing data. method
  • Divergence times for major millipede lineages were estimated using relaxed molecular clocks calibrated with two fossil constraints. method
  • Gonopods and other morphological character systems are of limited use for high-level millipede phylogeny because they may not be homologous across orders. mechanism
Experimental setups
Assay System Perturbation Readout Platform
Transcriptome sequencing (Illumina RNA-Seq) 9 exemplar millipede species plus a Lithobius centipede outgroup; anterior head region plus body rings of larger animals, whole bodies of smaller animals, preserved in RNAlater none Paired-end 50 bp raw reads (29.7–63.9 million per taxon) Illumina HiSeq, paired-end 50 bp chemistry; libraries prepared and sequenced at Hudson Alpha; RNA extracted with Qiagen RNEasy kit and Shredder columns
Read quality control/trimming Paired-end FASTQ read files from each of 9 sequenced taxa none Number of reads retained after quality-score-20 trimming, <30 bp removal, and first-9-base clipping FASTX Toolkit; FASTQC; sync_paired_end_reads.py
De novo transcriptome assembly Cleaned paired-end reads from 9 sequenced arthropod taxa none Number of assembled contigs representing unique mRNA transcripts Trinity pipeline (min_kmer_cov 2, run_butterfly, CPU 6, bflyHeapSpace 10G)
Transcriptome completeness estimation Assembled transcriptomes of each exemplar taxon none Number and percentage of 248 core eukaryotic genes (CEGs) recovered at complete and partial coverage CEGMA, using KOG/COG-derived core eukaryotic gene set
Orthology assignment and alignment/dataset optimization Trinity assemblies plus GenBank EST contigs for Ixodes scapularis and Archispirostreptus gigas; arthropod core ortholog set with Daphnia pulex as primary reference none Number of orthologous loci per taxon and final concatenated supermatrix dimensions (loci, characters, gaps, missing data) HaMStR (e-value 1e-20, representative option), MAFFT, SCaFoS, Gblocks, EMBOSS infoalign, Aliscore 1.0, ALICUT 2.0, FASconCAT
Maximum likelihood phylogenetic inference Concatenated amino acid supermatrix (221 loci) of 12 arthropod exemplars, partitioned by ortholog; also individual locus alignments and an outgroup-removed dataset outgroup removal (long-branch attraction/outgroup sensitivity test) ML tree topology and bootstrap support values RAxML v7.2.8, PROTGAMMAWAG model, 1,000 random addition sequence replicates, 1,000 bootstrap replicates
Bayesian phylogenetic inference and likelihood-based topology testing Same amino acid supermatrix; previously published tree topologies trimmed to the taxa included here alternative topology constraints from prior published hypotheses Posterior probabilities/consensus tree; per-topology likelihood scores compared by AU, NP, BP, PP, KH, SH, WKH, and WSH tests Phylobayes v3.3b (5 chains, 10,000 cycles, 20% burn-in, bpcomp); FastTree 2.1; CONSEL
Ancestral state reconstruction and molecular divergence dating Bayesian consensus phylogeny of millipede ordinal exemplars fossil calibration constraints (Pneumodesmus newmani 428 MYA; Sigmastria dilata 410 MYA; Diplopoda max ~480 MYA; root 560 MYA) Reconstructed ancestral states for gonopods (0/1/2), ozopores (0/1), and spinnerets (0/1); divergence dates with 95% confidence intervals Mesquite (parsimony model); Phylobayes 3.3 (UGAM and lognormal relaxed clock, 10,000 cycles, 2,000 burn-in)
Key results
  • The recovered phylogeny includes a clade never previously described: Juliformia + Merocheta + Stemmiulida.
  • The new topology is significantly better supported than previously published millipede ordinal phylogeny hypotheses given the data.
  • Ancestral reconstruction indicates spinnerets are not a reliable unifying characteristic for the Nematophora.
  • Dataset optimization reduced the supermatrix from 1,005 loci/532,002 characters to 221 loci/61,641 characters. 1005 to 221 loci; 532,002 to 61,641 aligned AA columns
  • Optimization decreased gaps and missing data proportions in the supermatrix. gaps 28.02% to 7.20%; missing data 28.73% to 10.16%
  • Quality screening retained the large majority of raw reads across taxa. average 74.18% (range 70.77%-76.39%)
  • CEGMA complete protein coverage varied widely among taxa, lowest in Prostemmiulus sp. and highest in Brachycybe lecontii. 33.47% to 81.85%, mean 57.48%
  • HaMStR ortholog recovery for newly sequenced taxa averaged 810.11 loci. mean 810.11 (range 594-924)
Key statistics
  • count 221 (Nuclear protein-coding loci in the final optimized supermatrix (1,005 pre-optimization))
  • count 61,641 (Aligned amino acid columns post-optimization (532,002 pre-optimization))
  • other gaps 7.20% post-optimization vs 28.02% pre-optimization; missing data 10.16% vs 28.73% (Supermatrix gap and missing-data proportions before and after optimization)
  • mean 74.18% (70.77%-76.39%) (Proportion of raw Illumina reads retained after FASTX Toolkit quality filtering)
  • mean 17,833.44 (10,524-33,692) (Trinity-assembled contigs per novel transcriptome; skewed upward by Lithobius)
  • mean complete coverage 57.48% (33.47%-81.85%); partial coverage 73.93% (56.85%-93.95%) (CEGMA transcriptome completeness against 248 core eukaryotic genes)
  • mean 810.11 (594-924) (HaMStR orthologs recovered per newly sequenced taxon)
  • count 12 exemplar arthropod specimens (9 millipedes and 3 outgroups) (Taxon sampling for the supermatrix; alignments required >=11 of 12 taxa)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper reconstructed an ordinal-level millipede phylogeny from 221 nuclear protein-coding loci using maximum likelihood (RAxML, 1000 bootstrap replicates) and Bayesian inference (Phylobayes, 5 chains x 10,000 cycles) on a concatenated supermatrix. Statistical comparison of the resulting topology against four previously published hypotheses was performed with likelihood-based topology tests (AU, NP, BP, PP, KH, SH, weighted KH, weighted SH) implemented in CONSEL. Ancestral states of three morphological characters were reconstructed with a parsimony model, and divergence times were estimated under relaxed-clock Bayesian dating with 95% confidence intervals from two fossil calibrations.

Replicationunclear Sample size12 exemplar taxa (9 millipedes plus 3 outgroups), reduced through orthology/alignment filtering from 1,005 to 221 loci (61,641 aligned amino acid columns); no formal power analysis is described Groupsthe study's resulting phylogeny versus four previously published millipede ordinal-level topology hypotheses Pairingna Randomization/blindingnot stated DispersionCI Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Likelihood-based tree topology tests (AU, NP, BP, Bayesian PP, KH, SH, weighted KH, weighted SH) via CONSEL comparing the study's ML/BI phylogeny to four previously published millipede ordinal-level topology hypotheses not stated
Maximum likelihood bootstrap support (RAxML, PROTGAMMAWAG model) node support on the ML supermatrix and individual-locus phylogenies 1,000 bootstrap replicates; 1,000 random addition sequence replicates not stated
Bayesian inference posterior probabilities (Phylobayes) node support on the BI supermatrix phylogeny (Figure 1) 5 independent chains, 10,000 cycles each, first 20% discarded as burn-in not stated
Parsimony-based ancestral state reconstruction (Mesquite) gonopod type, ozopore presence, and spinneret presence mapped onto the BI phylogeny not stated
Relaxed-clock Bayesian divergence dating (Phylobayes, UGAM and lognormal models) with 95% confidence intervals estimated timing of major millipede lineage divergences two fossil calibration points; 10,000 cycles per clock model, first 2,000 discarded as burn-in not stated
Approaches that could also have been used
  • Competing tree topologies were evaluated using eight CONSEL statistics (AU, NP, BP, PP, KH, SH, weighted KH, weighted SH) without a stated correction for evaluating multiple tests on the same comparisons.
    Could also: Designating one test (e.g., the AU test, which corrects for selection bias) as the primary criterion, or applying a formal multiple-comparison adjustment across tests, could also be used. — The different CONSEL statistics have distinct bias properties (e.g., SH tends to be conservative, AU corrects for selection bias), so treating one as primary or explicitly accounting for running several together is a standard way to summarize this kind of multi-test topology comparison.
  • Node support on the ML phylogeny was assessed with nonparametric bootstrapping (1,000 replicates).
    Could also: Transfer bootstrap expectation (TBE) or approximate likelihood ratio tests (aLRT) could also be used to summarize branch support. — TBE and aLRT were developed to give more informative support values on large or sparsely sampled phylogenomic matrices and are often reported alongside standard bootstrap values.
  • Ancestral states for gonopods, ozopores, and spinnerets were reconstructed using a parsimony model.
    Could also: A likelihood- or Bayesian-based ancestral state reconstruction (e.g., a Markov model of character evolution on the same BI tree) could also be used. — Model-based reconstruction incorporates branch-length information and yields probabilities for each ancestral state rather than a single most-parsimonious assignment, and is commonly reported alongside parsimony reconstructions for comparison.
  • Divergence times were estimated from a single Bayesian relaxed-clock analysis (Phylobayes, UGAM and lognormal models) calibrated with two fossil points.
    Could also: Cross-checking estimates with an independent dating program (e.g., BEAST or MCMCtree) or additional fossil calibrations could also be used. — Comparing divergence-time estimates across independent software and calibration schemes is a common way to gauge how sensitive date estimates are to modeling choices, complementing the reported credibility intervals.
  • Transcriptome completeness was summarized as percentages of complete/partial gene recovery using CEGMA, with ranges given across taxa.
    Could also: BUSCO-based completeness scoring could also be used as an alternative to CEGMA. — BUSCO uses more taxon-specific, updated ortholog sets and became a widely adopted successor to CEGMA for assessing transcriptome or genome completeness.
  • A single amino-acid substitution model (PROTGAMMAWAG) was applied uniformly across all partitions in the ML analysis.
    Could also: Partition-specific model selection (e.g., testing several candidate AA substitution models per locus with a tool such as ProtTest) could also be used. — Selecting models separately per partition can capture variation in substitution patterns among loci and is a commonly used complementary strategy in phylogenomic analyses with many loci.
Software: RAxML 7.2.8 · Phylobayes 3.3b · FastTree 2.1 · CONSEL · Mesquite · CEGMA

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

processed_reads_glomeridesmus
Reported
23,473,057 processed read pairs (paper Table 1)
Reproduced
23,074,205 processed read pairs (ratio 0.983)
within tolerance
processed_reads_brachycybe
Reported
26,009,589 processed read pairs (paper Table 1)
Reproduced
25,543,622 processed read pairs (ratio 0.982)
within tolerance
processed_reads_petaserpes
Reported
32,494,393 processed read pairs (paper Table 1)
Reproduced
31,854,025 processed read pairs (ratio 0.980)
within tolerance
processed_reads_pseudopolydesmus
Reported
30,638,665 processed read pairs (paper Table 1)
Reproduced
30,102,525 processed read pairs (ratio 0.983)
within tolerance
processed_reads_cleidogona
Reported
22,288,076 processed read pairs (paper Table 1)
Reproduced
21,751,427 processed read pairs (ratio 0.976)
within tolerance
processed_reads_abacion
Reported
30,238,012 processed read pairs (paper Table 1)
Reproduced
29,720,224 processed read pairs (ratio 0.983)
within tolerance
processed_reads_prostemmiulus
Reported
29,259,335 processed read pairs (paper Table 1)
Reproduced
28,590,967 processed read pairs (ratio 0.977)
within tolerance
processed_reads_cambala
Reported
27,334,077 processed read pairs (paper Table 1)
Reproduced
26,671,011 processed read pairs (ratio 0.976)
within tolerance
processed_reads_lithobius
Reported
46,366,497 processed read pairs (paper Table 1)
Reproduced
45,562,933 processed read pairs (ratio 0.983)
within tolerance
raw_reads_all_taxa
Reported
paper 'Raw Reads' column, Table 1, all 9 taxa
Reproduced
our RAW_PAIRS is consistently ~76-80% of paper's raw-read column despite near-exact processed-read match; most plausibly a column-definition ambiguity (mates vs pairs vs spots), not a pipeline defect
partial
contig_count_glomeridesmus
Reported
20,181 contigs (paper Table 1)
Reproduced
21,647 contigs (ratio 1.073)
within tolerance
contig_count_brachycybe
Reported
18,570 contigs (paper Table 1)
Reproduced
20,239 contigs (ratio 1.090)
within tolerance
contig_count_petaserpes
Reported
12,154 contigs (paper Table 1)
Reproduced
13,375 contigs (ratio 1.100)
within tolerance
contig_count_pseudopolydesmus
Reported
18,876 contigs (paper Table 1)
Reproduced
21,069 contigs (ratio 1.116)
within tolerance
contig_count_cleidogona
Reported
16,572 contigs (paper Table 1)
Reproduced
17,983 contigs (ratio 1.085)
within tolerance
contig_count_abacion
Reported
15,861 contigs (paper Table 1)
Reproduced
17,200 contigs (ratio 1.084)
within tolerance
contig_count_prostemmiulus
Reported
10,524 contigs (paper Table 1)
Reproduced
11,559 contigs (ratio 1.098)
within tolerance
contig_count_cambala
Reported
14,071 contigs (paper Table 1)
Reproduced
15,684 contigs (ratio 1.115)
within tolerance
contig_count_lithobius
Reported
33,692 contigs (paper Table 1)
Reproduced
38,974 contigs (ratio 1.157)
within tolerance
completeness_rank_order
Reported
CEGMA Complete% ranking across 9 taxa (paper Table 1): Brachycybe > Glomeridesmus > Pseudopolydesmus > Lithobius > Cambala > Cleidogona > Abacion > Petaserpes > Prostemmiulus
Reproduced
BUSCO Complete% (arthropoda_odb10) ranking: Brachycybe > Pseudopolydesmus > Glomeridesmus > Lithobius > Cambala > Abacion > Cleidogona > Petaserpes > Prostemmiulus (Spearman rho=0.967 vs paper ranking); tool substituted because CEGMA is deprecated/non-functional on current systems
within tolerance
hamstr_orthology_221_loci
Reported
HaMStR assignment of 221 nuclear protein-coding orthologous loci, used as input to the paper's supermatrix and ML phylogeny
Reproduced
not attempted: HaMStR (2013, Genewise-dependent) and its curated reference proteome are unavailable/unarchived; the maintained successor fDOG uses incompatible OrthoDB references; no orthology-assignment tool exists in any available conda env; a self-designed BLAST-RBH substitute was rejected as producing a non-comparable locus set (false precision)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 82/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

The upstream half of this study reproduces well: the same nine SRA accessions yield processed read-pair counts within ~2% of Table 1 (e.g. 23,074,205 vs 23,473,057 for Glomeridesmus), contig counts 7–16% higher in one consistent direction (assembler-version drift), and a BUSCO completeness ranking concordant with the paper's CEGMA ranking at rho=0.967. Two things keep this out of green. First, the paper's 'Raw Reads' column cannot be reconstructed from the deposited runs — ours is 76–80% of it despite the near-exact processed-read match — an unresolved counting-definition issue on the authors' reporting side. Second, and more importantly, the paper's actual result — the 221-locus supermatrix ML phylogeny — was not attempted at all because HaMStR and its curated reference proteome are unarchived and no compatible successor exists; the curators correctly refused a BLAST-RBH substitute as false precision. Nothing here looks fabricated or 'too perfect'; this is a well-documented partial reproduction whose limit is toolchain decay plus an undeposited reference set, not a discrepancy in the reported science.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.