Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Social complexity, life-history and lineage influence the molecular basis of castes in vespid wasps.

Nat Commun · 2023
L1 52/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
52/100
Reproducibility score
1.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 13% of all assessed papers rank 1014 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the authors' OWN containerized Nextflow+Perl+R pipeline (chriswyatt1/Trinity_to_SVM v1.0.0, container chriswyatt/perl_r_e1071) on openly-deposited GSE159973. CLEAN: deposit delivers exactly what is described - 9 Trinity assemblies (R1) + 18 caste-level RSEM tables (R2), internally consistent (RSEM rows==assembly transcripts), grade-A dataset. The SVM/edgeR analysis pipeline RUNS end-to-end and reproduces the METHOD (lm caste~expr p<0.05 feature selection -> SVM leave-one-out classifier + heatmaps/PCA/significance tables). The paper's HEADLINE gene counts (259/277/289 SVM; 57/353 edgeR) are NOT 1:1 reproducible from the shipped defaults because the repo pins only ONE example species-configuration (test=A.pallens vs 4 background) -> 1283/2707 p<0.05 orthogroups, not 259; the published per-society toolkits require the specific groupings in Sumner-lab/Multispecies_paper_ML not encoded here. R6 Fig4a orthogroup counts also mismatch: shipped noMac 9-species matrices give 2516 single-copy, not 1221, because Fig4a uses the outgroup-inclusive OrthoFinder run that is not deposited. All discrepancies are explainable (config / different deposited artifact), no fabrication signal. Not attempted: wet-lab, dN/dS PAML selection test, GO enrichment, phylogeny narrative.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 9c25ff77295a
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether a shared genetic toolkit of genes, co-opted from a common solitary ancestral ground plan, regulates caste differentiation across vespid wasp species representing different levels of social complexity (non-superorganismal to superorganismal).

Core claims
  • A shared genetic toolkit of caste-associated genes exists across vespid wasp species spanning different levels of social complexity finding
  • Fine-scale differences in predictive gene sets, functional enrichment, and rates of gene evolution are associated with level of social complexity, lineage (vespines vs polistines), and mode of colony founding finding
  • The concept of a universal shared genetic toolkit for sociality may be too simplistic to fully describe the major transition to sociality finding
  • Support vector machine models trained on brain transcriptomes can classify caste identity across species using selected gene sets method
  • 57 orthologous genes show caste-biased expression in four or more of nine species finding
  • Orthogroup OG0001418, predicted to be a vitellogenin family member linked to juvenile hormone signalling, is upregulated in queens in eight of nine species finding
  • Oocyte zinc finger 22 is consistently queen-biased across species finding
  • Brachygastra and Agelaia are outlier species that do not share the caste-specific expression patterns seen in the other species finding
Experimental setups
Assay System Perturbation Readout Platform
bulk brain RNA-seq / de novo transcriptome assembly 9 vespid wasp species (whole adult female brains, pooled queen and worker samples) none (natural caste phenotype) transcript per million (TPM) gene expression, caste-biased differential expression
orthology inference / phylogenomics (Orthofinder) single-copy orthologs across 9 wasp species none orthologous gene groups; phylogenetic tree of Hymenoptera Orthofinder
principal component analysis of gene expression brain transcriptomes, 9 species, before/after between-species normalisation none clustering by species vs caste; % variance explained by PCs
differential expression analysis brain transcriptomes per species, caste comparison none (queen vs worker) caste-biased differentially expressed orthologous genes edgeR
gene ontology enrichment analysis caste-biased DEG gene sets across species (n=562 genes) none overrepresented GO terms (Bonferroni corrected)
support vector machine classification with feature selection brain transcriptome gene expression matrices, 9 species (train on 8, test on 1) none caste classification certainty (0-1 scale)
morphometric analysis queens and workers from several colonies per species none evidence of morphological caste differentiation
Key results
  • After species-normalisation, brain transcriptome samples separate largely by queen vs worker caste along the top two principal components
  • 57 orthologous genes are caste-biased in four or more of nine species 57 genes
  • 353 orthologous genes show caste-biased expression in the same direction in at least two of nine species 353 genes
  • No ortholog showed consistent caste-biased differential expression across all nine species
  • Vitellogenin-family orthogroup OG0001418 is upregulated in queens 8 of 9 species
  • Oocyte zinc finger 22 is queen-biased
  • Pheromone, chitin/odorant binding and muscle-related GO terms are overrepresented among caste-biased genes shared across species
  • Observed overlap in caste-biased DEGs across species is greater than expected by chance
Key statistics
  • count 57 orthologous genes (caste-biased in ≥4 of 9 species)
  • count 353 orthologous genes (caste-biased in same direction in ≥2 of 9 species)
  • count 562 genes (gene set used for GO enrichment analysis (DE in ≥2 species))
  • pvalue unadjusted P < 0.05 (threshold used to call caste-biased DEGs)
  • pvalue Fisher two-sided P values (comparison of observed vs expected (1000 permutations) DEG overlap frequencies)
  • count 3-11 biological replicates per caste pool (RNA-seq pooled samples per species/caste)
  • count 9 species (vespid wasp species sampled, spanning Polistinae and Vespinae)
  • other SVM trained on 8 species, tested on held-out 9th species (leave-one-species-out cross-species caste classification validation)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used brain transcriptome data from nine vespid wasp species (queen vs worker pools per species) to test for a shared genetic toolkit for caste, combining conventional differential expression analysis (edgeR) with functional enrichment testing and a supervised machine-learning approach (support vector machines with linear-regression-based feature selection) to classify caste from gene expression using a leave-one-species-out training/testing design. Overlap of caste-biased genes across species was assessed with a permutation-based test (1000 permutations) and Fisher two-sided P values, and gene ontology enrichment used single-tailed, Bonferroni-corrected P values. Species-level variation was addressed via a between-species normalisation of TPM values prior to principal component analysis.

Replicationbiological Sample sizeEach species' caste pool (queen or worker) was made up of 3 to 11 biological replicates; SVM models were trained on eight species and tested on the ninth GroupsQueens vs workers (caste) across nine social wasp species, and comparisons across levels of social complexity, lineage, and colony-founding mode Pairingunclear Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBonferroni correction (for GO term enrichment); unadjusted P<0.05 was explicitly used when assessing consistency of caste-biased differential expression across all nine species
Statistical tests used
Test Applied to n Assumptions
edgeR differential expression analysis (per-species queen vs worker comparison) Identification of caste-biased orthologous genes across the nine species (Fig. 3a, c) Pooled RNA samples per caste per species, each pool made up of 3-11 biological replicates not stated
Permutation test (1000 permutations) with Fisher two-sided P values Testing whether observed overlap of caste-biased DEGs across species exceeds that expected by chance (Fig. 3a) not stated
Gene ontology (GO) overrepresentation test, single-tailed, Bonferroni corrected Functional enrichment of genes differentially expressed in at least two of nine species (Fig. 3b; n = 562 genes) 562 genes tested against a background of all genes tested across the nine species not stated
Support vector machine (SVM) classification with linear-regression-based feature selection Predicting queen/worker caste identity from global gene expression, trained on eight species and tested on the withheld ninth species (Figs. 4-5) Eight species used for training, one held out for testing, iterated across species not stated
Orthofinder-based ortholog inference and phylogenetic reconstruction Construction of single-copy orthogroups and a Hymenoptera phylogeny (Supplementary Fig. 1) na
Approaches that could also have been used
  • Consistency of caste-biased expression across the nine species was assessed using an unadjusted P<0.05 threshold rather than a multiplicity-corrected one.
    Could also: A false discovery rate correction (e.g., Benjamini-Hochberg) applied across the combined set of cross-species comparisons — FDR correction controls the expected proportion of false positives when many genes and species comparisons are evaluated together, which can complement threshold-based reporting when the number of simultaneous tests is large.
  • Gene ontology enrichment was tested with single-tailed P values corrected via the Bonferroni method.
    Could also: Benjamini-Hochberg FDR correction for GO enrichment — FDR-based correction is generally less conservative than Bonferroni and is commonly used in GO enrichment analyses involving many terms, offering additional power while still accounting for multiple comparisons.
  • Differential expression between castes was analyzed using edgeR, a negative-binomial GLM-based method for count data with few biological replicates per group.
    Could also: DESeq2 or limma-voom — These are similarly standard RNA-seq differential expression tools with comparable assumptions about count data dispersion; the choice among them often reflects pipeline compatibility or reviewer/field convention rather than a difference in validity.
  • Overlap in caste-biased genes across species was tested using a permutation-based approach (1000 permutations) with Fisher two-sided P values.
    Could also: An analytic hypergeometric or Fisher's exact test for gene set overlap — Analytic overlap tests can offer a computationally simpler complement to permutation-based approaches, though permutation methods can better account for non-independence among genes or species.
  • Caste was classified from global gene expression using a supervised SVM with linear-regression-based feature selection and a leave-one-species-out design.
    Could also: Alternative supervised classifiers such as random forests or regularized (e.g., LASSO) logistic regression — These approaches provide alternative ways to assess how well expression patterns predict caste and can offer complementary measures of feature importance, which may be useful alongside SVM-based classification.
  • Species-level variation was addressed by scaling TPM values to a -1 to 1 range (species-normalisation) prior to PCA.
    Could also: A mixed-effects model with species included as a random effect, or a variance-stabilizing transformation — Modeling species as a random effect can retain information about the magnitude of caste differences while still accounting for species-level variation, as an alternative to explicit rescaling before ordination.
Software: edgeR · Orthofinder · Support vector machine (SVM) implementation (package not specified in this excerpt)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36828829

Paper: Wyatt CDR et al. 2023. Social complexity, life-history and lineage influence the molecular basis of castes in vespid wasps. Nat Commun 14:1046. DOI 10.1038/s41467-023-36456-6 · PMCID PMC9958023.

Data: GEO GSE159973 — whole-brain RNA-seq, 9 vespid wasp species, 40 samples (pooled within castes for analysis). The deposit ships processed products: Trinity de-novo assemblies (GSE159973_FASTA_files.tar.gz) and RSEM isoform counts (GSE159973_RSEM_results.tar.gz), plus the raw reads (SRA).

Code (two layers):

  1. github.com/biocorecrg/transcriptome_assembly (the URL in the registry) — a generic Nextflow de-novo transcriptome assembly pipeline (Trimmomatic → Trinity → TransDecoder → RSEM quantify → annotation). This is the upstream assembly tool. Third-party-style generic tool (P16) — equally valid.
  2. github.com/chriswyatt1/Trinity_to_SVM @ v1.0.0 (Zenodo 10.5281/zenodo.7521827) — the authors' own downstream analysis pipeline (Nextflow + Perl Master.ML.pl + R). It downloads exactly the two GSE159973 supplementary tarballs, restructures them per species/caste, and runs the SVM caste classification + feature selection in container chriswyatt/perl_r_e1071. It even ships the OrthoFinder output (DATA/Orthofinder/Orthogroups*.csv), so orthology need not be rerun.

This is a strong reproduction target: a fully scripted, containerized pipeline acting on openly-deposited processed data.


IN SCOPE (pipeline-derived, attempted)

# Reported result Pipeline / tool Reproduction route
R1 Trinity de-novo assemblies for 9 species (transcripts per species) Trinity v2.8(.4) via biocore transcriptome_assembly Verify deposited FASTA assemblies (count transcripts/genes); optionally re-run assembly on one species' reads as a 1:1 spot check
R2 RSEM isoform quantification per species×caste RSEM v1.3.1 Verify deposited RSEM .isoforms.results (file count, columns, row counts)
R3 SVM caste-classification gene sets (core toolkit ≈259; simpler-society toolkit 277; complex-society toolkit 289) Trinity_to_SVM SVM (R e1071) + linear-regression feature selection Run Master.ML.pl/Nextflow on the deposited data → compare per-stage gene counts + classifier figure (Fig 4/5)
R4 edgeR caste-biased orthologs (57 genes in ≥4 species; 353 genes in ≥2 of 9 species) edgeR v3.26.5, dispersion 0.1, FDR<0.05 Run downstream R in Trinity_to_SVM on deposited counts + shipped orthology
R5 SVM∩edgeR overlap (35 genes; 14 uncharacterised) both above Derived once R3+R4 reproduced
R6 Orthogroup counts for dN/dS (1,221 single-copy; 1,831 ≤3-isoform; 5,536 relaxed; 1,971 used) OrthoFinder v2.2.7 (+diamond, muscle, FastTree) Recount from shipped Orthogroups*.csv files

OUT OF SCOPE (not attempted — not a deterministic pipeline output)

  • Wet-lab: brain dissection, RNA extraction, library prep, sequencing.
  • Biological/behavioural sampling and caste assignment in the field.
  • GO/synaptic-transport enrichment interpretation of the "top 400 genes" (downstream interpretation, depends on annotation DBs not pinned).
  • dN/dS (PAML) selection inference itself (heavy, many gene trees; the orthogroup counts feeding it (R6) are in scope, the selection test is not a first target).
  • Phylogeny / major-transition evolutionary inference (narrative).

Hard blocker note

All heavy compute runs on «our HPC» («infra») via «host» ssh. At repro start the VPN tunnel was down (ssh timeouts); per brief, not touched — control-plane work (paper, code, scope, claims, profiling design) done first, compute queued for when the tunnel returns.

Figures / tables: Fig. 5aFig. 5bFig. 4a
R1
Reported
9 Trinity de-novo brain transcriptome assemblies (1 per species)
Reproduced
9/9 Trinity FASTA present; 213,599-300,170 transcripts per species
exact
R2
Reported
RSEM isoform quantification per species x caste
Reproduced
18 RSEM .isoforms.results (9 spp x Queen/Worker), 8-col RSEM format, rows==assembly transcripts
exact
R3a
Reported
top 259 SVM-predictive caste genes (~20% from lm p<0.05)
Reproduced
default config: 1283/2707 single-copy orthogroups with lm p<0.05; SVM classifier + tables produced
partial
R3b
Reported
277 gene toolkit (simpler societies, Fig 5a)
Reproduced
config-specific grouping not pinned in default pipeline; not isolated
partial
R3c
Reported
289 gene toolkit (complex societies, Fig 5b)
Reproduced
config-specific grouping not pinned in default pipeline; not isolated
partial
R4a
Reported
57 edgeR caste-biased orthologs (>=4 species)
Reproduced
edgeR DE ran per species within pipeline; cross-species >=4 intersection is a downstream-config step not pinned in default
partial
R4b
Reported
353 edgeR caste-biased orthologs (>=2 of 9 species)
Reproduced
see R4a
partial
R6a
Reported
1,221 single-copy orthogroups (Fig 4a)
Reproduced
shipped 9-spp noMac matrix: 2516 single-copy in all 9 (raw); != 1221
did not match
R6b
Reported
1,831 orthogroups <=3 isoforms (Fig 4a)
Reproduced
present-in-all-9 with <=3 iso = 3994 (raw); != 1831
did not match
R6c
Reported
5,536 relaxed orthogroups (Fig 4a)
Reproduced
not directly reproducible from shipped noMac files
partial
R6d
Reported
1,971 orthogroups for dN/dS
Reproduced
not directly reproducible from shipped noMac files
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 52/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

63 k
tokens (I/O) · 3.4 M incl. cache
8 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.