Corpus 1,285 assessed · 1,186 scored · 647 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Transcriptomic analysis of the highly derived radial body plan of a sea urchin.

Genome Biol Evol · 2014
68/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
68/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1186 studies
🎯 Scores higher than 33% of all assessed papers rank 769 of 1186 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the full pipeline-derived chain of this paper (Trinity de novo assembly -> Bowtie/RSEM quantification -> edgeR differential expression -> Mfuzz clustering -> BLAST reciprocal-best-hit annotation -> pyEnrichment-style functional enrichment) from raw SRA reads (PRJNA243078, all 21 libraries) using a from-scratch, independently run pipeline plus the paper's own third-party tool (ofedrigo/pyEnrichment, whose statistical method was re-implemented in R because its bundled category-mapping data is human-only and its legacy Python2/syck dependency is unrunnable today). Results are a genuine mix of close/exact and clearly divergent: assembly min-contig-length matched EXACTLY (201 nt) and N50 was within ~5% (1685 vs 1769), but max contig length diverged ~2x, as expected from independently re-run stochastic Trinity assemblies. Of the two key edgeR stage-transition comparisons checked in depth, one (early_juvenile vs late_juvenile) reproduced within ~2-13% on every up/down/tested count; the other (advanced_rudiment vs early_juvenile) showed a real, unresolved mismatch -- ~2.2x more DE genes than reported and the dominant up/down direction flipped. The paper's documented '<5 counts excluded' filter was implemented and tested but was shown NOT to explain this gap (run_DE_analysis.pl/edgeR already applies its own internal per-comparison low-count filtering, so the pre-filter changed the overall matrix size 118491->94691 genes but left this comparison's tested-gene count unchanged at 41961/41328); the residual discrepancy most plausibly reflects the much larger total 'gene' universe from this independently-assembled from-scratch transcriptome versus the paper's own assembly + any undocumented redundancy-reduction steps, not a fixable bug in this pipeline -- this is reported honestly as an open reproducibility gap, not papered over. Mfuzz clustering (paper's exact stated parameters, c=16 fuzzifier m=1.55) reproduced almost exactly: 3340 vs 3316 genes retained at membership>=0.5 (<1% difference), the strongest single result in this study, on an input gene set that itself was ~17% smaller than the paper's (3654 vs 4421). RBB annotation reproduced within ~11% (12000 vs 13548 unique genes) despite substituting OrfPredictor+blastp (unavailable in a modern conda environment) with a blastx/tblastn reciprocal-best-hit approach against the SAME reference protein set and e-value threshold. Functional enrichment reproduced the paper's STATISTICAL METHOD (top-10%-by-significance one-sided Fisher exact test, BH FDR<=5%) and produced a comparable order of magnitude of significant categories (34 and 47 vs the paper's 44 total for one comparison), but on a fundamentally different, non-comparable category system: pyEnrichment's bundled PANTHER/GO category-mapping files were found to be human-only (0 of ~62k entries reference S. purpuratus), so the paper's legacy PANTHER v8.1 Biological Process scheme (coarse, ~hundreds of terms) had to be substituted with PANTHER's current official sea-urchin classification, which reports modern GO Biological Process terms (thousands of much more granular terms) -- explicitly graded 'partial', not attempted to be forced into a false exact match. NOT attempted / out of scope: any wet-lab, morphological, or manual-curation results in the paper (not pipeline-derived, per Hard Rule #1); no attempt was made to hand-tune the from-scratch Trinity assembly parameters to chase a closer match to the paper's exact gene/transcript counts, since Trinity version/parameter differences and stochastic assembly are a known, expected, and honestly-reported source of de novo transcriptome divergence rather than a pipeline bug. All large intermediates (raw/trimmed FASTQ ~117G, RSEM per-sample BAMs ~179G, Trinity assembly working directories ~640G) were deleted from «infra» scratch after final result extraction (934G -> 2.8G), since all are exactly regenerable from the retained SRA accessions, retained Trinity.fasta assembly, and retained job scripts; only the final assembly FASTA, quantification matrices, DE/Mfuzz/annotation/enrichment result tables, and job scripts/logs were kept.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-09
Rubric version
not recorded
Assessed by
Last updated
2026-08-09

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper asks how developmental changes in gene expression underlie the highly derived radial (pentameral) adult body plan and metamorphosis of echinoderms, focusing on whether larval neurogenesis genes are repurposed during adult nervous system development or whether a suite of adult-specific neurogenesis transcription factors is deployed.

Core claims
  • A de novo reference transcriptome for Heliocidaris erythrogramma spanning larval, metamorphic, and postmetamorphic stages provides a genomic resource for studying radial body plan evolution. resource
  • The most extensive changes in transcript abundance occur during and after metamorphosis, indicating these phases are transcriptionally more complex than earlier ones. finding
  • Distinct sets of regulatory and effector proteins are used in pre- versus postmetamorphic development, though the separation of gene sets is far from absolute. finding
  • Larval neurogenesis transcription factors are also deployed during adult nervous system development, alongside an adult-specific set of neurogenesis transcription factors. finding
  • Developmental genes are expressed in three broad phases: larval regulators declining across the time-course, skeletal/mesodermal genes peaking prior to metamorphosis, and neural development genes activated after metamorphosis. finding
  • Late juveniles show enrichment of metabolic, energy-generation, and translation/ribosomal protein genes, consistent with increased demands of growth and metabolism. finding
  • Fuzzy c-means clustering combined with PANTHER categorical enrichment and multidimensional-scaling 'bubble plots' resolves stage-specific expression programs across the life cycle. method
Experimental setups
Assay System Perturbation Readout Platform
de novo reference transcriptome assembly sea urchin Heliocidaris erythrogramma, multiple developmental stages none assembled reference transcript sequences; 13,017 genes examined
RNA-seq developmental time-course Heliocidaris erythrogramma, seven stages: 24, 28, 32, 36, 40 hpf (premetamorphic) and 96 hpf, 10 dpf (postmetamorphic), biological replicates none (developmental time-course) gene expression levels (counts per million, FPKM), log2 fold-change, differential expression at FDR 5%
principal components analysis of expression Heliocidaris erythrogramma, expression averaged across biological replicates of seven stages none proportion of variance explained by PC1/PC2; separation of pre- vs postmetamorphic stages
categorical enrichment analysis (hypergeometric test) Heliocidaris erythrogramma advanced rudiment (40 hpf), early juvenile (96 hpf), late juvenile (10 dpf) transcriptomes none enriched PANTHER biological process categories among top 10% of differentially expressed genes at FDR 5% PANTHER ontology database
fuzzy c-means clustering with Wilcoxon rank test enrichment Heliocidaris erythrogramma developmental time-course, differentially expressed genes (log2FC ≥ 2, FDR ≤ 0.001) none 16 expression clusters, membership scores, enriched PANTHER biological process categories per cluster
targeted expression profiling of neurogenesis transcription factors Heliocidaris erythrogramma time-course; gene sets defined from the Strongylocentrotus purpuratus genome (Burke et al. 2006) none FPKM expression trajectories of 35 larval neurogenesis TFs and 18 putative (non-larval) neurogenesis TFs
multidimensional scaling of category distance matrix (bubble plots) PANTHER biological process categories enriched in H. erythrogramma stage comparisons none 2D layout of category relatedness; significance by bubble color, occupancy by bubble size
Key results
  • PC1 separates pre- from postmetamorphic stages, while PC2 explains nearly all variance among premetamorphic stages; overall variation recapitulates the developmental time-course. PC1 = 72% of variation; PC2 = 18%
  • Across metamorphosis (40 to 96 hpf), 4,544 genes were differentially expressed at FDR 5%, with 2,322 downregulated and 2,222 upregulated. 4,544 genes (33.5% of those examined)
  • During postmetamorphic development (96 hpf to 10 dpf), 2,679 genes were downregulated and 2,960 upregulated at FDR 5%. 2,679 down / 2,960 up
  • A cluster of 63 highly expressed genes further increased in late juveniles; all but two (Sm37, a spicule matrix protein, and Chrna9, a cholinergic receptor) encode ribosomal proteins. 63 genes
  • Advanced rudiment stage is enriched for RNA/DNA metabolism, RNA processing, and cell division categories, whereas early juveniles are enriched for cell signaling, development, neurogenesis, transport, and reproduction.
  • All 35 known larval nervous system transcription factors are also expressed in postmetamorphic stages; among 18 putative non-larval neurogenesis TFs, most peak after metamorphosis (nurr1, phox2, pou4f, pitx3, lxh3-4) or stay low, while a few peak premetamorphically (barh1, meis, myt, sp8). 35 larval TFs; 18 putative adult TFs
  • Fuzzy c-means clustering retained 3,316 genes with membership score ≥ 0.5, yielding stage-specific profiles: translation/RNA processing declining (Cluster 2), cell cycle/DNA replication rising (Cluster 5), developmental genes peaking premetamorphically (Cluster 6), neural/immune/signaling genes rising post-metamorphosis (Cluster 13), and metabolism/translation peaking in late juveniles (Cluster 16). 3,316 of 4,421 clustered genes retained
  • Enriched categories were clearly delineated between advanced rudiment and early juvenile, but less well separated between early and late juveniles, and ~33% of clustered genes were unclassified. almost 33% unclassified
Key statistics
  • other PC1 explains 72% of variation; PC2 explains 18% (PC analysis of gene expression averaged across replicates for seven developmental stages)
  • count 4,544 differentially expressed genes (33.5%), 2,322 down and 2,222 up (40 hpf (advanced rudiment) vs 96 hpf (early juvenile) at FDR 5%)
  • count 2,679 downregulated; 2,960 upregulated (96 hpf (early juvenile) vs 10 dpf (late juvenile) at FDR 5%)
  • count 13,017 genes (total genes assessed in MA plots of expression across the time-course)
  • fold_change log2FC ≥ 2 (at least 4-fold), FDR ≤ 0.001 (thresholds for genes entered into fuzzy c-means clustering)
  • count 3,316 clustered genes retained (membership score ≥ 0.5) in 16 clusters (fuzzy c-means clustering of developmental expression patterns)
  • count 63 genes in cluster; all but two encode ribosomal proteins (highly expressed cluster further increasing from 96 hpf to 10 dpf)
  • count 31 enriched categories in early juveniles; 13 enriched categories in late juveniles (PANTHER hypergeometric enrichment, early vs late juvenile comparison at FDR 5%)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used RNA-seq to profile gene expression across seven developmental stages of the sea urchin Heliocidaris erythrogramma, examining global variation via principal components analysis and assessing differential expression between key stage-pair comparisons (e.g., pre- vs postmetamorphic) using count-based tests controlled at a 5% false discovery rate. Functional category enrichment among differentially expressed genes was tested with a hypergeometric test against the PANTHER ontology database, and temporal expression patterns were further summarized by fuzzy c-means clustering, with per-cluster functional enrichment assessed by a Wilcoxon signed-rank test. Results are reported mainly as gene counts, FDR-thresholded significance calls, and log2 fold-changes rather than exact p-values or confidence intervals.

Replicationbiological Sample sizeBiological replicates are mentioned (expression values were averaged across biological replicates for the PC analysis), but the number of replicates per stage and any power/sample-size justification are not stated in the provided text. GroupsSeven developmental stages (24, 28, 32, 36, 40 hpf larval/premetamorphic; 96 hpf and 10 dpf postmetamorphic), with primary contrasts of advanced rudiment (40 hpf) vs early juvenile (96 hpf) and early juvenile (96 hpf) vs late juvenile (10 dpf) Pairingunclear Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFalse discovery rate (FDR) control at 5% (the specific FDR procedure, e.g., Benjamini-Hochberg, is not named in the text)
Statistical tests used
Test Applied to n Assumptions
Differential expression testing (RNA-seq count-based comparison; specific statistical model/software not named in the provided text) Pairwise comparisons of gene expression between developmental stages (e.g., 40 hpf advanced rudiment vs 96 hpf early juvenile; 96 hpf vs 10 dpf late juvenile) across 13,017 genes not stated
Hypergeometric test Categorical enrichment of PANTHER biological process categories among the top 10% of differentially expressed genes in each stage comparison not stated
Wilcoxon signed-rank test Enrichment of PANTHER biological process categories within each of the 16 fuzzy c-means expression clusters (supplementary table S5) 3,316 clustered genes (membership score ≥ 0.5) drawn from 4,421 genes with log2 fold-change ≥ 2 and FDR ≤ 0.001 not stated
Principal components analysis Global variation in gene expression across the seven developmental stages (fig. 1B), based on expression averaged across biological replicates na
Approaches that could also have been used
  • Differential expression between stages was reported as significance calls at a 5% FDR threshold, without exact p-values or confidence intervals for fold-changes
    Could also: Reporting exact (or FDR-adjusted) p-values alongside fold-change confidence intervals — This would let readers gauge the strength of evidence for each gene beyond a binary significant/non-significant call and assess the precision of the estimated fold-change.
  • Functional category enrichment was tested with a hypergeometric test applied only to the top 10% of differentially expressed genes in each comparison
    Could also: Gene set enrichment analysis (GSEA) using the full ranked gene list — A rank-based, threshold-free approach can detect coordinated but modest shifts in a category's genes that might not individually clear a hard top-10% cutoff.
  • Gene expression across the time-course was analyzed largely through separate pairwise stage comparisons (e.g., 40 hpf vs 96 hpf, 96 hpf vs 10 dpf)
    Could also: A time-course/trajectory model (e.g., spline-based methods such as maSigPro or ImpulseDE2) — Such models treat developmental time as continuous and can characterize overall expression trajectories directly, rather than combining information from discrete pairwise contrasts.
  • Genes were grouped into 16 clusters via fuzzy c-means clustering with a fixed membership-score threshold (ms ≥ 0.5) for inclusion
    Could also: Model-based clustering (e.g., Gaussian mixture models via Mclust) or hierarchical clustering with a data-driven cut height — These approaches can estimate the number of clusters and membership thresholds from the data itself, which is a complementary way to select clustering parameters.
  • The PC analysis was performed on gene expression averaged across biological replicates rather than on individual replicate measurements
    Could also: Performing PCA on individual replicate-level expression values — This would additionally visualize within-stage (replicate-to-replicate) variability alongside between-stage variability, giving a fuller picture of the data's overall structure.
  • Measures of dispersion (e.g., SD, SEM, or CI) for expression values are not reported in the main text
    Could also: Reporting a dispersion measure such as SD, SEM, or a 95% CI alongside mean/average expression values — This would communicate the spread or uncertainty of expression estimates at each stage, which is especially informative when replicate numbers are limited.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

assembly_n50
Reported
Trinity de novo assembly of 21 RNA-seq libraries (7 stages x 3 reps): N50 = 1769 nt (longest-isoform/'unigene' set)
Reproduced
N50 = 1685 nt (longest-isoform-per-Trinity-gene set, 118,491 genes / 251,066 transcripts total; get_longest_isoform_seq_per_trinity_gene.pl)
within tolerance
assembly_min_contig
Reported
Minimum contig length = 201 nt
Reproduced
Minimum contig length = 201 nt (exact match)
exact
assembly_max_contig
Reported
Maximum contig length = 9873 nt
Reproduced
Maximum contig length = 19043 nt (~1.9x paper's value)
did not match
de_early_juvenile_vs_late_juvenile
Reported
early_juvenile(96hpf) vs late_juvenile(10dpf): 2679 genes down + 2960 up at FDR<=5% (edgeR)
Reproduced
tested=41328, sig(FDR<=0.05)=5827, up=3005, down=2822 (own from-scratch Trinity assembly + Bowtie/RSEM/edgeR pipeline, min-count-5 filtered matrix)
within tolerance
de_advanced_rudiment_vs_early_juvenile
Reported
advanced_rudiment(40hpf) vs early_juvenile(96hpf): 13017 genes tested, 4544 DE (33.5%) at FDR<=5% (2322 down + 2222 up)
Reproduced
tested=41961, sig(FDR<=0.05)=10136, up=5806, down=4330 -- ~2.2x more DE genes than reported AND the dominant direction is flipped (paper: down>up; reproduced: up>down)
did not match
mfuzz_de_geneset_size
Reported
Union of genes with |log2FC|>=2 & FDR<=0.001 across 6 adjacent-stage transitions used as Mfuzz input: 4421 genes
Reproduced
3654 genes (same filter, own edgeR output)
within tolerance
mfuzz_retained_genes
Reported
Mfuzz soft clustering (c=16, fuzzifier m=1.55) on TMM-normalized stage means; 3316 genes retained at max membership score >=0.5 (incl. 1102 of unknown function)
Reproduced
3340 genes retained at max membership score >=0.5 (same c=16, m=1.55 parameters)
exact
rbb_annotation_count
Reported
13548 unique genes annotated via reciprocal-best-BLAST-hit (RBB) vs S. purpuratus RefSeq v3.1 proteins (OrfPredictor + blastp, e<1e-10)
Reproduced
12000 unique genes via RBB (blastx forward + tblastn reverse, since OrfPredictor+blastp is unavailable in a modern conda env; documented substitution of ORF-calling method, same reference protein set GCF_000002235.3, same e-value threshold)
within tolerance
functional_enrichment_categories
Reported
pyEnrichment top-10%-by-significance one-sided Fisher exact enrichment (BH FDR<=5%) vs legacy PANTHER v8.1 Biological Process scheme: early_juvenile vs late_juvenile gave 31 categories up-in-early + 13 up-in-late = 44 total
Reproduced
Same statistical method (top-10% by FDR vs rest, one-sided Fisher exact, BH FDR<=0.05) applied against PANTHER's current official sea-urchin classification (PTHR19.0_sea_urchin, modern GO Biological Process terms, since pyEnrichment's bundled category-mapping files are human-only -- 0/62k entries reference S. purpuratus): 34 GO-BP categories at FDR<=0.05 for early_juvenile vs late_juvenile; 47 for advanced_rudiment vs early_juvenile. Category system is fundamentally different (few hundred coarse legacy PANTHER BP terms vs thousands of granular modern GO BP terms), so counts are not directly comparable even though the same order of magnitude was reproduced.
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.