Genetic parallels in biomineralization of the calcareous sponge Sycon ciliatum and stony corals.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Bulk RNA-seq DEG reproduction (osculum vs basal wall, |log2FC|>=2 & padj<0.01) for the calcareous sponge Sycon ciliatum. The paper's own analysis scripts are stated to be on GitHub but no resolvable URL exists; the repo named in the brief (Stylophora_single_cell_atlas) is the coral atlas (Levy 2021) whose UMI counts this paper REUSED, not its own code. Per P16 I reimplemented the paper's described standard pipeline on its own public data: ENA PRJEB78728 (15 runs, 5 individuals x OSC/IN/OUT) quantified with Salmon against an HBWS01 reference + exact tx2gene built from the authors' deposited annotation (Zenodo 14755899, Sci_HBWS01_all_info.tsv), then tximport->DESeq2 (design ~individual+group2). Result: 1175 DEGs total (paper 1575; ratio 0.75) of which 497 up in osculum (paper 829) and 678 up in basal, over 14623 genes tested. Same data, same thresholds, same analysis class, same-order-of-magnitude counts => PARTIAL (described well enough to reproduce qualitatively, not 1:1). Mean Salmon mapping rate ~50%, consistent with the reference being HBWS01 CDS (ORFs only, no UTR) rather than full transcripts; that plus unpublished exact DESeq2 params (prefilter/independent-filtering/IN-OUT pooling) is the most likely driver of the lower counts. No fabrication signal: the reported numbers are plausibly derivable from the deposited data+method. NOT attempted (out of scope / 80-20): proteomics (35 spicule matrix proteins, calcarin LC-MS/MS), AlphaFold structures, calcarin family curation (C3=17), WGCNA midnightblue module (C4=196, secondary), and the coral single-cell reuse (C6).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-15 ⛓ b2b61f25a28a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper tests whether the molecular and genetic machinery underlying calcite spicule biomineralization in the calcareous sponge Sycon ciliatum shares parallels with the aragonite biomineralization toolkit of stony corals, hypothesizing that gene duplication and neofunctionalization shaped biomineralization in both lineages independently.
- ★ 829 genes are overexpressed in the oscular region of increased calcite spicule formation in S. ciliatum finding
- ★ 17 galaxin-analogous proteins, named calcarins (Cal1-Cal17), are localized in the spicule matrix and expressed in sclerocytes finding
- ★ Calcarins are secreted proteins (signal peptides) structurally analogous to coral galaxins, sharing di-cysteine beta-hairpin disulfide-bridge architecture mechanism
- ★ Calcarin expression varies temporally and spatially, specific to certain spicule types and sclerocyte types (founder vs thickener cells), indicating fine-tuned gene regulation controls biomineralization finding
- ★ Tandem gene arrangements and expression changes indicate gene duplication and neofunctionalization shaped S. ciliatum biomineralization, paralleling corals mechanism
- ★ Carbonate biomineralization evolved in parallel in calcitic S. ciliatum and aragonitic corals, exemplifying convergent evolution of reef-builders finding
- Calcarins appear unique to calcareous sponges and absent from other sponge lineages finding
- A combined transcriptomic, genomic, proteomic, and in situ hybridization approach identifies spicule-formation genes in S. ciliatum method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Differential gene expression (bulk RNA-seq) analysis | Sycon ciliatum (calcareous sponge), five specimens | none (regional comparison: apical oscular region vs inner/outer basal body walls) | differentially expressed genes (log2-fold change >=2, padj<0.01) | — |
| GO-term enrichment analysis | S. ciliatum genes/transcripts | none | enriched biological-process GO terms via closest Uniprot hit annotation | Uniprot |
| Genome/proteome prediction and BLASTp homology search | S. ciliatum predicted proteins | none | galaxin-like proteins (calcarins) identified by sequence similarity | BLASTp |
| Protein structure prediction | S. ciliatum calcarins and coral (A. millepora) galaxins | none | predicted tertiary structure / disulfide-bridged beta-hairpins | AlphaFold |
| Chromogenic RNA in situ hybridization (CISH) | S. ciliatum tissue | none | spatial mRNA expression of Cal1-Cal8, Spiculin, Triactinin, SciCA1 in sclerocytes | — |
| Hairpin chain reaction fluorescence in situ hybridization (HCR-FISH) | S. ciliatum (incl. regenerated asconoid juvenile stage) | none | spatial/temporal calcarin expression in founder vs thickener cells | — |
| Proteomic analysis of spicule matrix | S. ciliatum calcite spicules | none | proteins occluded/localized in spicule matrix (calcarins) | — |
| Orthogroup / comparative phylogenetic analysis | sponges and corals (calcarin, galaxin-like, galaxin transcripts) | none | orthogroup assignment and transcript counts | — |
- – 1575 genes differentially expressed between oscular region and basal body walls 1575 genes
- ▲ 829 genes overexpressed in the oscular region (spicule-forming zone), including known biomineralization genes CA1, CA2, AE-like1, AE-like2, Diactinin, Triactinin, Spiculin 829 genes
- ▲ 14 secreted proteins with higher oscular expression show similarity to coral galaxin/galaxin-like proteins; 17 calcarins total identified from genome 14 overexpressed; 17 total
- – Calcarins contain 10-23 di-cysteine residues separated by 10-15 amino acids, forming disulfide-bridged four-amino-acid beta-hairpins like galaxins 10-23 di-cysteines
- – Calcarins show spicule-type-specific expression: Cal2 in diactine founder cells, Cal1 in triactine/tetractine founder cells, Spiculin in thickener cells
- – Cal1 expression in founder cells ceases as they transform into thickener cells and Spiculin expression begins, with transient co-expression marking the transition
- – Choanoderm chambers form at approximately 90° angle to the radial tube axis ~90°
- – Diactine formation involves two sclerocytes and triactine formation involves six 2 vs 6 cells
- count 1575 differentially expressed genes (DGE oscular vs basal body wall, log2FC>=2, padj<0.01)
- count 829 genes overexpressed in oscular region (upregulated in spicule-forming zone)
- count 17 calcarins (Cal1-Cal17) (galaxin-like proteins identified from S. ciliatum genome by BLASTp)
- count 14 secreted galaxin-similar proteins (significantly higher expression in oscular region)
- fold_change log2-fold change >=2 (DGE significance threshold)
- pvalue padj<0.01 (adjusted p-value threshold for DGE)
- count 10-23 di-cysteine residues per calcarin, separated by 10-15 amino acids (calcarin/galaxin structural motif)
- other ~90° angle of choanoderm chambers to tube axis (S. ciliatum body architecture)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used differential gene expression (DGE) analysis to compare transcriptomes of the apical oscular region versus inner and outer body walls across five Sycon ciliatum specimens, applying log2-fold change and adjusted p-value thresholds to identify genes overexpressed in the region of higher spicule formation. Complementary GO-term enrichment analysis was performed on the overexpressed gene set, annotated via closest Uniprot BLAST hits. Structural and sequence analyses (BLASTp, AlphaFold monomer predictions) characterized novel galaxin-like proteins (calcarins), and RNA in situ hybridization (chromogenic CISH and HCR-FISH) validated spatial and temporal expression patterns in sclerocytes. Outcomes were reported primarily as gene counts, fold-change thresholds, and qualitative expression patterns rather than continuous summary statistics.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential gene expression analysis — specific package not named in available text; padj notation is consistent with DESeq2 Wald test or edgeR quasi-likelihood F-test (negative binomial model); thresholds log2FC ≥2 and padj <0.01 | Apical oscular region vs. inner and outer body walls; yielded 1575 DEGs total, 829 overexpressed in oscular region | 5 specimens | not stated |
| GO-term overrepresentation / enrichment analysis; specific test (e.g., Fisher's exact, hypergeometric) not named in available text | Set of 829 genes overexpressed in the oscular region, annotated via closest Uniprot BLAST hit | 829 upregulated genes as query set; background not described in available text | not stated |
| BLASTp sequence similarity search | Identification of galaxin-like proteins (calcarins) from predicted S. ciliatum proteome; 17 calcarins identified | — | na |
| AlphaFold monomer structure prediction with visual inspection of beta-hairpin and disulfide-bridge topology | Structural comparison of calcarins vs. coral galaxins | — | na |
-
The specific DGE package and normalization strategy are not named in the available text; padj is reported as the significance criterion↳ Could also: Explicitly reporting the package (e.g., DESeq2 with Wald test, edgeR with QL-F test), version, normalization method (e.g., median-of-ratios, TMM), and dispersion estimation strategy would fully specify the analysis — Different DGE tools apply different dispersion estimation and normalization strategies that can yield different gene lists, especially at small n (here n=5); explicit reporting supports reproducibility and facilitates meta-analysis
-
GO-term enrichment was performed by annotating S. ciliatum genes via their closest Uniprot BLAST hit, then testing the overexpressed set for overrepresentation↳ Could also: Graph-aware enrichment methods (e.g., topGO with the 'elim' or 'weight' algorithm, or clusterProfiler with redundancy filtering) could reduce inflation from nested GO terms; transfer annotation uncertainty could be addressed by requiring reciprocal best hits or an e-value cutoff — Closest-hit annotation can propagate incorrect GO terms when the nearest hit is only distantly related; graph-aware methods reduce over-counting of hierarchically nested terms and improve interpretability of enriched biological processes
-
A fixed log2-fold change threshold (≥2) was applied alongside the adjusted p-value to define the DEG set↳ Could also: Ranked-list methods such as fgsea/GSEA, or the independent hypothesis weighting (IHW) framework, analyze all genes without a hard fold-change cutoff — Threshold-free ranked methods capture coordinated shifts across gene sets and can detect biologically meaningful signals from genes that fall just below an arbitrary fold-change cutoff, potentially relevant for subtle biomineralization-related expression changes noted in the discussion
-
Structural similarity between calcarins and galaxins was assessed via AlphaFold monomer predictions and qualitative visual inspection of beta-hairpin and disulfide-bridge patterns↳ Could also: Quantitative structural comparison using TM-score (TM-align), Dali server Z-scores, or FATCAT would complement the visual comparison with a scale-free numerical similarity metric — Quantitative alignment scores provide an objective, reproducible measure of structural similarity that can be compared across protein pairs and reported alongside the predicted structures, supporting evolutionary inference
-
No formal power analysis or minimum-detectable-effect statement accompanied the choice of n=5 biological specimens for DGE↳ Could also: An a priori power analysis (e.g., using RNASeqPower or powsimR) or a post-hoc sensitivity statement reporting the minimum detectable log2FC at 80% power for n=5 could be included — Reporting power or sensitivity contextualizes the gene list for readers — particularly relevant when comparing a non-model organism with a small sample size — and helps distinguish truly non-differentially-expressed genes from those that may have been undetectable at the given sample size
-
The evolutionary relationship between calcarins and coral galaxins was inferred from sequence similarity (BLASTp), structural topology (AlphaFold), and orthogroup assignment↳ Could also: A formal phylogenetic analysis (maximum likelihood or Bayesian, with bootstrap or posterior probability support) of the full galaxin/calcarin/galaxin-like protein family would provide explicit statistical support for the convergent-evolution inference — Phylogenetic trees with branch support values quantify the confidence in topological claims (e.g., independent origins vs. deep homology), directly addressing the paper's central evolutionary hypothesis with a statistical framework
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 40922549 (eLife 2025, DOI 10.7554/elife.106239)
Title: Genetic parallels in biomineralization of the calcareous sponge Sycon ciliatum and stony corals. Repo in brief: github.com/sebepedroslab/Stylophora_single_cell_atlas — this is NOT the paper's own analysis code. It is the coral single-cell atlas (Levy et al., Cell 2021) from which this paper reused raw UMI counts + celltype assignments for the coral comparison. The paper's own analysis scripts are stated to be "in a GitHub repository" but no resolvable URL is given in the article or PMC, and no Sycon/biomineralization repo exists in the sebepedroslab org (16 repos checked, none match). Per brief rule P16, reproducing by running the described standard pipeline on the paper's own public data is equally valid → we do that (independent reimplementation of the stated method).
Pipeline-derived results (candidate claims)
| # | Reported result | Pipeline | In scope? | Reason |
|---|---|---|---|---|
| C1 | 1575 DEGs (osculum vs basal wall, |log2FC|≥2, padj<0.01) | Salmon → DESeq2 | YES (primary) | Fully specified; public data + reference |
| C2 | 829 DEGs overexpressed in oscular region (subset of C1) | Salmon → DESeq2 | YES | Same run as C1 |
| C3 | 17 calcarins (Cal1–Cal17) | manual/OrthoFinder + domain curation | partial/out | Curation + proteomics-driven; not a single deterministic pipeline output |
| C4 | 196 genes in WGCNA 'midnightblue' meta-module | WGCNA | secondary | Depends on many soft params (soft-power, merge cut); harder to hit 1:1 |
| C5 | 35 acid-insoluble spicule matrix proteins; 7–8 calcarins in matrix | LC-MS/MS + MASCOT/Scaffold | OUT (wet-lab) | Proteomics, not a reproducible compute pipeline from public raw |
| C6 | 980 calicoblast cells express 14 known coral SOMPs | reuse of Stylophora atlas (Seurat) | out (80/20) | Re-analysis of external single-cell atlas; lower-hanging target is C1/C2 |
Primary reproduction target = C1 + C2 (the clearly-specified, low-hanging bulk-RNA-seq DEG count). C4 (WGCNA) attempted only if C1/C2 land cleanly and budget allows.
Inputs (all public)
- Reads: ENA PRJEB78728, body-part dataset, 15 runs ERR13472820–ERR13472834 (5 individuals GW30948/30951/30956/30957/30959 × 3 regions OSC/IN/OUT). ~6 GB compressed.
- Reference transcriptome: Caglar et al. 2021 S. ciliatum transcriptome = GenBank TSA HBWS01
(Zenodo 14755899 names all annotation files
Sci_HBWS01_*; paper: "high-quality transcriptome … better BUSCO"). Fetch fasta from ENA/NCBI on the compute node. - Annotations: Zenodo 14755899 →
Sci_HBWS01_all_info.tsv(transcript→gene/GO/domain map).
Method (as described in paper)
- Trim/filter reads (paper: "trimmed and filtered"; use fastp, default-ish).
- Salmon quant against HBWS01 transcriptome (decoy-aware not specified for transcriptome-only → plain index).
- tximport → DESeq2. Contrast = osculum (OSC) vs basal wall (IN + OUT).
Design accounts for the 5 paired individuals:
~ individual + group. - DEGs: |log2FC|≥2 & padj<0.01. Count total (expect 1575) and up-in-OSC (expect 829).
Out of scope (not attempted), and why
- Proteomics (C5), AlphaFold structures, in-situ/wet-lab calcarin localisation — not pipeline-from-public-raw.
- Exact calcarin family curation (C3) — manual + orthology + structural reasoning; not 1:1 deterministic.
- Coral single-cell re-analysis (C6) — external atlas; primary low-hanging target is the bulk DEG count.
Honesty notes
- The exact DESeq2 script (trimming params, whether IN+OUT pooled vs modeled separately, prefilter,
independent filtering, lfcShrink) is not published, so an exact 1575/829 match is not guaranteed;
a within-order-of-magnitude / same-direction count is a legitimate
partial/within-toloutcome. - If HBWS01 fasta cannot be retrieved, fall back to the Zenodo BRAKER gene predictions (GCA_964019385) as reference and flag the reference substitution expl
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a solid partial reproduction: an independent reimplementation of the paper's described bulk RNA-seq pipeline on the authors' own public reads (ENA PRJEB78728) and deposited annotation returned 1175 DEGs (paper 1575; ratio 0.75) and 497 up-in-osculum (paper 829; ratio 0.60) — same order of magnitude, no fabrication signal. The deviation sits on our methodology / data-version side (a CDS-only HBWS01 reference giving ~50% mapping, plus unpublished exact DESeq2 params), compounded by an authors-side gap — the stated analysis code has no resolvable URL. Severity is moderate: total magnitude holds but the OSC/basal direction skew reverses (paper majority up in OSC, reproduction majority up in basal), so the body-part expression conclusion is only partially confirmed.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.