Genetic parallels in biomineralization of the calcareous sponge Sycon ciliatum and stony corals.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Bulk RNA-seq DEG reproduction (osculum vs basal wall, |log2FC|>=2 & padj<0.01) for the calcareous sponge Sycon ciliatum. The paper's own analysis scripts are stated to be on GitHub but no resolvable URL exists; the repo named in the brief (Stylophora_single_cell_atlas) is the coral atlas (Levy 2021) whose UMI counts this paper REUSED, not its own code. Per P16 I reimplemented the paper's described standard pipeline on its own public data: ENA PRJEB78728 (15 runs, 5 individuals x OSC/IN/OUT) quantified with Salmon against an HBWS01 reference + exact tx2gene built from the authors' deposited annotation (Zenodo 14755899, Sci_HBWS01_all_info.tsv), then tximport->DESeq2 (design ~individual+group2). Result: 1175 DEGs total (paper 1575; ratio 0.75) of which 497 up in osculum (paper 829) and 678 up in basal, over 14623 genes tested. Same data, same thresholds, same analysis class, same-order-of-magnitude counts => PARTIAL (described well enough to reproduce qualitatively, not 1:1). Mean Salmon mapping rate ~50%, consistent with the reference being HBWS01 CDS (ORFs only, no UTR) rather than full transcripts; that plus unpublished exact DESeq2 params (prefilter/independent-filtering/IN-OUT pooling) is the most likely driver of the lower counts. No fabrication signal: the reported numbers are plausibly derivable from the deposited data+method. NOT attempted (out of scope / 80-20): proteomics (35 spicule matrix proteins, calcarin LC-MS/MS), AlphaFold structures, calcarin family curation (C3=17), WGCNA midnightblue module (C4=196, secondary), and the coral single-cell reuse (C6).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-15 ⛓ b2b61f25a28a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether the calcareous sponge Sycon ciliatum uses genetic mechanisms for calcite spicule biomineralization that are similar to those used by stony corals for aragonite skeleton formation, consistent with a shared, pre-adapted 'biomineralization toolkit' inherited from a common ancestor.
- ★ 829 genes are overexpressed in regions of increased calcite spicule formation in S. ciliatum, including known sclerocyte-specific biomineralization genes. finding
- ★ 17 calcarins (Cal1-Cal17), proteins analogous to coral galaxins, were identified in the S. ciliatum genome, are secreted, localize to the spicule matrix, and are expressed in sclerocytes. finding
- ★ Calcarin expression varies temporally and spatially and is specific to certain spicule types and sclerocyte stages, indicating fine-tuned gene regulation controls biomineralization. finding
- ★ Tandem gene arrangements and expression changes suggest gene duplication and neofunctionalization significantly shaped S. ciliatum's biomineralization, similar to corals. mechanism
- ★ Carbonate biomineralization evolved in parallel in calcitic S. ciliatum and aragonitic corals. finding
- Calcarins share a common structural motif with galaxins: beta-hairpins formed from di-cysteine residues linked by disulfide bridges, as predicted by AlphaFold. finding
- Previously known sclerocyte-specific biomineralization genes (CA1, CA2, AE-like1, AE-like2, Diactinin, Triactinin, Spiculin) are among those overexpressed in the oscular region. finding
- Calcarins appear unique to calcareous sponges and absent from other sponge lineages. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| differential gene expression (RNA-seq) analysis | Sycon ciliatum, apical oscular region vs. inner/outer body wall | none (regional comparison) | differentially expressed genes (log2FC, padj) | — |
| GO-term enrichment analysis | Sycon ciliatum transcripts (Uniprot-annotated closest hits) | none | enriched biological process GO terms among oscular-overexpressed genes | — |
| BLASTp similarity search | Sycon ciliatum predicted genome proteins | none | identification of galaxin-similar proteins (calcarins) | BLASTp |
| AlphaFold structure prediction | S. ciliatum calcarins and coral (e.g. Acropora millepora) galaxins/galaxin-like proteins | none | predicted monomer tertiary structure, disulfide bridge/beta-hairpin arrangement | AlphaFold |
| chromogenic in situ hybridization (CISH) | Sycon ciliatum sclerocytes | none | spatial localization of Cal1-Cal8 mRNA expression | — |
| hairpin chain reaction fluorescence in situ hybridization (HCR-FISH) | Sycon ciliatum sclerocytes/spicule-forming regions | none | co-localization of Cal1-Cal8, Spiculin, Triactinin, SciCA1 expression across founder/thickener cells and developmental stages | — |
| skeletal matrix proteomics | cnidarian (coral/octocoral) skeletal matrix | none | detection/absence of galaxin-like proteins in skeletal matrix proteome | — |
| orthogroup/comparative genomics analysis | sponges and corals (multiple species) | none | number and orthogroup assignment of calcarin, galaxin-like, and galaxin transcripts | — |
- – 1575 differentially expressed genes identified between oscular region and body wall (log2-fold change ≥2, padj<0.01) log2FC≥2, padj<0.01
- ▲ 829 genes overexpressed in the oscular growth region, including CA1, CA2, AE-like1, AE-like2, Diactinin, Triactinin, Spiculin
- – 17 calcarins (Cal1-17) identified via BLASTp as galaxin-similar proteins in S. ciliatum genome, secreted (signal peptide present) 17 proteins
- ▲ 14 secreted proteins with significantly higher expression in the oscular region showed similarity to coral galaxin/galaxin-like proteins 14 proteins
- – Calcarin expression is sclerocyte-type and stage specific (e.g., Cal1 in triactine/tetractine founder cells ceasing upon transition to thickener cells; Cal2 in diactine founder cells; Cal7 in early triactine founder cells only)
- ▲ GO enrichment of oscular-overexpressed genes for skeletal system development, ossification, and bone mineralization terms
- – Galaxin-like proteins in orthogroups shared with calcarins were not detected in the skeletal matrix proteomes of the respective cnidarian species
- count 1575 differentially expressed genes (DGE analysis, oscular vs. body wall regions, 5 specimens)
- count 829 genes overexpressed in oscular region (subset of DEGs upregulated in spicule-forming growth zone)
- fold_change log2-fold change ≥2 (DEG inclusion threshold)
- pvalue padj<0.01 (DEG significance threshold)
- count 17 calcarins (Cal1-Cal17) (galaxin-similar proteins identified by BLASTp in S. ciliatum genome)
- count 14 secreted proteins similar to galaxin with higher oscular expression (subset of calcarins with significant oscular overexpression)
- count 10-23 di-cysteine residues per region, separated by 10-15 amino acids (structural feature of calcarins/galaxins)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used differential gene expression (DGE) analysis to compare transcriptomes of the apical oscular region versus inner and outer body walls across five Sycon ciliatum specimens, applying log2-fold change and adjusted p-value thresholds to identify genes overexpressed in the region of higher spicule formation. Complementary GO-term enrichment analysis was performed on the overexpressed gene set, annotated via closest Uniprot BLAST hits. Structural and sequence analyses (BLASTp, AlphaFold monomer predictions) characterized novel galaxin-like proteins (calcarins), and RNA in situ hybridization (chromogenic CISH and HCR-FISH) validated spatial and temporal expression patterns in sclerocytes. Outcomes were reported primarily as gene counts, fold-change thresholds, and qualitative expression patterns rather than continuous summary statistics.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential gene expression analysis — specific package not named in available text; padj notation is consistent with DESeq2 Wald test or edgeR quasi-likelihood F-test (negative binomial model); thresholds log2FC ≥2 and padj <0.01 | Apical oscular region vs. inner and outer body walls; yielded 1575 DEGs total, 829 overexpressed in oscular region | 5 specimens | not stated |
| GO-term overrepresentation / enrichment analysis; specific test (e.g., Fisher's exact, hypergeometric) not named in available text | Set of 829 genes overexpressed in the oscular region, annotated via closest Uniprot BLAST hit | 829 upregulated genes as query set; background not described in available text | not stated |
| BLASTp sequence similarity search | Identification of galaxin-like proteins (calcarins) from predicted S. ciliatum proteome; 17 calcarins identified | — | na |
| AlphaFold monomer structure prediction with visual inspection of beta-hairpin and disulfide-bridge topology | Structural comparison of calcarins vs. coral galaxins | — | na |
-
The specific DGE package and normalization strategy are not named in the available text; padj is reported as the significance criterion↳ Could also: Explicitly reporting the package (e.g., DESeq2 with Wald test, edgeR with QL-F test), version, normalization method (e.g., median-of-ratios, TMM), and dispersion estimation strategy would fully specify the analysis — Different DGE tools apply different dispersion estimation and normalization strategies that can yield different gene lists, especially at small n (here n=5); explicit reporting supports reproducibility and facilitates meta-analysis
-
GO-term enrichment was performed by annotating S. ciliatum genes via their closest Uniprot BLAST hit, then testing the overexpressed set for overrepresentation↳ Could also: Graph-aware enrichment methods (e.g., topGO with the 'elim' or 'weight' algorithm, or clusterProfiler with redundancy filtering) could reduce inflation from nested GO terms; transfer annotation uncertainty could be addressed by requiring reciprocal best hits or an e-value cutoff — Closest-hit annotation can propagate incorrect GO terms when the nearest hit is only distantly related; graph-aware methods reduce over-counting of hierarchically nested terms and improve interpretability of enriched biological processes
-
A fixed log2-fold change threshold (≥2) was applied alongside the adjusted p-value to define the DEG set↳ Could also: Ranked-list methods such as fgsea/GSEA, or the independent hypothesis weighting (IHW) framework, analyze all genes without a hard fold-change cutoff — Threshold-free ranked methods capture coordinated shifts across gene sets and can detect biologically meaningful signals from genes that fall just below an arbitrary fold-change cutoff, potentially relevant for subtle biomineralization-related expression changes noted in the discussion
-
Structural similarity between calcarins and galaxins was assessed via AlphaFold monomer predictions and qualitative visual inspection of beta-hairpin and disulfide-bridge patterns↳ Could also: Quantitative structural comparison using TM-score (TM-align), Dali server Z-scores, or FATCAT would complement the visual comparison with a scale-free numerical similarity metric — Quantitative alignment scores provide an objective, reproducible measure of structural similarity that can be compared across protein pairs and reported alongside the predicted structures, supporting evolutionary inference
-
No formal power analysis or minimum-detectable-effect statement accompanied the choice of n=5 biological specimens for DGE↳ Could also: An a priori power analysis (e.g., using RNASeqPower or powsimR) or a post-hoc sensitivity statement reporting the minimum detectable log2FC at 80% power for n=5 could be included — Reporting power or sensitivity contextualizes the gene list for readers — particularly relevant when comparing a non-model organism with a small sample size — and helps distinguish truly non-differentially-expressed genes from those that may have been undetectable at the given sample size
-
The evolutionary relationship between calcarins and coral galaxins was inferred from sequence similarity (BLASTp), structural topology (AlphaFold), and orthogroup assignment↳ Could also: A formal phylogenetic analysis (maximum likelihood or Bayesian, with bootstrap or posterior probability support) of the full galaxin/calcarin/galaxin-like protein family would provide explicit statistical support for the convergent-evolution inference — Phylogenetic trees with branch support values quantify the confidence in topological claims (e.g., independent origins vs. deep homology), directly addressing the paper's central evolutionary hypothesis with a statistical framework
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 40922549 (eLife 2025, DOI 10.7554/elife.106239)
Title: Genetic parallels in biomineralization of the calcareous sponge Sycon ciliatum and stony corals. Repo in brief: github.com/sebepedroslab/Stylophora_single_cell_atlas — this is NOT the paper's own analysis code. It is the coral single-cell atlas (Levy et al., Cell 2021) from which this paper reused raw UMI counts + celltype assignments for the coral comparison. The paper's own analysis scripts are stated to be "in a GitHub repository" but no resolvable URL is given in the article or PMC, and no Sycon/biomineralization repo exists in the sebepedroslab org (16 repos checked, none match). Per brief rule P16, reproducing by running the described standard pipeline on the paper's own public data is equally valid → we do that (independent reimplementation of the stated method).
Pipeline-derived results (candidate claims)
| # | Reported result | Pipeline | In scope? | Reason |
|---|---|---|---|---|
| C1 | 1575 DEGs (osculum vs basal wall, |log2FC|≥2, padj<0.01) | Salmon → DESeq2 | YES (primary) | Fully specified; public data + reference |
| C2 | 829 DEGs overexpressed in oscular region (subset of C1) | Salmon → DESeq2 | YES | Same run as C1 |
| C3 | 17 calcarins (Cal1–Cal17) | manual/OrthoFinder + domain curation | partial/out | Curation + proteomics-driven; not a single deterministic pipeline output |
| C4 | 196 genes in WGCNA 'midnightblue' meta-module | WGCNA | secondary | Depends on many soft params (soft-power, merge cut); harder to hit 1:1 |
| C5 | 35 acid-insoluble spicule matrix proteins; 7–8 calcarins in matrix | LC-MS/MS + MASCOT/Scaffold | OUT (wet-lab) | Proteomics, not a reproducible compute pipeline from public raw |
| C6 | 980 calicoblast cells express 14 known coral SOMPs | reuse of Stylophora atlas (Seurat) | out (80/20) | Re-analysis of external single-cell atlas; lower-hanging target is C1/C2 |
Primary reproduction target = C1 + C2 (the clearly-specified, low-hanging bulk-RNA-seq DEG count). C4 (WGCNA) attempted only if C1/C2 land cleanly and budget allows.
Inputs (all public)
- Reads: ENA PRJEB78728, body-part dataset, 15 runs ERR13472820–ERR13472834 (5 individuals GW30948/30951/30956/30957/30959 × 3 regions OSC/IN/OUT). ~6 GB compressed.
- Reference transcriptome: Caglar et al. 2021 S. ciliatum transcriptome = GenBank TSA HBWS01
(Zenodo 14755899 names all annotation files
Sci_HBWS01_*; paper: "high-quality transcriptome … better BUSCO"). Fetch fasta from ENA/NCBI on the compute node. - Annotations: Zenodo 14755899 →
Sci_HBWS01_all_info.tsv(transcript→gene/GO/domain map).
Method (as described in paper)
- Trim/filter reads (paper: "trimmed and filtered"; use fastp, default-ish).
- Salmon quant against HBWS01 transcriptome (decoy-aware not specified for transcriptome-only → plain index).
- tximport → DESeq2. Contrast = osculum (OSC) vs basal wall (IN + OUT).
Design accounts for the 5 paired individuals:
~ individual + group. - DEGs: |log2FC|≥2 & padj<0.01. Count total (expect 1575) and up-in-OSC (expect 829).
Out of scope (not attempted), and why
- Proteomics (C5), AlphaFold structures, in-situ/wet-lab calcarin localisation — not pipeline-from-public-raw.
- Exact calcarin family curation (C3) — manual + orthology + structural reasoning; not 1:1 deterministic.
- Coral single-cell re-analysis (C6) — external atlas; primary low-hanging target is the bulk DEG count.
Honesty notes
- The exact DESeq2 script (trimming params, whether IN+OUT pooled vs modeled separately, prefilter,
independent filtering, lfcShrink) is not published, so an exact 1575/829 match is not guaranteed;
a within-order-of-magnitude / same-direction count is a legitimate
partial/within-toloutcome. - If HBWS01 fasta cannot be retrieved, fall back to the Zenodo BRAKER gene predictions (GCA_964019385) as reference and flag the reference substitution expl
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a solid partial reproduction: an independent reimplementation of the paper's described bulk RNA-seq pipeline on the authors' own public reads (ENA PRJEB78728) and deposited annotation returned 1175 DEGs (paper 1575; ratio 0.75) and 497 up-in-osculum (paper 829; ratio 0.60) — same order of magnitude, no fabrication signal. The deviation sits on our methodology / data-version side (a CDS-only HBWS01 reference giving ~50% mapping, plus unpublished exact DESeq2 params), compounded by an authors-side gap — the stated analysis code has no resolvable URL. Severity is moderate: total magnitude holds but the OSC/basal direction skew reverses (paper majority up in OSC, reproduction majority up in basal), so the body-part expression conclusion is only partially confirmed.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.