Identification and Mechanisms of Osteocyte Subsets Involved in the Pathological Progression of Osteoporosis.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL — described well enough only for its one public-data claim; the central single-cell results are NOT reproducible because the authors withheld the data. Data Availability Statement reads verbatim 'Research data are not shared', and the Methods give NO accession for the mouse-femur sham/OVX scRNA-seq that every headline result derives from (CellRanger v7.0.1 -> Seurat v4.2.0 -> Harmony -> DoubletFinder -> CellChat v1.6.1 -> Monocle2). So C2 (24851 cells/23491 genes), C3 (6 osteocyte subsets, BHR-Ocy markers, 701 osteocytes), C4 (the listed repo CellChat's Sema5a-Plxna1 axis) and C5 (Fig 2H proportions) are honest non-attempts (data_restricted): the listed third-party tool (CellChat) cannot be exercised because its required processed Seurat object is not shared. What IS reproducible 1:1 is the single public-data claim: GSE230665 (human Agilent GPL10332, 12 PMOP vs 3 control femur) used in Fig S12 for 'Sema5a significantly higher in PMOP than control'. On «our HPC» («job», conda GEOquery 2.70.0 + limma 3.58.1 / R 4.3.3, all data on «infra») the standard pipeline gives SEMA5A probe 21646 mean OP=2639 vs CTRL=527, higher in OP, Welch p=0.00109, limma logFC=2.02 adj.P=0.015 -> direction AND significance reproduced exactly (qualitative claim, no number to hit, hence within-tol). NOT attempted (and why): the entire single-cell pipeline + CellChat + Monocle, because the underlying data is not public. Fabrication note: C1 is independently supported by the public dataset (no fabrication signal); C2-C5 are NOT independently checkable because the data was withheld, so no fabrication judgement is made -- but flag that the paper's main computational findings are not reproducible as published (data not shared, no accession, listed repo not runnable on available data). Verdict provisional; human audit sheet in AUDIT.md.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 57assessed: 2026-06-15 ⛓ 42db7254c73c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDistinct osteocyte subpopulations play specific roles in the onset and progression of osteoporosis; the study tests whether osteocyte subset heterogeneity, particularly a bone homeostasis regulatory subset (BHR-Ocys), drives pathological bone loss and through what molecular mechanism.
- ★ Six distinct osteocyte subsets (C1-C6) exist in mouse bone, identified by single-cell sequencing. finding
- ★ The Egfr+ Il1r1+ Sema5a+ osteocyte subset (BHR-Ocys) is a key subpopulation regulating bone homeostasis via osteoblasts and osteoclasts. finding
- ★ The proportion of BHR-Ocys is increased in OVX (postmenopausal osteoporotic) mice and contributes to bone loss. finding
- ★ BHR-Ocys-derived Sema5a binds osteoclast receptor Plxna1 to promote osteoclast differentiation and bone resorption. mechanism
- ★ The Sema5a-Plxna1 axis activates PI3K/AKT/Myc signaling to drive osteoclastogenesis and pathological bone resorption. mechanism
- Osteocyte-derived Sema5a affects osteoclast function but has no significant effect on osteoblast differentiation or mineralization. finding
- Sema5a requires RANKL co-stimulation to promote osteoclast differentiation; Sema5a alone is insufficient. finding
- The Sema5a-Plxna1 axis is a potential therapeutic target for osteoporosis prevention and treatment. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA sequencing | mouse femoral bone tissue (sham vs OVX) | ovariectomy (OVX) | cell cluster identity and subset gene expression / proportions | — |
| flow cytometry sorting | Dmp1-Cre+;tdTomato mouse femur/tibia osteocytes (sham vs OVX) | OVX | ratio of tdTomato+ Egfr+ Il1r1+ BHR-Ocys | — |
| Western blot | sorted BHR-Ocys vs non-BHR-Ocys from OVX mice | none | RANKL/Tnfsf11, Spp1, Csf-1, SOST/sclerostin protein levels | — |
| TRAP staining (osteoclast coculture/differentiation) | BMDMs cocultured with BHR-Ocys or non-BHR-Ocys; BMDMs +/- recombinant Sema5a | coculture / recombinant Sema5a + RANKL + M-CSF | TRAP+ multinucleated (>=3 nuclei) osteoclast number and bone resorption area | — |
| immunofluorescence staining | cultured BHR-Ocys vs non-BHR-Ocys osteocytes | none | Sema5a fluorescence intensity | — |
| micro-CT and DXA | WT vs Sema5a-CKO (Dmp1-Cre;Sema5a fl/fl) mice, +/- OVX | osteocyte-specific Sema5a knockout + OVX | BMD, BV/TV, Tb.N, Tb.Th, Tb.Sp, Ct.Th; serum CTX-1; Oc.S/BS | — |
| Plxna1 knockdown / Lyz2-Cre conditional knockout (Plxna1-CKO) | BMDMs with Plxna1 knockdown; Plxna1-CKO mice +/- OVX | Plxna1 knockdown / osteoclast-specific Plxna1 knockout | osteoclast differentiation, bone resorption, Oc.S/BS, serum CTX-1 | — |
| RT-qPCR / ALP staining / ARS staining | MC3T3-E1 osteoblasts | recombinant Sema5a treatment | ALP, Runx2, OCN, OPN expression; ALP activity; mineralization | — |
- ▲ Proportion of BHR-Ocys significantly increased in OVX mice compared to sham
- ▲ RANKL, Spp1, Csf-1, and sclerostin protein levels significantly higher in BHR-Ocys than non-BHR-Ocys
- ▲ BHR-Ocys promoted BMDM differentiation into osteoclasts whereas non-BHR-Ocys did not
- ▲ Sema5a-pretreated BMDMs showed significant osteoclastogenesis with suboptimal RANKL vs untreated control
- ▼ Sema5a-CKO mice after OVX showed decreased Oc.S/BS and lower serum CTX-1 vs WT
- – Sema5a-CKO mice after OVX showed increased femur/total BMD, BV/TV, Tb.N, Tb.Th, Ct.Th and decreased Tb.Sp
- ▼ Plxna1 knockdown abolished Sema5a-induced BMDM differentiation and bone resorption
- – Sema5a had no significant effect on osteoblast markers (ALP, Runx2, OCN, OPN), differentiation, or mineralization
- count 701 cells in total across 8 osteocyte subclusters (osteocyte subgroup single-cell analysis)
- count 6 osteocyte subsets retained (C1-C6) after excluding erythroid contaminants (osteocyte subset identification)
- count n = 6 per group (flow cytometry BHR-Ocy ratio in sham vs OVX)
- count each Western blot sample = BHR-Ocys from femurs/tibias of 10 mice, 6 samples per group (BHR-Ocy protein quantification)
- count n = 3 (TRAP+ osteoclast quantification in coculture)
- other RANKL 5 ng/mL (suboptimal), M-CSF 10 ng/mL, Sema5a 10 µg/mL (osteoclast induction conditions)
- pvalue p < 0.05, p < 0.01, p < 0.001 (significance thresholds throughout figures)
- count ≈100 cells counted per group, 3 fields per sample, n = 3 (Sema5a immunofluorescence intensity)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study combines single-cell RNA sequencing (scRNA-seq) of mouse femoral bone with in vitro co-culture assays and in vivo genetic knockout models (OVX-induced osteoporosis) to characterize osteocyte subset heterogeneity. Primary group comparisons (sham vs. OVX; BHR-Ocys vs. non-BHR-Ocys; WT vs. conditional knockouts) were analyzed with two-tailed unpaired Student's t-tests for two-group contrasts and one-way/two-way ANOVA with Tukey-Kramer post-hoc testing for multi-group contrasts. Results were reported as mean ± SD with threshold-based p-value symbols (* p<0.05, ** p<0.01, *** p<0.001); exact p-values, effect sizes, and confidence intervals were not provided.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-tailed unpaired Student's t-test | Flow cytometry quantification of BHR-Ocy proportion in sham vs. OVX mice (Figure 2H) | n=6 per group | not stated |
| two-tailed unpaired Student's t-test | Western blot quantification of RANKL, Spp1, Csf-1, and sclerostin in BHR-Ocys vs. non-BHR-Ocys (Figures 3B–E); each sample pooled from 10 mice, 6 samples per group | n=6 per group | not stated |
| two-tailed unpaired Student's t-test | TRAP-positive multinucleated cell count in BHR-Ocy vs. non-BHR-Ocy co-culture with BMDMs (Figure 3H) | n=3 per group | not stated |
| two-tailed unpaired Student's t-test or ANOVA with post-hoc Tukey-Kramer test | In vivo Sema5a-CKO experiments: Oc.S/BS, serum CTX-1, femur BMD, total BMD, BV/TV, Tb.N, Tb.Th, Tb.Sp, Ct.Th across WT/CKO × sham/OVX groups (Figures 4H–Q) | n=6 per group | not stated |
| Pearson correlation | Heatmap of pairwise correlations among 8 osteocyte subclusters (Figure S1B) | 701 cells total across subclusters | not stated |
| Gene Ontology (GO) enrichment analysis | Characteristic genes of BHR-Ocys vs. other osteocyte subsets (Figure 2E; Figure S3); method for significance not stated in provided text | — | not stated |
-
Cell type proportions (ratio of BHR-Ocys among all osteocytes) were compared between sham and OVX groups using a two-tailed unpaired t-test on flow cytometry values↳ Could also: Compositional data analysis methods such as Dirichlet regression (DirichletReg in R) or the scCODA framework could also be used for scRNA-seq-derived cell-type proportions — Cell-type proportions are compositional (they sum to 1 within each sample), which violates the independence assumption of a t-test on individual proportions; compositional methods account for this constraint and can model uncertainty in estimated proportions, particularly when total cell counts per sample are small
-
Single-cell RNA sequencing differential gene expression between osteocyte subclusters was used to define subcluster identities and BHR-Ocy marker genes; the specific DE test was not stated in the provided text↳ Could also: Pseudobulk approaches (e.g., DESeq2 or edgeR on aggregated per-animal counts per cluster) could also be used alongside or instead of single-cell-level tests such as Wilcoxon rank-sum or MAST — Pseudobulk methods treat biological replicates (animals) rather than individual cells as the unit of replication, which better controls type-I error inflation from treating thousands of cells as independent observations; they are increasingly recommended for scRNA-seq differential expression when multiple biological samples are available
-
Multiple independent two-group Student's t-tests were applied across several figures and comparisons (Western blot, flow cytometry, TRAP counts) without a stated family-wise error correction↳ Could also: A single linear model or one-way ANOVA framework with a Bonferroni or Benjamini-Hochberg FDR correction applied across the family of comparisons could also be used — When many independent t-tests are conducted within a study, the probability of at least one false positive increases; a pre-specified correction strategy applied to the full family of tests would provide an explicit bound on the overall error rate
-
Data dispersion was reported exclusively as mean ± SD, with n as small as 3 for some in vitro assays↳ Could also: Individual data points plotted alongside the mean ± SD (e.g., dot plots or strip plots overlaid on bar or box plots) could also be used, as recommended by many journals — When n=3, the SD describes only three observations and the mean may not be representative; showing individual values allows readers to assess the full distribution, detect outliers, and evaluate whether the summary statistics are informative
-
Statistical significance was reported only as threshold-based symbols (* p<0.05, ** p<0.01, *** p<0.001) without exact p-values or effect size metrics↳ Could also: Reporting exact p-values together with a standardized effect size (e.g., Cohen's d for t-tests, partial η² for ANOVA) and 95% confidence intervals could also be done — Exact p-values allow readers and meta-analysts to assess the strength of evidence on a continuous scale; effect sizes convey biological magnitude independently of sample size and are necessary to judge whether statistically significant differences are also practically meaningful
-
Western blot protein quantification compared BHR-Ocys vs. non-BHR-Ocys with an unpaired t-test; each Western blot sample consisted of cells pooled from 10 mice, with 6 pooled samples per group↳ Could also: A non-parametric Mann-Whitney U test could also be considered for this comparison — With n=6 biological replicates per group and Western blot data that may not be normally distributed (due to pooling and gel-to-gel variability), a non-parametric alternative does not require the normality assumption of the t-test and can be more robust when the sample size is small
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41250977
Paper: Jiang et al. 2025, Adv Sci. "Identification and Mechanisms of Osteocyte Subsets Involved in the Pathological Progression of Osteoporosis." PMID 41250977 · PMCID PMC12850396 · DOI 10.1002/advs.202513427. Listed code: github.com/sqjin/CellChat (third-party tool). Listed data: GSE230665.
The decisive availability fact
- Data Availability Statement (verbatim): "Research data are not shared."
- The Methods give no accession (no GEO/SRA/GSA/OMIX) for the study's own mouse femur sham/OVX single-cell RNA-seq — the dataset that ALL the headline computational results derive from. CellRanger v7.0.1 → Seurat v4.2.0 → Harmony → DoubletFinder → CellChat v1.6.1 → Monocle2 were all run on that private scRNA-seq matrix.
- The only public, obtainable dataset the paper uses is GSE230665 (human Agilent microarray, GPL10332; 12 postmenopausal-osteoporosis vs 3 healthy-control femur samples). It is used for a single validation claim.
In scope AND feasible (public data → attempted)
| id | result | location | pipeline |
|---|---|---|---|
| C1 | Sema5a expression significantly higher in PMOP patients vs control | Fig S12, text | standard GEOquery+limma on GSE230665 |
This is the only pipeline-derived result reproducible from public data. We run the standard third-party pipeline (GEOquery to fetch GSE230665 + GPL10332 annotation; limma moderated t-test + Welch t-test on SEMA5A probes) — equivalent to what the authors describe ("data obtained from GSE230665").
In scope but NOT feasible — DATA NOT SHARED (not attempted, drop reason)
All central single-cell results depend on the authors' unshared mouse scRNA-seq raw data ("Research data are not shared"; no accession):
- 24,851 cells / 23,491 genes after QC (<500 or >7500 genes, >10% mito, doublets)
- 6 osteocyte subsets C1–C6; BHR-Ocy = Egfr+ Il1r1+ Sema5a+; 701 osteocytes
- CellChat v1.6.1 cell–cell communication (Sema5a–Plxna1 axis, PI3K/AKT/Myc)
- Monocle2 pseudotime trajectory of osteocyte subsets
- BHR-Ocy proportion ↑ in OVX vs sham (Fig 2H)
Because the raw scRNA-seq is unobtainable, the CellChat tool (the listed repo) cannot be exercised on the paper's data. Reproducing CellChat would require a processed Seurat object with the authors' cell annotations, which is not shipped.
Out of scope (wet-lab / not a pipeline)
Sema5a-CKO and Plxna1-CKO mouse phenotyping (BV/TV, Tb.N/Th/Sp, Ct.Th; Fig 4), ImageJ histomorphometry, in-vitro osteoclast assays. Not attempted.
Expected outcome
partial — reproduce C1 (the one public-data claim) 1:1 in direction +
significance; everything central is a data_restricted/data_unavailable
non-attempt because the authors did not share the scRNA-seq data.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a partial reproduction driven by data withholding. The single peripheral public-data claim (C1, Sema5a higher in PMOP, Fig S12) reproduces 1:1 on GSE230665 (OP 2639 vs CTRL 527, p=0.00109, adj.P=0.0154), with no fabrication signal. The four central single-cell claims (24851 cells, 701 osteocytes, six osteocyte subsets, the CellChat Sema5a-Plxna1 axis, Fig 2H) are non-reproducible because the authors gave no accession and state 'Research data are not shared' — the problem sits on the authors'/data-availability side, not in a demonstrated computational error. Severity is indeterminate rather than severe: where testable the science held, but the headline results cannot be independently checked as published.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.