Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identification and Mechanisms of Osteocyte Subsets Involved in the Pathological Progression of Osteoporosis.

Adv Sci (Weinh) · 2025
L1 57/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
57/100
Reproducibility score
1.0 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 17% of all assessed papers rank 965 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL — described well enough only for its one public-data claim; the central single-cell results are NOT reproducible because the authors withheld the data. Data Availability Statement reads verbatim 'Research data are not shared', and the Methods give NO accession for the mouse-femur sham/OVX scRNA-seq that every headline result derives from (CellRanger v7.0.1 -> Seurat v4.2.0 -> Harmony -> DoubletFinder -> CellChat v1.6.1 -> Monocle2). So C2 (24851 cells/23491 genes), C3 (6 osteocyte subsets, BHR-Ocy markers, 701 osteocytes), C4 (the listed repo CellChat's Sema5a-Plxna1 axis) and C5 (Fig 2H proportions) are honest non-attempts (data_restricted): the listed third-party tool (CellChat) cannot be exercised because its required processed Seurat object is not shared. What IS reproducible 1:1 is the single public-data claim: GSE230665 (human Agilent GPL10332, 12 PMOP vs 3 control femur) used in Fig S12 for 'Sema5a significantly higher in PMOP than control'. On «our HPC» («job», conda GEOquery 2.70.0 + limma 3.58.1 / R 4.3.3, all data on «infra») the standard pipeline gives SEMA5A probe 21646 mean OP=2639 vs CTRL=527, higher in OP, Welch p=0.00109, limma logFC=2.02 adj.P=0.015 -> direction AND significance reproduced exactly (qualitative claim, no number to hit, hence within-tol). NOT attempted (and why): the entire single-cell pipeline + CellChat + Monocle, because the underlying data is not public. Fabrication note: C1 is independently supported by the public dataset (no fabrication signal); C2-C5 are NOT independently checkable because the data was withheld, so no fabrication judgement is made -- but flag that the paper's main computational findings are not reproducible as published (data not shared, no accession, listed repo not runnable on available data). Verdict provisional; human audit sheet in AUDIT.md.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 57
    assessed: 2026-06-15 ⛓ 42db7254c73c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Distinct osteocyte subpopulations play specific roles in the onset and progression of osteoporosis; the study tests whether osteocyte subset heterogeneity, particularly a bone homeostasis regulatory subset (BHR-Ocys), drives pathological bone loss and through what molecular mechanism.

Core claims
  • Six distinct osteocyte subsets (C1-C6) exist in mouse bone, identified by single-cell sequencing. finding
  • The Egfr+ Il1r1+ Sema5a+ osteocyte subset (BHR-Ocys) is a key subpopulation regulating bone homeostasis via osteoblasts and osteoclasts. finding
  • The proportion of BHR-Ocys is increased in OVX (postmenopausal osteoporotic) mice and contributes to bone loss. finding
  • BHR-Ocys-derived Sema5a binds osteoclast receptor Plxna1 to promote osteoclast differentiation and bone resorption. mechanism
  • The Sema5a-Plxna1 axis activates PI3K/AKT/Myc signaling to drive osteoclastogenesis and pathological bone resorption. mechanism
  • Osteocyte-derived Sema5a affects osteoclast function but has no significant effect on osteoblast differentiation or mineralization. finding
  • Sema5a requires RANKL co-stimulation to promote osteoclast differentiation; Sema5a alone is insufficient. finding
  • The Sema5a-Plxna1 axis is a potential therapeutic target for osteoporosis prevention and treatment. resource
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA sequencing mouse femoral bone tissue (sham vs OVX) ovariectomy (OVX) cell cluster identity and subset gene expression / proportions
flow cytometry sorting Dmp1-Cre+;tdTomato mouse femur/tibia osteocytes (sham vs OVX) OVX ratio of tdTomato+ Egfr+ Il1r1+ BHR-Ocys
Western blot sorted BHR-Ocys vs non-BHR-Ocys from OVX mice none RANKL/Tnfsf11, Spp1, Csf-1, SOST/sclerostin protein levels
TRAP staining (osteoclast coculture/differentiation) BMDMs cocultured with BHR-Ocys or non-BHR-Ocys; BMDMs +/- recombinant Sema5a coculture / recombinant Sema5a + RANKL + M-CSF TRAP+ multinucleated (>=3 nuclei) osteoclast number and bone resorption area
immunofluorescence staining cultured BHR-Ocys vs non-BHR-Ocys osteocytes none Sema5a fluorescence intensity
micro-CT and DXA WT vs Sema5a-CKO (Dmp1-Cre;Sema5a fl/fl) mice, +/- OVX osteocyte-specific Sema5a knockout + OVX BMD, BV/TV, Tb.N, Tb.Th, Tb.Sp, Ct.Th; serum CTX-1; Oc.S/BS
Plxna1 knockdown / Lyz2-Cre conditional knockout (Plxna1-CKO) BMDMs with Plxna1 knockdown; Plxna1-CKO mice +/- OVX Plxna1 knockdown / osteoclast-specific Plxna1 knockout osteoclast differentiation, bone resorption, Oc.S/BS, serum CTX-1
RT-qPCR / ALP staining / ARS staining MC3T3-E1 osteoblasts recombinant Sema5a treatment ALP, Runx2, OCN, OPN expression; ALP activity; mineralization
Key results
  • Proportion of BHR-Ocys significantly increased in OVX mice compared to sham
  • RANKL, Spp1, Csf-1, and sclerostin protein levels significantly higher in BHR-Ocys than non-BHR-Ocys
  • BHR-Ocys promoted BMDM differentiation into osteoclasts whereas non-BHR-Ocys did not
  • Sema5a-pretreated BMDMs showed significant osteoclastogenesis with suboptimal RANKL vs untreated control
  • Sema5a-CKO mice after OVX showed decreased Oc.S/BS and lower serum CTX-1 vs WT
  • Sema5a-CKO mice after OVX showed increased femur/total BMD, BV/TV, Tb.N, Tb.Th, Ct.Th and decreased Tb.Sp
  • Plxna1 knockdown abolished Sema5a-induced BMDM differentiation and bone resorption
  • Sema5a had no significant effect on osteoblast markers (ALP, Runx2, OCN, OPN), differentiation, or mineralization
Key statistics
  • count 701 cells in total across 8 osteocyte subclusters (osteocyte subgroup single-cell analysis)
  • count 6 osteocyte subsets retained (C1-C6) after excluding erythroid contaminants (osteocyte subset identification)
  • count n = 6 per group (flow cytometry BHR-Ocy ratio in sham vs OVX)
  • count each Western blot sample = BHR-Ocys from femurs/tibias of 10 mice, 6 samples per group (BHR-Ocy protein quantification)
  • count n = 3 (TRAP+ osteoclast quantification in coculture)
  • other RANKL 5 ng/mL (suboptimal), M-CSF 10 ng/mL, Sema5a 10 µg/mL (osteoclast induction conditions)
  • pvalue p < 0.05, p < 0.01, p < 0.001 (significance thresholds throughout figures)
  • count ≈100 cells counted per group, 3 fields per sample, n = 3 (Sema5a immunofluorescence intensity)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study combines single-cell RNA sequencing (scRNA-seq) of mouse femoral bone with in vitro co-culture assays and in vivo genetic knockout models (OVX-induced osteoporosis) to characterize osteocyte subset heterogeneity. Primary group comparisons (sham vs. OVX; BHR-Ocys vs. non-BHR-Ocys; WT vs. conditional knockouts) were analyzed with two-tailed unpaired Student's t-tests for two-group contrasts and one-way/two-way ANOVA with Tukey-Kramer post-hoc testing for multi-group contrasts. Results were reported as mean ± SD with threshold-based p-value symbols (* p<0.05, ** p<0.01, *** p<0.001); exact p-values, effect sizes, and confidence intervals were not provided.

Replicationbiological Sample sizeSample sizes stated per figure legend (n=3 or n=6 per group for in vitro/in vivo assays; Western blot samples pooled from 10 mice each, 6 samples per group); no formal power calculation described Groupssham vs. OVX mice; BHR-Ocys vs. non-BHR-Ocys; WT vs. Sema5a-CKO (±OVX); WT vs. Plxna1-CKO (±OVX); Sema5a-treated vs. untreated BMDMs Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionTukey-Kramer post-hoc test following ANOVA (for multi-group in vivo comparisons); no correction stated for multiple independent t-tests across figures; no FDR method stated for scRNA-seq DEG analysis in provided text
Statistical tests used
Test Applied to n Assumptions
two-tailed unpaired Student's t-test Flow cytometry quantification of BHR-Ocy proportion in sham vs. OVX mice (Figure 2H) n=6 per group not stated
two-tailed unpaired Student's t-test Western blot quantification of RANKL, Spp1, Csf-1, and sclerostin in BHR-Ocys vs. non-BHR-Ocys (Figures 3B–E); each sample pooled from 10 mice, 6 samples per group n=6 per group not stated
two-tailed unpaired Student's t-test TRAP-positive multinucleated cell count in BHR-Ocy vs. non-BHR-Ocy co-culture with BMDMs (Figure 3H) n=3 per group not stated
two-tailed unpaired Student's t-test or ANOVA with post-hoc Tukey-Kramer test In vivo Sema5a-CKO experiments: Oc.S/BS, serum CTX-1, femur BMD, total BMD, BV/TV, Tb.N, Tb.Th, Tb.Sp, Ct.Th across WT/CKO × sham/OVX groups (Figures 4H–Q) n=6 per group not stated
Pearson correlation Heatmap of pairwise correlations among 8 osteocyte subclusters (Figure S1B) 701 cells total across subclusters not stated
Gene Ontology (GO) enrichment analysis Characteristic genes of BHR-Ocys vs. other osteocyte subsets (Figure 2E; Figure S3); method for significance not stated in provided text not stated
Approaches that could also have been used
  • Cell type proportions (ratio of BHR-Ocys among all osteocytes) were compared between sham and OVX groups using a two-tailed unpaired t-test on flow cytometry values
    Could also: Compositional data analysis methods such as Dirichlet regression (DirichletReg in R) or the scCODA framework could also be used for scRNA-seq-derived cell-type proportions — Cell-type proportions are compositional (they sum to 1 within each sample), which violates the independence assumption of a t-test on individual proportions; compositional methods account for this constraint and can model uncertainty in estimated proportions, particularly when total cell counts per sample are small
  • Single-cell RNA sequencing differential gene expression between osteocyte subclusters was used to define subcluster identities and BHR-Ocy marker genes; the specific DE test was not stated in the provided text
    Could also: Pseudobulk approaches (e.g., DESeq2 or edgeR on aggregated per-animal counts per cluster) could also be used alongside or instead of single-cell-level tests such as Wilcoxon rank-sum or MAST — Pseudobulk methods treat biological replicates (animals) rather than individual cells as the unit of replication, which better controls type-I error inflation from treating thousands of cells as independent observations; they are increasingly recommended for scRNA-seq differential expression when multiple biological samples are available
  • Multiple independent two-group Student's t-tests were applied across several figures and comparisons (Western blot, flow cytometry, TRAP counts) without a stated family-wise error correction
    Could also: A single linear model or one-way ANOVA framework with a Bonferroni or Benjamini-Hochberg FDR correction applied across the family of comparisons could also be used — When many independent t-tests are conducted within a study, the probability of at least one false positive increases; a pre-specified correction strategy applied to the full family of tests would provide an explicit bound on the overall error rate
  • Data dispersion was reported exclusively as mean ± SD, with n as small as 3 for some in vitro assays
    Could also: Individual data points plotted alongside the mean ± SD (e.g., dot plots or strip plots overlaid on bar or box plots) could also be used, as recommended by many journals — When n=3, the SD describes only three observations and the mean may not be representative; showing individual values allows readers to assess the full distribution, detect outliers, and evaluate whether the summary statistics are informative
  • Statistical significance was reported only as threshold-based symbols (* p<0.05, ** p<0.01, *** p<0.001) without exact p-values or effect size metrics
    Could also: Reporting exact p-values together with a standardized effect size (e.g., Cohen's d for t-tests, partial η² for ANOVA) and 95% confidence intervals could also be done — Exact p-values allow readers and meta-analysts to assess the strength of evidence on a continuous scale; effect sizes convey biological magnitude independently of sample size and are necessary to judge whether statistically significant differences are also practically meaningful
  • Western blot protein quantification compared BHR-Ocys vs. non-BHR-Ocys with an unpaired t-test; each Western blot sample consisted of cells pooled from 10 mice, with 6 pooled samples per group
    Could also: A non-parametric Mann-Whitney U test could also be considered for this comparison — With n=6 biological replicates per group and Western blot data that may not be normally distributed (due to pooling and gel-to-gel variability), a non-parametric alternative does not require the normality assumption of the t-test and can be more robust when the sample size is small
Software: CellChat · Monocle (pseudotime trajectory analysis) · scRNA-seq analysis platform (t-SNE/UMAP; clustering)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 63/100
partly built on non-reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41250977

Paper: Jiang et al. 2025, Adv Sci. "Identification and Mechanisms of Osteocyte Subsets Involved in the Pathological Progression of Osteoporosis." PMID 41250977 · PMCID PMC12850396 · DOI 10.1002/advs.202513427. Listed code: github.com/sqjin/CellChat (third-party tool). Listed data: GSE230665.

The decisive availability fact

  • Data Availability Statement (verbatim): "Research data are not shared."
  • The Methods give no accession (no GEO/SRA/GSA/OMIX) for the study's own mouse femur sham/OVX single-cell RNA-seq — the dataset that ALL the headline computational results derive from. CellRanger v7.0.1 → Seurat v4.2.0 → Harmony → DoubletFinder → CellChat v1.6.1 → Monocle2 were all run on that private scRNA-seq matrix.
  • The only public, obtainable dataset the paper uses is GSE230665 (human Agilent microarray, GPL10332; 12 postmenopausal-osteoporosis vs 3 healthy-control femur samples). It is used for a single validation claim.

In scope AND feasible (public data → attempted)

id result location pipeline
C1 Sema5a expression significantly higher in PMOP patients vs control Fig S12, text standard GEOquery+limma on GSE230665

This is the only pipeline-derived result reproducible from public data. We run the standard third-party pipeline (GEOquery to fetch GSE230665 + GPL10332 annotation; limma moderated t-test + Welch t-test on SEMA5A probes) — equivalent to what the authors describe ("data obtained from GSE230665").

In scope but NOT feasible — DATA NOT SHARED (not attempted, drop reason)

All central single-cell results depend on the authors' unshared mouse scRNA-seq raw data ("Research data are not shared"; no accession):

  • 24,851 cells / 23,491 genes after QC (<500 or >7500 genes, >10% mito, doublets)
  • 6 osteocyte subsets C1–C6; BHR-Ocy = Egfr+ Il1r1+ Sema5a+; 701 osteocytes
  • CellChat v1.6.1 cell–cell communication (Sema5a–Plxna1 axis, PI3K/AKT/Myc)
  • Monocle2 pseudotime trajectory of osteocyte subsets
  • BHR-Ocy proportion ↑ in OVX vs sham (Fig 2H)

Because the raw scRNA-seq is unobtainable, the CellChat tool (the listed repo) cannot be exercised on the paper's data. Reproducing CellChat would require a processed Seurat object with the authors' cell annotations, which is not shipped.

Out of scope (wet-lab / not a pipeline)

Sema5a-CKO and Plxna1-CKO mouse phenotyping (BV/TV, Tb.N/Th/Sp, Ct.Th; Fig 4), ImageJ histomorphometry, in-vitro osteoclast assays. Not attempted.

Expected outcome

partial — reproduce C1 (the one public-data claim) 1:1 in direction + significance; everything central is a data_restricted/data_unavailable non-attempt because the authors did not share the scRNA-seq data.

Figures / tables: Fig S12Fig 2Fig 3Fig 2H
C1
Reported
Sema5a significantly higher in PMOP patients than control (Fig S12, qualitative)
Reproduced
SEMA5A probe 21646 on GSE230665: mean OP=2638.99 vs CTRL=526.95 (higher in OP, ~5x); Welch t=4.374 p=0.00109; limma logFC=2.023 adj.P=0.01538
within tolerance
C2
Reported
24851 cells / 23491 genes after QC (mouse scRNA-seq)
Reproduced
not attempted
partial
C3
Reported
6 osteocyte subsets C1-C6; BHR-Ocy=Egfr+Il1r1+Sema5a+; 701 osteocytes
Reproduced
not attempted
partial
C4
Reported
CellChat Sema5a-Plxna1 signaling axis (PI3K/AKT/Myc)
Reproduced
not attempted
partial
C5
Reported
BHR-Ocy proportion increased in OVX vs sham (Fig 2H)
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 57/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is a partial reproduction driven by data withholding. The single peripheral public-data claim (C1, Sema5a higher in PMOP, Fig S12) reproduces 1:1 on GSE230665 (OP 2639 vs CTRL 527, p=0.00109, adj.P=0.0154), with no fabrication signal. The four central single-cell claims (24851 cells, 701 osteocytes, six osteocyte subsets, the CellChat Sema5a-Plxna1 axis, Fig 2H) are non-reproducible because the authors gave no accession and state 'Research data are not shared' — the problem sits on the authors'/data-availability side, not in a demonstrated computational error. Severity is indeterminate rather than severe: where testable the science held, but the headline results cannot be independently checked as published.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

95.7 k
tokens (I/O) · 4.9 M incl. cache
9 min
runtime · 0.01 CPU-h
2.5 GB
peak RAM
1
HPC jobs
hummel
machine