Deconvolution of the hematopoietic stem cell microenvironment reveals a high degree of specialization and conservation.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡Reported values were only indirectly comparable
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the data-provenance claim, partly for the rest. This is a re-analysis paper; assigned dataset GSE128423 (Baryawno 2019), samples GSM3674224-29. CLEAN 1:1 result: the reported dataset size 38,443 cells is reproduced EXACTLY -- 42,000 raw deposited barcodes (7000/sample) reduce to 38,443 after the conventional initial QC threshold (>=200 genes/cell), fully derivable from the public data, no fabrication signal. Using Seurat 5.3.0 (third-party tool, P16) on GSE128423 the endothelial (Pecam1/Cdh5/Egfl7) and mesenchymal LepR+/Cxcl12-high (Lepr/Cxcl12/Col1a1) niche compartments clearly separate and express the paper's reported markers (qualitative match). NOT attempted (the honest ~20%): the integrated EC=9,587/MSC=5,291 counts and the 14/11 subcluster structure -- these require SCTransform integration of three datasets (GSE128423+GSE108891+in-house SCP1747) plus IKAP and a divide-and-conquer Louvain at unspecified resolutions; cluster counts are highly version/parameter sensitive and need data beyond the assigned accession. All heavy compute ran on «our HPC»; raw data and conda env stay on «infra»; only small results copied here.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 60assessed: 2026-06-15 ⛓ b7faaebcd0cf
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan integrating multiple public and proprietary scRNA-seq datasets through a tailored bioinformatic pipeline resolve the full spectrum of cellular states and differentiation stages in the murine HSC-regulatory bone marrow microenvironment, and to what extent are these features conserved in the human bone marrow?
- ★ A customized bootstrapping/random-forest divide-and-conquer clustering pipeline integrating three scRNA-seq datasets robustly resolves cell states despite high cell-to-cell similarity within compartments method
- ★ 14 intermediate endothelial cell states/subclusters and 11 mesenchymal differentiation stages are identified in the bone marrow microenvironment finding
- ★ Integration of datasets yields deeper subpopulation characterization than any single dataset alone, recovering markers undetectable in individual datasets finding
- ★ Bone marrow endothelial and mesenchymal cells display a previously unrecognized higher level of functional specialization finding
- ★ Microenvironmental cellular features identified in mouse are substantially conserved in human bone marrow, suggesting shared layers of hematopoietic microenvironmental regulation between species finding
- ★ The resource provides the most comprehensive cell atlas to date of the murine HSC-regulatory bone marrow microenvironment resource
- Cluster signatures derived from the integration can annotate independent datasets (validated with Baccin et al. via SingleR) method
- Endothelial subclusters can be discriminated into arteries and sinusoids with both known and novel markers finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA sequencing (scRNA-seq) integration/clustering | mouse bone marrow microenvironment (endothelial and mesenchymal cells) | none | cell subcluster/state and differentiation stage identity, gene markers | — |
| scRNA-seq (in-house dataset generation) | mouse bone marrow stromal cells lacking hematopoietic markers | none | transcriptional profiles (13,402 cells) | — |
| scRNA-seq (pilot) | human bone marrow | none | transcriptional profiles integrated with mouse atlas to assess conservation | — |
| gene set enrichment / Gene Ontology, Reactome, KEGG analysis | mouse BM endothelial and mesenchymal subclusters | none | enriched pathways/functional annotation per cluster (top 50 markers) | — |
| automated annotation with SingleR | independent mouse BM dataset (Baccin et al., 2020) | none | validation of cluster signature transferability | SingleR |
| robustness/cluster validation (random-forest + bootstrapping) | integrated mouse endothelial and mesenchymal cells | none | recall per cell and #correct dominant-cluster metric | — |
- – 14 robust subclusters identified in the bone marrow endothelium 14 subclusters
- – 11 robust subpopulations/differentiation stages identified in the mesenchyme 11 subpopulations
- – For subclusters B3.4, A2.1, A2.6 (endothelium) and C2.1, C3 (mesenchyme), over 50% of markers could not be detected by each dataset separately >50% of markers
- – Only the Baryawno dataset alone robustly identified all subclusters except B3.4 in endothelium, but at lower resolution than the integrated dataset
- ▲ Arterial clusters showed high expression of Ly6a, Ly6c1, Igfbp3, Vim; sinusoids expressed Adamts5, Stab2, Il6st, Ubd
- ▲ Novel markers Igfbp7 and Ppia identified for arteries and Cd164/Blvrb for sinusoids via integrative differential expression
- – SingleR annotation discriminated most described cellular states in an independent dataset despite lower cell numbers
- – Only clusters B2 and D3 lacked cells from all three datasets; large entropy levels indicated good dataset mixing
- count 6626 cells (Tikhonova et al. (2019) dataset)
- count 38443 cells (Baryawno et al. (2019) dataset)
- count 13,402 cells (in-house dataset)
- count N = 9587 (endothelial cells labeled across datasets)
- count N = 5291 (mesenchymal cells labeled across datasets)
- count 14 (endothelial subclusters identified)
- count 11 (mesenchymal subpopulations/differentiation stages)
- other over 50% (percent of markers undetectable per single dataset for certain subclusters)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper describes a computational pipeline integrating three mouse bone marrow scRNA-seq datasets (two public, one in-house) to characterize endothelial and mesenchymal cell populations. The primary analytical strategy combined Louvain high-resolution clustering with a custom divide-and-conquer bootstrapping/random-forest framework to assess cluster robustness, yielding 14 endothelial and 11 mesenchymal subclusters. Cluster identity was annotated using differential expression marker analysis and gene set enrichment against GO, Reactome, and KEGG databases; an independent dataset annotated with SingleR served as external validation.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Louvain community detection (high-resolution clustering) | Initial and iterative sub-clustering of endothelial (N=9587) and mesenchymal (N=5291) integrated cells | 9587 endothelial cells; 5291 mesenchymal cells (post-QC, integrated) | not stated |
| Random forest classification + bootstrapping (robustness evaluation, adapted from Tasic et al. 2016) | Validation of each proposed subcluster across all levels of the divide-and-conquer hierarchy; recall-per-cell and #Correct-per-cluster metrics computed | null | not stated |
| Differential expression analysis (specific tool(s) not named in excerpt; described as 'several differential expression tools') | Identification of top marker genes per endothelial and mesenchymal subcluster (Tables S2, S3) | null | not stated |
| Gene set enrichment analysis (GSEA) against GO, Reactome, and KEGG databases | Functional annotation of each subcluster using top 50 markers per cluster (Table S4, Figure 4E) | null | not stated |
| SingleR automated cell-type annotation | External validation: cluster signatures projected onto the independent Baccin et al. 2020 dataset (Figure S6) | null | not stated |
| Shannon entropy analysis | Quantification of dataset-of-origin distribution within each final subcluster to assess integration bias (Table S1, Figure S5) | null | not stated |
-
Cluster robustness was assessed with a custom random-forest + bootstrapping framework adapted from Tasic et al. 2016, with two bespoke metrics (recall-per-cell and #Correct-per-cluster)↳ Could also: Consensus clustering methods such as SC3 (single-cell consensus clustering) or stability-based resampling as implemented in clusterboot (R) could also quantify subcluster robustness — These off-the-shelf consensus approaches provide standardized robustness scores with known statistical properties and would facilitate direct comparison across studies using the same benchmarking framework
-
Louvain community detection was used for graph-based clustering at multiple resolution levels↳ Could also: The Leiden algorithm (Traag et al. 2019, already cited) or HDBSCAN could also be applied, as both are designed for hierarchical or density-aware partitioning of high-dimensional single-cell data — Leiden guarantees well-connected communities (a known limitation of Louvain) and is increasingly the default in single-cell pipelines; reporting both would allow comparison of sensitivity to algorithm choice
-
Differential expression marker identification was performed with multiple unnamed tools, with concordance noted qualitatively↳ Could also: Explicitly named and benchmarked tools such as DESeq2 (pseudo-bulk), edgeR, or MAST (mixed model) could also be applied with stated multiple-testing corrections (e.g. Benjamini-Hochberg FDR) — Naming the DE tools and their FDR thresholds would allow readers to reproduce the marker lists and assess the stability of conclusions across methods
-
Dataset-of-origin bias within clusters was summarized using Shannon entropy averages per cluster (Table S1)↳ Could also: A formal mixing metric such as the Local Inverse Simpson's Index (LISI, as in Harmony) or kBET could also quantify integration quality per cell rather than per cluster — Cell-level mixing metrics provide a more granular and statistically grounded characterization of integration success and are now widely adopted as benchmarks in the scRNA-seq integration literature
-
External validation used SingleR reference-based label transfer onto the Baccin et al. dataset↳ Could also: Seurat label transfer, scANVI, or scArches (reference-mapping) could also project the integrated atlas onto the held-out dataset and provide uncertainty scores per cell — Uncertainty/confidence scores from these tools would allow quantification of how reliably each subcluster can be identified in independent data, supplementing the qualitative concordance shown in Figure S6
-
Integration of three datasets was performed by using Tikhonova as a reference anchor without specifying the correction algorithm applied to batch effects across studies↳ Could also: Harmony, Seurat CCA/RPCA, or scVI could also be applied as dedicated batch-correction/integration algorithms, with quantitative benchmarking of integration quality — Reporting which batch-correction method was used and providing integration-quality metrics (e.g. LISI) would help readers gauge the degree to which remaining dataset-of-origin effects influence the final cluster assignments
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35494238
Paper: Ye J, Calvo IA, Cenzano I, ... Prosper F, Gomez-Cabrero D. "Deconvolution of the hematopoietic stem cell microenvironment reveals a high degree of specialization and conservation." iScience 25(5):104225 (2022). DOI 10.1016/j.isci.2022.104225 · PMID 35494238 · PMCID PMC9046238.
What kind of paper
A re-analysis / meta-analysis paper. It does NOT generate the mouse data it analyses; it deconvolutes the bone-marrow-niche (BMN) microenvironment from existing public mouse scRNA-seq datasets plus an in-house human/mouse dataset:
- GSE128423 (Baryawno et al. 2019, Cell) — mouse BM stroma, samples GSM3674224–GSM3674229, reported as 38,443 cells. ← our assigned data
- GSE108891 (Tikhonova et al. 2019) — 6,626 cells.
- In-house (Single Cell Portal SCP1747) — 13,402 cells.
Code
- BRIEF "code" link =
github.com/satijalab/seurat— this is the generic third-party tool (Seurat) auto-harvested from Methods. Per BRIEF rule P16, applying this tool to the paper's data is an equally valid reproduction. - Authors' own code =
github.com/TranslationalBioinformaticsUnit/BMN_characterization(R-only:Clustering.R,Using_Bootstrapping_functions.R,singleR.R,Added_value1.R). README has no pinned versions, no expected cell/cluster counts, no runnable end-to-end entrypoint — preprocessing/integration is described in prose only.
Pipeline (from Methods)
R 3.6.3/4.0.3 · Seurat 4.0.0/3.2.3 · SCTransform · IKAP clustering ·
pairwise Seurat IntegrateData · PCA+UMAP · divide-and-conquer Louvain
(hierarchical) · bootstrap random-forest robustness · SingleR 1.4.1 ·
clusterProfiler 3.18.1 · CellRanger 6.0.1.
IN SCOPE (attempted — clear, low-hanging, pipeline/data-derived)
- C1 — Input cell count of GSE128423 (GSM3674224–29) = 38,443.
Directly checkable: download
GSE128423_RAW.tar, count cell barcodes in the six deposited filtered matrices, sum. Pure data-provenance check, no clustering ambiguity. This is the cleanest 1:1 data point. - C2 — Major BMN compartments recoverable by a standard Seurat pipeline. Run Seurat (P16 third-party tool) on the 6 GSE128423 homeostasis samples (QC → SCTransform → cluster) and confirm the endothelial and mesenchymal/MSC niche compartments separate out and express the paper's reported marker genes (EC: Pecam1/Cdh5; MSC: Lepr, Cxcl12). Qualitative + approximate cell counts.
OUT OF SCOPE (NOT attempted — the hard ~20%, with reason)
- Exact integrated counts 9,587 EC / 5,291 MSC: require 3-dataset SCTransform integration (GSE128423+GSE108891+in-house) — multi-dataset, not GSE128423 alone.
- 14 EC subclusters / 11 MSC subclusters: depend on IKAP + divide-and-conquer Louvain at unspecified resolutions; cluster counts are highly version/parameter sensitive — not a clean 1:1 target.
- Human analysis (907+658 EC, 249 MSC), conservation enrichment scores, bootstrap robustness, SingleR cross-species, GSEA GO terms: depend on in-house SCP1747 data and bespoke scripts — out of scope.
Possible-fabrication watch
None flagged a priori. C1 is the one number on our assigned dataset that is exactly derivable from the shipped public data; we report the measured value vs 38,443 and let the human auditor judge.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
For the only quantitatively comparable endpoint, the reported 38,443 cells of GSE128423 (GSM3674224-29) is reproduced exactly (42,000 raw barcodes − 3,557 with <200 genes), fully derivable from public data with no fabrication signal, and the EC (Cdh5/Pecam1/Egfl7) and LepR+/Cxcl12-high MSC compartments clearly re-separate with the paper's markers. The honest ~20% — integrated counts (9587 EC/5291 MSC) and the 14/11 subcluster structure — was not attempted because it requires SCTransform integration of three datasets (incl. in-house SCP1747) plus IKAP/Louvain at unspecified resolutions. This gap is a scope/methodology limitation on our side (and data availability for the in-house set), not an authors' defect; severity for what was compared is negligible, but the central 'high specialization' conclusion is only partially confirmed.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.