Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Deconvolution of the hematopoietic stem cell microenvironment reveals a high degree of specialization and conservation.

iScience · 2022
L1 60/100 PQI 79
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
60/100
Reproducibility score
0.8 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 21% of all assessed papers rank 918 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for the data-provenance claim, partly for the rest. This is a re-analysis paper; assigned dataset GSE128423 (Baryawno 2019), samples GSM3674224-29. CLEAN 1:1 result: the reported dataset size 38,443 cells is reproduced EXACTLY -- 42,000 raw deposited barcodes (7000/sample) reduce to 38,443 after the conventional initial QC threshold (>=200 genes/cell), fully derivable from the public data, no fabrication signal. Using Seurat 5.3.0 (third-party tool, P16) on GSE128423 the endothelial (Pecam1/Cdh5/Egfl7) and mesenchymal LepR+/Cxcl12-high (Lepr/Cxcl12/Col1a1) niche compartments clearly separate and express the paper's reported markers (qualitative match). NOT attempted (the honest ~20%): the integrated EC=9,587/MSC=5,291 counts and the 14/11 subcluster structure -- these require SCTransform integration of three datasets (GSE128423+GSE108891+in-house SCP1747) plus IKAP and a divide-and-conquer Louvain at unspecified resolutions; cluster counts are highly version/parameter sensitive and need data beyond the assigned accession. All heavy compute ran on «our HPC»; raw data and conda env stay on «infra»; only small results copied here.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 60
    assessed: 2026-06-15 ⛓ b7faaebcd0cf
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can integrating multiple public and proprietary scRNA-seq datasets through a tailored bioinformatic pipeline resolve the full spectrum of cellular states and differentiation stages in the murine HSC-regulatory bone marrow microenvironment, and to what extent are these features conserved in the human bone marrow?

Core claims
  • A customized bootstrapping/random-forest divide-and-conquer clustering pipeline integrating three scRNA-seq datasets robustly resolves cell states despite high cell-to-cell similarity within compartments method
  • 14 intermediate endothelial cell states/subclusters and 11 mesenchymal differentiation stages are identified in the bone marrow microenvironment finding
  • Integration of datasets yields deeper subpopulation characterization than any single dataset alone, recovering markers undetectable in individual datasets finding
  • Bone marrow endothelial and mesenchymal cells display a previously unrecognized higher level of functional specialization finding
  • Microenvironmental cellular features identified in mouse are substantially conserved in human bone marrow, suggesting shared layers of hematopoietic microenvironmental regulation between species finding
  • The resource provides the most comprehensive cell atlas to date of the murine HSC-regulatory bone marrow microenvironment resource
  • Cluster signatures derived from the integration can annotate independent datasets (validated with Baccin et al. via SingleR) method
  • Endothelial subclusters can be discriminated into arteries and sinusoids with both known and novel markers finding
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA sequencing (scRNA-seq) integration/clustering mouse bone marrow microenvironment (endothelial and mesenchymal cells) none cell subcluster/state and differentiation stage identity, gene markers
scRNA-seq (in-house dataset generation) mouse bone marrow stromal cells lacking hematopoietic markers none transcriptional profiles (13,402 cells)
scRNA-seq (pilot) human bone marrow none transcriptional profiles integrated with mouse atlas to assess conservation
gene set enrichment / Gene Ontology, Reactome, KEGG analysis mouse BM endothelial and mesenchymal subclusters none enriched pathways/functional annotation per cluster (top 50 markers)
automated annotation with SingleR independent mouse BM dataset (Baccin et al., 2020) none validation of cluster signature transferability SingleR
robustness/cluster validation (random-forest + bootstrapping) integrated mouse endothelial and mesenchymal cells none recall per cell and #correct dominant-cluster metric
Key results
  • 14 robust subclusters identified in the bone marrow endothelium 14 subclusters
  • 11 robust subpopulations/differentiation stages identified in the mesenchyme 11 subpopulations
  • For subclusters B3.4, A2.1, A2.6 (endothelium) and C2.1, C3 (mesenchyme), over 50% of markers could not be detected by each dataset separately >50% of markers
  • Only the Baryawno dataset alone robustly identified all subclusters except B3.4 in endothelium, but at lower resolution than the integrated dataset
  • Arterial clusters showed high expression of Ly6a, Ly6c1, Igfbp3, Vim; sinusoids expressed Adamts5, Stab2, Il6st, Ubd
  • Novel markers Igfbp7 and Ppia identified for arteries and Cd164/Blvrb for sinusoids via integrative differential expression
  • SingleR annotation discriminated most described cellular states in an independent dataset despite lower cell numbers
  • Only clusters B2 and D3 lacked cells from all three datasets; large entropy levels indicated good dataset mixing
Key statistics
  • count 6626 cells (Tikhonova et al. (2019) dataset)
  • count 38443 cells (Baryawno et al. (2019) dataset)
  • count 13,402 cells (in-house dataset)
  • count N = 9587 (endothelial cells labeled across datasets)
  • count N = 5291 (mesenchymal cells labeled across datasets)
  • count 14 (endothelial subclusters identified)
  • count 11 (mesenchymal subpopulations/differentiation stages)
  • other over 50% (percent of markers undetectable per single dataset for certain subclusters)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper describes a computational pipeline integrating three mouse bone marrow scRNA-seq datasets (two public, one in-house) to characterize endothelial and mesenchymal cell populations. The primary analytical strategy combined Louvain high-resolution clustering with a custom divide-and-conquer bootstrapping/random-forest framework to assess cluster robustness, yielding 14 endothelial and 11 mesenchymal subclusters. Cluster identity was annotated using differential expression marker analysis and gene set enrichment against GO, Reactome, and KEGG databases; an independent dataset annotated with SingleR served as external validation.

Replicationmixed Sample sizeTotal cells per source dataset stated (Tikhonova: 6626; Baryawno: 38443; In-house: 13402); post-QC labeled cells stated per compartment (EC: 9587; MSC: 5291); no formal power calculation described Groups14 endothelial subclusters vs. each other; 11 mesenchymal subclusters vs. each other; arterial vs. sinusoidal; mouse vs. human (cross-species inference); individual datasets vs. integrated dataset (false-negative/positive marker recovery) Pairingna Randomization/blindingnot stated Dispersionnone Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
Louvain community detection (high-resolution clustering) Initial and iterative sub-clustering of endothelial (N=9587) and mesenchymal (N=5291) integrated cells 9587 endothelial cells; 5291 mesenchymal cells (post-QC, integrated) not stated
Random forest classification + bootstrapping (robustness evaluation, adapted from Tasic et al. 2016) Validation of each proposed subcluster across all levels of the divide-and-conquer hierarchy; recall-per-cell and #Correct-per-cluster metrics computed null not stated
Differential expression analysis (specific tool(s) not named in excerpt; described as 'several differential expression tools') Identification of top marker genes per endothelial and mesenchymal subcluster (Tables S2, S3) null not stated
Gene set enrichment analysis (GSEA) against GO, Reactome, and KEGG databases Functional annotation of each subcluster using top 50 markers per cluster (Table S4, Figure 4E) null not stated
SingleR automated cell-type annotation External validation: cluster signatures projected onto the independent Baccin et al. 2020 dataset (Figure S6) null not stated
Shannon entropy analysis Quantification of dataset-of-origin distribution within each final subcluster to assess integration bias (Table S1, Figure S5) null not stated
Approaches that could also have been used
  • Cluster robustness was assessed with a custom random-forest + bootstrapping framework adapted from Tasic et al. 2016, with two bespoke metrics (recall-per-cell and #Correct-per-cluster)
    Could also: Consensus clustering methods such as SC3 (single-cell consensus clustering) or stability-based resampling as implemented in clusterboot (R) could also quantify subcluster robustness — These off-the-shelf consensus approaches provide standardized robustness scores with known statistical properties and would facilitate direct comparison across studies using the same benchmarking framework
  • Louvain community detection was used for graph-based clustering at multiple resolution levels
    Could also: The Leiden algorithm (Traag et al. 2019, already cited) or HDBSCAN could also be applied, as both are designed for hierarchical or density-aware partitioning of high-dimensional single-cell data — Leiden guarantees well-connected communities (a known limitation of Louvain) and is increasingly the default in single-cell pipelines; reporting both would allow comparison of sensitivity to algorithm choice
  • Differential expression marker identification was performed with multiple unnamed tools, with concordance noted qualitatively
    Could also: Explicitly named and benchmarked tools such as DESeq2 (pseudo-bulk), edgeR, or MAST (mixed model) could also be applied with stated multiple-testing corrections (e.g. Benjamini-Hochberg FDR) — Naming the DE tools and their FDR thresholds would allow readers to reproduce the marker lists and assess the stability of conclusions across methods
  • Dataset-of-origin bias within clusters was summarized using Shannon entropy averages per cluster (Table S1)
    Could also: A formal mixing metric such as the Local Inverse Simpson's Index (LISI, as in Harmony) or kBET could also quantify integration quality per cell rather than per cluster — Cell-level mixing metrics provide a more granular and statistically grounded characterization of integration success and are now widely adopted as benchmarks in the scRNA-seq integration literature
  • External validation used SingleR reference-based label transfer onto the Baccin et al. dataset
    Could also: Seurat label transfer, scANVI, or scArches (reference-mapping) could also project the integrated atlas onto the held-out dataset and provide uncertainty scores per cell — Uncertainty/confidence scores from these tools would allow quantification of how reliably each subcluster can be identified in independent data, supplementing the qualitative concordance shown in Figure S6
  • Integration of three datasets was performed by using Tikhonova as a reference anchor without specifying the correction algorithm applied to batch effects across studies
    Could also: Harmony, Seurat CCA/RPCA, or scVI could also be applied as dedicated batch-correction/integration algorithms, with quantitative benchmarking of integration quality — Reporting which batch-correction method was used and providing integration-quality metrics (e.g. LISI) would help readers gauge the degree to which remaining dataset-of-origin effects influence the final cluster assignments
Software: Seurat · SingleR · Louvain/Traag implementation (igraph or leidenalg ecosystem) · Random forest classifier (package not specified)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
10
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE128423 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE108891 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM2915575 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM2915576 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM2915577 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM3674224 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM3674225 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM3674226 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM3674227 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM3674228 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM3674229 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35494238

Paper: Ye J, Calvo IA, Cenzano I, ... Prosper F, Gomez-Cabrero D. "Deconvolution of the hematopoietic stem cell microenvironment reveals a high degree of specialization and conservation." iScience 25(5):104225 (2022). DOI 10.1016/j.isci.2022.104225 · PMID 35494238 · PMCID PMC9046238.

What kind of paper

A re-analysis / meta-analysis paper. It does NOT generate the mouse data it analyses; it deconvolutes the bone-marrow-niche (BMN) microenvironment from existing public mouse scRNA-seq datasets plus an in-house human/mouse dataset:

  • GSE128423 (Baryawno et al. 2019, Cell) — mouse BM stroma, samples GSM3674224–GSM3674229, reported as 38,443 cells. ← our assigned data
  • GSE108891 (Tikhonova et al. 2019) — 6,626 cells.
  • In-house (Single Cell Portal SCP1747) — 13,402 cells.

Code

  • BRIEF "code" link = github.com/satijalab/seurat — this is the generic third-party tool (Seurat) auto-harvested from Methods. Per BRIEF rule P16, applying this tool to the paper's data is an equally valid reproduction.
  • Authors' own code = github.com/TranslationalBioinformaticsUnit/BMN_characterization (R-only: Clustering.R, Using_Bootstrapping_functions.R, singleR.R, Added_value1.R). README has no pinned versions, no expected cell/cluster counts, no runnable end-to-end entrypoint — preprocessing/integration is described in prose only.

Pipeline (from Methods)

R 3.6.3/4.0.3 · Seurat 4.0.0/3.2.3 · SCTransform · IKAP clustering · pairwise Seurat IntegrateData · PCA+UMAP · divide-and-conquer Louvain (hierarchical) · bootstrap random-forest robustness · SingleR 1.4.1 · clusterProfiler 3.18.1 · CellRanger 6.0.1.

IN SCOPE (attempted — clear, low-hanging, pipeline/data-derived)

  • C1 — Input cell count of GSE128423 (GSM3674224–29) = 38,443. Directly checkable: download GSE128423_RAW.tar, count cell barcodes in the six deposited filtered matrices, sum. Pure data-provenance check, no clustering ambiguity. This is the cleanest 1:1 data point.
  • C2 — Major BMN compartments recoverable by a standard Seurat pipeline. Run Seurat (P16 third-party tool) on the 6 GSE128423 homeostasis samples (QC → SCTransform → cluster) and confirm the endothelial and mesenchymal/MSC niche compartments separate out and express the paper's reported marker genes (EC: Pecam1/Cdh5; MSC: Lepr, Cxcl12). Qualitative + approximate cell counts.

OUT OF SCOPE (NOT attempted — the hard ~20%, with reason)

  • Exact integrated counts 9,587 EC / 5,291 MSC: require 3-dataset SCTransform integration (GSE128423+GSE108891+in-house) — multi-dataset, not GSE128423 alone.
  • 14 EC subclusters / 11 MSC subclusters: depend on IKAP + divide-and-conquer Louvain at unspecified resolutions; cluster counts are highly version/parameter sensitive — not a clean 1:1 target.
  • Human analysis (907+658 EC, 249 MSC), conservation enrichment scores, bootstrap robustness, SingleR cross-species, GSEA GO terms: depend on in-house SCP1747 data and bespoke scripts — out of scope.

Possible-fabrication watch

None flagged a priori. C1 is the one number on our assigned dataset that is exactly derivable from the shipped public data; we report the measured value vs 38,443 and let the human auditor judge.

Figures / tables: Fig 1
C1
Reported
38443 cells (GSE128423 GSM3674224-29)
Reproduced
38443 after standard Seurat initial filter (min.cells=3, min.features=200); 42000 raw deposited barcodes (7000/sample)
exact
C2a
Reported
Endothelial markers Pecam1/CD31, Cdh5, Icam2, Egfl7 mark an EC population
Reproduced
EC markers jointly high in clusters 8/15/17/23/26 (cl8: Cdh5=2.20, Pecam1=1.89, Egfl7=2.99, Icam2=1.22)
partial
C2b
Reported
Mesenchymal/MSC markers Cxcl12, Lepr, Angpt1, Col1a1 mark an MSC population
Reproduced
MSC markers jointly high in clusters 1/26/28 (cl1: Lepr=1.47, Cxcl12=5.92, Angpt1=0.44 = LepR+ MSC)
partial
C3
Reported
Integrated EC=9587 / MSC=5291
Reproduced
NOT ATTEMPTED (out of scope: 3-dataset SCTransform integration)
partial
C4
Reported
14 EC subclusters / 11 MSC subclusters
Reproduced
NOT ATTEMPTED (out of scope: IKAP + divide-and-conquer Louvain, version/parameter sensitive)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 60/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

For the only quantitatively comparable endpoint, the reported 38,443 cells of GSE128423 (GSM3674224-29) is reproduced exactly (42,000 raw barcodes − 3,557 with <200 genes), fully derivable from public data with no fabrication signal, and the EC (Cdh5/Pecam1/Egfl7) and LepR+/Cxcl12-high MSC compartments clearly re-separate with the paper's markers. The honest ~20% — integrated counts (9587 EC/5291 MSC) and the 14/11 subcluster structure — was not attempted because it requires SCTransform integration of three datasets (incl. in-house SCP1747) plus IKAP/Louvain at unspecified resolutions. This gap is a scope/methodology limitation on our side (and data availability for the in-house set), not an authors' defect; severity for what was compared is negligible, but the central 'high specialization' conclusion is only partially confirmed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

152.6 k
tokens (I/O) · 9.2 M incl. cache
20 min
runtime · 0.02 CPU-h
4.9 GB
peak RAM
2
HPC jobs
hummel
machine