comBO: A combined human bone and lympho-myeloid bone marrow organoid for preclinical modeling of hematopoietic disorders.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
comBO (Cell Stem Cell 2026, GSE287648) — described well enough to test the deposited end-products 1:1. Pipeline (CellRanger->CellBender->Souporcell->Seurat->Harmony) is clearly stated; full re-run from FASTQ was deliberately NOT attempted (the hard ~20%: 12-sample alignment+ambient+demux, and CellBender's raw matrices are not deposited, only SRA FASTQ). Instead we counted cells in the shipped fully-annotated Harmony-integrated Seurat objects (readRDS+ncol on «our HPC»/«infra») and compared to the two reported integrated dataset sizes. RESULT: the Fig 4 multiple-myeloma dataset reproduces EXACTLY to the cell (45,706 haematopoietic + 14,577 stromal). The Fig 1E characterization dataset (reported 29,837 haem / 15,837 stroma) does NOT match any deposited object: closest day20+d35 haematopoietic object = 25,356 (delta -4,481), closest stromal object = 15,092 (delta -745). Likely a re-processed/re-filtered deposit ('-v2' objects) post-dating the Fig 1E count rather than fabrication; flagged for human audit. NOT attempted: FASTQ->counts pipeline, clustering/annotation re-derivation, CellChat MIF networks, miloR differential abundance, all wet-lab claims. Verdict is provisional and must be independently checked.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 65assessed: 2026-06-14 ⛓ c0705eeeb902
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether a scalable induced pluripotent stem cell (iPSC)-derived bone marrow organoid (comBO) can be engineered to simultaneously capture lympho-myeloid hematopoiesis and multi-lineage stromal diversity (osteogenic, vascular, adipogenic), and whether such a model can recapitulate multiple myeloma-driven niche remodeling to identify therapeutically targetable disease mechanisms.
- ★ comBO is a single iPSC-derived organoid differentiation that generates osteolineage, vascular, lymphoid, and myeloid compartments together finding
- ★ Granular microgel scaffolds replace bulk hydrogel embedding, improving scalability and reproducibility of comBO production method
- ★ comBO hematopoietic and stromal cells show transcriptional homology to adult human bone marrow, higher than to fetal bone marrow or prior organoid models finding
- ★ comBO sustains long-term lympho-myeloid HSPC potential across serial organoid re-seeding assays without exogenous cytokines finding
- ★ Myeloma-engrafted comBO chimeroids (MM-comBO) recapitulate patient-like niche remodeling, including stromal inflammation, metabolic reprogramming of MSCs, and arrested osteoblast maturation finding
- ★ Macrophage inhibitory factor (MIF) signaling between myeloma cells and organoid niche components is a central mediator of myeloma-driven inflammation mechanism
- ★ MIF inhibition reduces inflammatory responses and myeloma cell proliferation in MM-comBO finding
- comBO-derived CD7+ T-lineage progenitors can mature into CD3+TCRαβ+ CD4/CD8 single-positive T cells when engrafted into an artificial thymus organoid finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| flow cytometry | hiPSC-derived comBO (days 21 and 35) | none (differentiation protocol) | hematopoietic, lymphoid, myeloid, and stromal lineage marker expression | — |
| single-cell RNA sequencing (scRNA-seq) | comBO organoids, days 20 and 35, pooled from 4 differentiations | none | cell-type identity, gene expression, lineage trajectory (Palantir pseudotime, Moscot) | — |
| 3D confocal imaging / immunofluorescence | comBO organoids | none | CD34+ vasculature, CD271+ perivascular cells, BODIPY+ adipocytes, osteocalcin/Osterix/CD7/TdT localization | confocal microscope |
| Von Kossa and alizarin red staining; H&E histology | comBO organoids | none | progressive mineralization / bone-like morphology | — |
| colony-forming unit (CFU) assay | comBO-derived CD34+ cells vs. peripheral blood CD34+ controls | none | clonogenic potential | — |
| serial organoid re-seeding assay | primary healthy-donor CD34+ cells and comBO-derived (mScarlet-tagged) CD34+ cells seeded into comBO | serial re-seeding across 3 passages, growth-factor-free media | CD34+ cell fold change and lympho-myeloid differentiation over 8 weeks | flow cytometry |
| artificial thymus organoid (ATO) engraftment | comBO-derived CD7+ T-lineage progenitors | engraftment into ATO | CD3, TCRα/β expression and CD4+/CD8+ single-positive cell emergence | flow cytometry |
| patient cell engraftment (chimeroid) with scRNA-seq, flow cytometry, and imaging | comBO engrafted with CD138+ cells from 3 high-risk multiple myeloma patients (MM-comBO) vs. un-engrafted and healthy CD34+-engrafted controls | myeloma cell engraftment | myeloma cell proliferation/persistence, niche cell differential abundance (MiloR), GSEA/DEG stromal and hematopoietic remodeling | — |
- – Single comBO differentiation generated HSPCs, myeloid cells (erythroid, megakaryocyte, monocyte, EBM), lymphoid cells (T-lineage progenitors, NK cells), MSCs, osteolineage, and endothelial cells by day 21
- – scRNA-seq of comBO yielded 29,837 hematopoietic and 15,837 stromal cells, annotated into adult bone marrow-like cell types 29,837 hematopoietic + 15,837 stromal cells
- ▲ Day 35 comBO data integrated with adult bone marrow atlas showed consistently higher transcriptional homology to adult vs. fetal bone marrow atlas data and other organoid models
- – Granular microgel scaffolds supported vascular sprouting and equivalent lineage/architectural complexity to bulk hydrogels while using far less material 80% reduction in bulk gel volume
- – CD34+ cell numbers increased in all 3 healthy donors after first and second re-seedings; 2 of 3 continued expanding after a third re-seeding, one declined 1st seeding: 1.26x/1.21x/1.64x; 2nd: 2.82x/5.35x/7.52x; 3rd: 1.19x/5.35x/-0.49x (decrease)
- – MM-comBO scRNA-seq identified 45,706 hematopoietic and 14,577 stromal cells including a myeloma-specific plasma cell cluster 45,706 hematopoietic + 14,577 stromal cells
- ▲ Differential abundance analysis (MiloR) showed significant enrichment of lymphoid progenitors, T cells, mast cells, osteolineage cells, adipocytes, and endothelium in MM-comBO vs. controls
- ▲ DEG analysis showed upregulation of stromal activation, angiogenic, adipogenic, and inflammatory genes (e.g., NFKBIA, CCL2, JUNB, SOCS3) in MM-comBO stromal populations
- count 29,837 hematopoietic cells and 15,837 stromal cells (scRNA-seq of comBO pooled from 4 independent differentiations (days 20 and 35))
- fold_change HD1 1.26x, HD2 1.21x, HD3 1.64x (CD34+ cell fold increase after first re-seeding of healthy donor cells)
- fold_change HD1 2.82x, HD2 5.35x, HD3 7.52x (CD34+ cell fold increase after second (secondary) re-seeding)
- fold_change HD1 1.19x increase, HD2 5.35x increase, HD3 0.49x decrease (CD34+ cell fold change after third (tertiary) re-seeding)
- count 45,706 hematopoietic cells and 14,577 stromal cells (scRNA-seq of MM-comBO chimeroids engrafted with patient myeloma cells)
- other 80% reduction in bulk gel volume (material savings using granular microgel vs. bulk hydrogel scaffold)
- other 5 ng/mL EPO and 1 ng/mL TPO (minimal cytokine supplementation supporting myeloma cell expansion to day 28 post-engraftment)
- count 3 patients with high-risk multiple myeloma (CD138+ bone marrow aspirate donors used to generate MM-comBO chimeroids)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
comBO is characterized through flow cytometry, confocal imaging, and single-cell RNA sequencing (scRNA-seq) across multiple independent differentiations and donor samples. Transcriptional homology to adult bone marrow was assessed by atlas integration, and lineage trajectories were inferred via Palantir pseudotime and Moscot-based optimal transport. Myeloma-induced remodeling of the niche was quantified by differential abundance analysis (MiloR), differentially expressed gene (DEG) analysis, and gene set enrichment analysis (GSEA). Functional endpoints such as serial re-seeding capacity are reported as per-donor fold-changes without aggregate summary statistics or formal inferential testing; the full statistical methods section was not available in the provided text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential abundance analysis (MiloR; KNN-graph neighbourhood-based) | Cell-type abundance differences between MM-comBO and un-engrafted / healthy-donor-engrafted control organoids (Figures 4I, 4J) | 45,706 hematopoietic cells and 14,577 stromal cells across MM-comBO and control scRNA-seq datasets | not stated |
| Gene set enrichment analysis (GSEA) | Hallmark gene set enrichment in MM-comBO versus controls across cell clusters (Figures 5A, 5B) | — | not stated |
| Differentially expressed gene (DEG) analysis | Upregulated genes in CD271+ MSC, NES+ MSC, osteolineage 1, osteoclasts, and lymphoid progenitors in MM-comBO versus controls (Figure 5C) | — | not stated |
| Palantir pseudotime / diffusion-map trajectory inference | HSC/MPP differentiation trajectories toward myeloid and lymphoid lineages (Figure 1G[i]) | 29,837 hematopoietic cells pooled from 4 independent differentiations at days 20 and 35 | na |
| Moscot optimal-transport trajectory inference | Time-point-informed lineage trajectory sampling across day 20 and day 35 samples (Figure 1G[ii]) | Cells from 4 independent differentiations at days 20 and 35 | na |
| Cell-cycle phase scoring / classification | Cycling status (G1, S, G2M) of myeloma plasma cells in MM-comBO (Figure 4H[i, ii]) | — | na |
-
scRNA-seq data from 4 independent differentiations were pooled into a single dataset for DEG and cluster-level analysis↳ Could also: Pseudo-bulk aggregation per biological replicate (e.g., summing raw counts per donor per cluster) followed by DESeq2 or edgeR Wald/likelihood-ratio tests on pseudo-bulk counts could also be applied — Pseudo-bulk approaches treat the biological replicate as the unit of analysis, which accounts for within-replicate cell-cell correlation that standard single-cell DE methods assume away; this tends to reduce inflated false-positive rates when replicates are available
-
Differential cell-type abundance between myeloma-engrafted and control organoids was assessed with MiloR, a neighbourhood-graph-based method↳ Could also: scCODA (Bayesian Dirichlet-multinomial model) or Propeller (limma-based transformation of cell-type proportions) could also quantify differential abundance — scCODA explicitly models the compositional constraint of cell-type proportions and provides posterior credible intervals; Propeller sits within a familiar linear-model framework; convergence across methods can strengthen confidence in reported abundance shifts
-
Serial re-seeding HSPC maintenance was reported as individual per-donor fold-changes across three donors without an aggregate inferential test↳ Could also: A linear mixed-effects model with donor as a random effect and seeding round as a fixed effect, or a repeated-measures ANOVA, could also summarize the trend across rounds — Formal modeling would partition donor-level variability from the overall seeding effect and yield an estimate of consistency across donors, complementing the individual trajectories already shown
-
Pathway-level differences in MM-comBO were characterized with hallmark GSEA applied to cluster-level aggregate expression↳ Could also: Single-sample GSEA (ssGSEA) or AUCell could also score pathway activity at per-cell resolution within each cluster — Per-cell scoring allows statistical comparison of pathway activity distributions across conditions using standard tests and can reveal within-cluster heterogeneity that aggregate GSEA scores may smooth over
-
Transcriptional homology to adult bone marrow was assessed qualitatively by UMAP co-embedding with a published atlas↳ Could also: Quantitative label-transfer confidence scores (e.g., Seurat TransferData, scANVI posterior probabilities) or Pearson correlation of mean cluster gene-expression profiles could also measure homology numerically — Numeric similarity metrics provide a continuous, reproducible score that is less susceptible to UMAP projection distortion and facilitates direct comparison across organoid models or time points
-
Lineage trajectories were inferred using Palantir and Moscot, both of which rely on diffusion-based or optimal-transport frameworks↳ Could also: RNA velocity (scVelo or UniTVelo) or Monocle 3 (reversed graph embedding) could also infer differentiation directionality from spliced/unspliced transcript ratios or graph-based topology — RNA velocity uses a complementary source of signal (kinetic splicing rates) rather than transcriptional similarity alone; convergent results across mechanistically distinct methods increase confidence in the inferred developmental ordering
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-41734765 (comBO, Cell Stem Cell 2026)
Pipeline (Methods → "Analysis of single cell RNA sequencing")
Raw scRNA-seq (10x 3' v3.1) processed with: CellRanger v7.0.0 → CellBender v0.3.0 (ambient RNA) → Souporcell v3.0.0 (genotype demux) → Seurat v5.1.0 (QC/cluster/annotate) → Harmony v1.2.0 (integration) → CellChat v1.6.1 (LR), miloR v2.0 (differential abundance). Authors' scripts: https://github.com/khanaswimm/comBO.v1 Brief's "code" link (broadinstitute/CellBender) = ONE tool in this chain. Data: GEO GSE287648 — 12 scRNA-seq samples + fully annotated integrated R objects.
IN SCOPE (pipeline-derived, cheaply + faithfully checkable from SHIPPED data)
The series supplementary ships the END-PRODUCT annotated, Harmony-integrated R objects. The paper reports exact cell counts for the integrated objects. These are 1:1 verifiable by loading the object and counting cells:
- C1 florg hematopoietic = 29,837 cells (Fig 1E / text) -> florg_haem.rds
- C2 florg stromal = 15,837 cells (Fig 1E / text) -> florg_stroma.rds
- C3 MM hematopoietic = 45,706 cells (Fig 4D / text) -> MM_haem.rds
- C4 MM stromal = 14,577 cells (Fig 4E / text) -> MM_stroma.rds
These are the clean, low-hanging pipeline outputs (80%). Counting cells in the shipped integrated object directly reproduces the reported dataset sizes.
OUT OF SCOPE (the hard ~20%, not attempted — see AUDIT/ROOM_RESULT)
- Re-running CellRanger→CellBender→Souporcell from raw FASTQ (SRA PRJNA1214192): full 12-sample alignment+ambient-removal+demux = days of compute; the named CellBender step needs raw_feature_bc_matrix.h5 which is NOT shipped (per-sample supplementary = NONE; only FASTQ via SRA). Not the 80%.
- Re-deriving clustering/annotation, CellChat MIF networks, miloR DA, MIF-driver biology — multi-step, parameter-sensitive; wet-lab claims (flow, imaging) are out of scope by definition.
Approach
Heavy step = loading multi-GB integrated Seurat objects (needs RAM) -> «our HPC» SLURM job, data on «infra». Compare object cell counts to reported C1–C4.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
For the Fig 4 multiple-myeloma dataset the reported counts reproduce to the exact cell (45,706 haematopoietic + 14,577 stromal), confirming the deposited GEO objects are the real analysis end-products. The Fig 1E characterization counts do not match any deposited object: 29,837 haem vs a closest 25,356 (-15%) and 15,837 vs 15,092 stroma (-4.7%). The discrepancy sits at the input/filtering level and most plausibly reflects a re-processed '-v2' deposit post-dating the figure rather than fabrication, though we did not re-run the FASTQ→counts pipeline and chose the closest object ourselves. Overall a solid partial reproduction with an explainable, one-directional shortfall and no fabrication signal.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.