Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

An integrated single-cell and spatial proteotranscriptomics atlas of fibroblast-driven immunoregulation within the human adult oral cavity.

Cell Press Blue · 2026
85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1, within tolerance). Scope: the atlas paper's registry 'code' is the third-party Wellcome-Sanger cellgeni/reprocess_public_10x STARsolo pipeline (BRIEF P2: third-party tool on the paper's data = equally valid) and its registry 'data' GSE217720 is a public scRNA-seq dataset (Ko et al., JEM 2022) the atlas reanalyses — NOT the atlas's own ~246k-cell deposition. Per 80/20 I took the cleanest reprocessable target: the smallest human sample GSM6725455 (SRR22254651+52, 97.41M read pairs). Ran STARsolo verbatim (STAR 2.7.10a, 2020-A GRCh38, auto-chemistry -> v3, Forward strand, EmptyDrops_CR) fully on «our HPC»/«infra» («job», completed): env build + 11.4GB ref + 30GB STAR index + fastq download + auto-chemistry/strand detection + final solo run. Result: 906 STARsolo cells vs 893 deposited = +1.46% — excellent cross-caller agreement (STARsolo EmptyDrops_CR re-derived from raw reads vs authors' deposited CellRanger matrix; kartei expectation 10-20%). All QC concordant (97.6% valid barcodes, 93.8% genome-mapped, median 2386 GeneFull/cell, seq sat 0.869). No value fabricated; the reproduced count is re-derived end-to-end from raw reads. NOT attempted (out of scope / 80-20 floor): atlas's own 246k-cell integration + spatial proteomics + fibroblast niches (authors' bespoke downstream, raw data in CELLxGENE/HCA not GSE217720); plus the heavier GSM6725456 human rep2 and GSM6725454 mouse samples (available, documented as bonus-not-done).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 85
    assessed: 2026-06-19 ⛓ 04484cfd1593
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether fibroblasts orchestrate conserved immune-signaling strategies across distinct oral mucosae and glands, with niche-specific subtypes encoding basal, para-inflammatory, pathologic, and/or immunoregulatory roles, and how this fibroblast-anchored architecture is remodeled in chronic oral disease.

Core claims
  • Fibroblasts act as central regulators of structural immunity in the human oral cavity, forming peri-epithelial hubs enriched in effector cytokines finding
  • Eight harmonized fibroblast subtypes exist across oral tissues: universal, immune/inflammatory, peri-epithelial, peri-vascular, peri-neural, APC-like, stress-responsive, and myofibroblast finding
  • Stress-responsive fibroblasts partition into two transcriptionally distinct groups: type I (enriched in salivary glands) and type II (enriched in mucosae) finding
  • Mucosal stress-responsive fibroblasts are putative immunoregulatory hubs mapped via spatial multiomics ligand-receptor programs finding
  • In chronic periodontitis, niche-aware integration reveals fibroblast rewiring into inflammatory and reparative niches with expansion of MHC-I+, MHC-II+, and PD-L1+ fibroblasts near tertiary lymphoid structures finding
  • AstroSuite, an AI-enabled spatial analysis toolkit (hist2omics, TACIT, Constellation, STARComm, Astrograph), was developed to define neighborhoods and interaction modules from spatial multiomics data method
  • Fibroblasts and vascular endothelial cells display the highest ligand-receptor interaction potential among structural cell types, with immune cells acting predominantly as receivers finding
  • Harmony outperformed four other batch-correction methods (Scanorama, FastMNN, scANVI, Seurat) for integrating the scRNA-seq atlas method
Experimental setups
Assay System Perturbation Readout Platform
scRNA-seq human oral mucosa, salivary gland, and dental pulp (11 public + 3 new datasets) none (healthy adult tissue) cell-type annotation, clustering, niche-specific heterogeneity
spatial proteomics (40-plex) + same-section spatial transcriptomics (300-plex) human FFPE oral tissue: tongue, gingiva, buccal mucosa, parotid, submandibular, minor salivary gland none cell type/neighborhood identification, ligand-receptor spatial mapping PhenoCycler-Fusion + Xenium
spatial proteomics (40-plex) + sequential-section spatial transcriptomics (300-plex) human FFPE oral tissue, same six niches none cell type/neighborhood identification, ligand-receptor spatial mapping PhenoCycler-Fusion + MERSCOPE
ligand-receptor communication modeling (MultiNicheNet) scRNA-seq atlas, 8 glandular and mucosal niches none recurrent structural-to-immune signaling programs across samples MultiNicheNet
ligand-receptor interaction quantification (CellPhoneDB) integrated scRNA-seq atlas, structural cell types none overall L-R interaction potential per cell type CellPhoneDB
ligand-receptor pathway analysis (CellChat) integrated scRNA-seq atlas, epithelia/fibroblasts/vasculature none shared ECM and immunomodulatory signaling circuits CellChat
whole-slide image cell segmentation spatial-proteomics tissue sections (six oral niches) none single-cell segmentation for downstream annotation Cellpose3, H&E-guided human-in-the-loop training
drug-target mapping (Drug2Cell) disease vs. healthy fibroblast/immune interfaces chronic periodontitis (disease) vs. healthy context-specific therapeutic target identification Drug2Cell
Key results
  • Final integrated scRNA-seq atlas comprised 246,102 cells from 70 samples spanning 13 macroniches
  • Harmony achieved the best batch-mixing/biological-conservation balance among five tested methods score 0.82
  • >25,000 fibroblasts from 10 niches subclustered into eight transcriptionally distinct subtypes
  • Fibroblasts and VECs ranked highest in L-R interaction potential; immune cells were predominantly receivers
  • Macrophages were significantly more abundant in mucosae than glands p = 0.012
  • CD8+ T cells were significantly enriched in glands relative to mucosae p = 0.006
  • Eight tissue cellular neighborhoods (TCNs A-H) identified spatially; mucosae exhibited all eight while glands harbored only five (A, C-F)
  • Spatial proteomics runs (36 tissues) yielded ~2 million segmented cells for neighborhood and cell-state annotation
Key statistics
  • count 246,102 cells; 70 samples; 13 macroniches (final integrated scRNA-seq atlas after QC)
  • pvalue p = 0.012 (macrophage proportion higher in mucosae vs. glands)
  • pvalue p = 0.006 (CD8+ T cell proportion higher in glands vs. mucosae)
  • count >25,000 fibroblasts from 10 niches (fibroblast subclustering into 8 subtypes)
  • count ~2 million segmented cells across 36 tissues (combined spatial-proteomics runs)
  • other batch-correction score 0.82 (Harmony scIB performance benchmark)
  • count >250,000 single-cell transcriptomes and >4 million spatially resolved cells across 13 niches (overall atlas scale (summary))
  • count 21 unique human samples, ~4 million cells, 6 niches (spatial proteotranscriptomics atlases combined)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper presents an integrated single-cell and spatial proteotranscriptomic atlas of human oral and craniofacial tissues, combining scRNA-seq from >246,000 cells (70 samples, 13 niches) with spatial proteomics (PhenoCycler-Fusion) and spatial transcriptomics (Xenium/MERSCOPE) covering ~4 million cells across 6 niches. Batch correction was benchmarked across five methods using scIB composite metrics, with Harmony selected. Cell-cell communication was characterized using MultiNicheNet, CellPhoneDB, and CellChat, while spatial neighborhoods were identified via a custom AstroSuite pipeline; two nominal p-values are reported for cell-type proportion comparisons between tissue types without identification of the underlying statistical test.

Replicationbiological Sample sizescRNA-seq: 70 samples, 13 macroniches, 246,102 cells after QC; spatial multiomics: 21 unique human samples, 6 niches, ~4 million cells total; spatial proteomics runs: 36 tissue sections, ~2 million segmented cells GroupsOral mucosae (tongue, gingiva, buccal) vs. salivary glands (parotid, submandibular, minor); healthy oral tissues vs. chronic periodontitis for niche-aware disease integration Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
scIB composite score (aggregating batch mixing, biological conservation, and cluster separation sub-metrics) Benchmarking of five batch-correction methods (Harmony, Scanorama, FastMNN, scANVI, Seurat; Figures S1A–S1C) 246,102 cells from 70 samples not stated
Statistical test not specified (p = 0.012) Comparison of macrophage (innate immune cell) proportions between mucosae and glands (Figure S3G) ~2 million segmented cells from 36 spatial-proteomics tissue sections; per-donor n not stated not stated
Statistical test not specified (p = 0.006) Comparison of CD8+ T cell (adaptive immune cell) proportions between mucosae and glands (Figure S3G) ~2 million segmented cells from 36 spatial-proteomics tissue sections; per-donor n not stated not stated
Co-occurrence analysis (specific method not named) Spatial proximity of immune cells to structural cells, mucosae vs. glands (Figure S3H) not stated
MultiNicheNet (adapted) Ligand-receptor signaling modeled across eight glandular and mucosal niches from pooled scRNA-seq samples (Figures 2A–2B, Table S2) Multiple donors per tissue group; exact per-group n not stated in available text not stated
CellPhoneDB Quantification of overall L-R interaction potential across structural and immune cell types (Figure 2C, Table S3) not stated
CellChat Identification of shared signaling circuits among epithelia, fibroblasts, and vasculature (Figure 2E) not stated
Approaches that could also have been used
  • Two p-values are reported for cell-type proportion comparisons between mucosae and glands, but the underlying statistical test is not named.
    Could also: Explicitly name and report the test (e.g., Wilcoxon rank-sum, permutation test, or a compositional model such as scCODA or DirichletReg) along with its assumptions. — Cell-type proportions per donor are compositional (sum to 1) and may be overdispersed; naming the test lets readers assess whether it accounts for these properties and for donor-level correlation.
  • Multiple cell-type proportion comparisons between tissue types are reported, but no multiplicity correction is stated.
    Could also: Apply Benjamini-Hochberg FDR correction across all cell-type proportion tests performed simultaneously. — When multiple cell types are tested, FDR control provides a standard framework for communicating the expected proportion of false discoveries among the reported findings.
  • Batch correction was selected by comparing five methods on an aggregate scIB composite score (Harmony, 0.82).
    Could also: Report individual scIB sub-metric scores (batch mixing, biological conservation, cluster separation) for each of the five methods alongside the composite. — Sub-metric breakdowns can reveal trade-offs — for instance, a method that excels at batch removal may sacrifice biological signal — that a single aggregate value may not surface.
  • Cell-cell communication was inferred using three separate tools (MultiNicheNet, CellPhoneDB, CellChat), each with its own scoring method and null model.
    Could also: A meta-aggregation framework such as LIANA could be used to harmonize and rank L-R interactions across multiple tools within a single, consensus-scored output. — Consensus aggregation assigns ranks that reflect agreement across tools with different assumptions, and can reduce the interpretive complexity of reconciling three separate result sets.
  • Spatial tissue cellular neighborhoods (TCNs) are described and compared between mucosae and glands in terms of enriched cell types and hub connectivity, but no formal enrichment statistics or effect sizes are reported.
    Could also: Quantify TCN differences using permutation-based enrichment scores (e.g., from Squidpy, spatialDE, or Banksy) with accompanying effect sizes or confidence intervals. — Formal statistics and effect sizes allow readers to assess the magnitude and reproducibility of neighborhood differences beyond visual or qualitative description.
  • The dataset is noted to be skewed toward younger, female donors with limited racial and ethnic diversity, but donor demographic variables are described rather than included as covariates in the statistical models.
    Could also: Include donor age and sex as covariates in differential expression or L-R models (e.g., within MultiNicheNet's mixed-effects framework), or present sensitivity analyses stratified by these variables. — Modeling known demographic covariates helps distinguish niche-driven biological variation from donor-level variation attributable to age or sex, and can strengthen confidence in findings that survive covariate adjustment.
Software: Harmony · Seurat · scANVI · FastMNN · Scanorama · MultiNicheNet · CellPhoneDB · CellChat · AstroSuite (custom suite: TACIT, Constellation, STARComm, Astrograph) · Cellpose3 · Drug2Cell

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-42147490

Paper: Matuck BF et al. (2026) An integrated single-cell and spatial proteotranscriptomics atlas of fibroblast-driven immunoregulation within the human adult oral cavity. Cell Press Blue. PMID 42147490 / PMCID PMC13179517 / DOI 10.1016/j.cpblue.2026.100007.

Registry code artifact: https://github.com/cellgeni/reprocess_public_10x (Wellcome Sanger / cellgeni STARsolo reprocessing pipeline — a third-party tool, not the authors' own analysis code; per BRIEF P2 this is an equally valid reproduction target: apply the tool to the paper's data with the described parameters).

Registry data accession: GEO GSE217720.

What GSE217720 actually is (verified via GEO/ENA, 2026-06-14)

GSE217720 is not the atlas paper's own deposition (that atlas is deposited in CELLxGENE / HCA, ~246k cells, 70 samples + spatial). GSE217720 is a public 3rd-party scRNA-seq dataset — Ko et al., "Distinct Fibroblast Progenitor Subpopulation Expedites Regenerative Mucosal Healing by Immunomodulation" (JEM 2022, doi:10.1084/jem.20221350) — that the oral-cavity atlas reanalyses/integrates as one of its public input datasets. It is small and self-contained:

GSM title organism SRX ENA runs size
GSM6725454 mouse_palate rep1 (CD45-) mouse SRX18231179 SRR22254653 (1 run, 144M reads, interleaved) ~10 GB
GSM6725455 human_palatal_ruga rep1 (all cells) human SRX18231180 SRR22254651+SRR22254652 (2 runs, ~97M pairs) ~6 GB
GSM6725456 human_palatal_ruga rep2 (all cells) human SRX18231181 4 runs incl. 2× 760M-read (~100 GB) ~110 GB

Each GSM ships the authors' own deposited count matrix (*_barcodes.tsv.gz / *_features.tsv.gz / *_matrix.mtx.gz) = ground-truth cell calls to compare against.

IN SCOPE (pipeline-derived, attempted)

Apply the cellgeni reprocess_public_10x STARsolo workflow (STAR 2.7.10a, CellRanger 2020-A GRCh38 reference, auto-chemistry whitelist detection, --soloCellFilter EmptyDrops_CR, exact flags copied verbatim from the repo's scripts/starsolo_10x_auto.sh) to GSM6725455 (human_palatal_ruga rep1), the smallest human sample, and compare the pipeline's per-sample QC metrics (STARsolo Estimated Number of Cells, Median GeneFull per Cell, reads mapped) 1:1 against the deposited GEO matrix for the same sample (cells = #barcodes in the deposited filtered matrix). This is a genuine cross-tool reprocessing reproduction: does the STARsolo tool recover a comparable cell population to the originally deposited (CellRanger-based) matrix?

Clear data points (few, honest): #cells (STARsolo EmptyDrops_CR vs deposited matrix), median genes/cell, reads-mapped fraction, detected 10x chemistry.

OUT OF SCOPE (not attempted — and why)

  • The atlas's own 246,102-cell integration, spatial proteomics (~2–4M cells), fibroblast subclustering, niche/macroniche definitions, CellxGENE deposition — these are the authors' bespoke downstream analyses; the registry code artifact is only the upstream reprocessing tool, and the atlas raw data is in HCA, not in the registry's GSE217720. Reproducing the full atlas is far beyond the 80/20 low-hanging target and not what the cited pipeline produces.
  • GSM6725456 (human rep2): two 760M-read runs (~110 GB) — deliberately skipped under the 80/20 rule (cost ≫ value of a 2nd human replicate). Noted, not hidden.
  • GSM6725454 (mouse): different reference; optional stretch only.

Comparison target provenance

"Reported value" for each claim = a property of the deposited GSE217720 matrix (authors' processed output) and/or the STARsolo run's own Summary.csv — both are machine-derived from shipped data, so any mismatch is an honest tool-vs-tool delta, not a fabrication signal. Flagged explicitly in AUDIT.md.

C1
Reported
893 cells in the deposited GEO filtered matrix for GSM6725455 (human palatal ruga rep1) of GSE217720 — the public scRNA-seq input (Ko et al., JEM 2022) reanalysed by the atlas paper. Ground-truth verified: zcat barcodes.tsv.gz | wc -l = 893.
Reproduced
906 cells. Re-ran the third-party cellgeni/reprocess_public_10x STARsolo pipeline (STAR 2.7.10a, verbatim flags, 10x 2020-A GRCh38 ref, auto-detected v3 chemistry, EmptyDrops_CR) end-to-end on the raw ENA fastqs (SRR22254651+SRR22254652, 97.41M read pairs). STARsolo GeneFull 'Estimated Number of Cells' = 906; filtered/barcodes.tsv = 906; Gene feature = 895. +13 cells / +1.46% vs deposited.
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

371.6 k
tokens (I/O) · 26.6 M incl. cache
285 min
runtime · 2.72 CPU-h
49.7 GB
peak RAM
1
HPC jobs
hummel
machine