Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

SPRR1B+ keratinocytes prime oral mucosa for rapid wound healing via STAT3 activation.

Commun Biol · 2024
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to PARTIALLY reproduce the bulk-RNA-seq branch of Fig.1 using the listed data (GSE97615) and listed tool (clusterProfiler). 1:1 on the GO themes (C3) and ~1:1 on the DEG count (C1: 479 vs 485, -1.2%). The P2=146 sub-module (C2) did NOT reproduce: the paper's 'unsupervised clustering' of the 485 DEGs into P1/P2 is under-specified (no algorithm/distance/k), and a reasonable ward.D2 k=2 split gives 307/172. Important honesty flag: the paper attributes the 485 DEGs to an injured-vs-uninjured contrast, but GSE97615 contains ONLY post-wound samples (no uninjured baseline); the only computable contrast (oral vs skin) reproduces 479 and matches the paper's own description of P2 as constitutively expressed in oral but absent in skin — likely imprecise methods wording rather than a fabricated number, flagged for human review. NOT attempted (out of scope): all single-cell analyses (GSE164241 + PRJCA006797, Seurat/Harmony, 9 cell types, KC1-KC5), SCENIC TF regulons (STAT3/KLF5/PRDM1/GRHL1), STRING/Cytoscape/CytoHubba PPI hubs, Monocle2 pseudotime, and ChIP-seq (SRP070705, bowtie2/MACS2) — these use different data and different tools than the listed clusterProfiler, and are heavy.

💻 Code ↗ 🗄 Data: GSE97615

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 78
    assessed: 2026-06-14 ⛓ 0eb58393a3cb
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Oral mucosa heals faster and with less scarring than skin, and the paper tests whether a specific keratinocyte subpopulation, constitutively present in unwounded oral mucosa and induced via STAT3 activation, primes oral tissue for rapid wound healing.

Core claims
  • A shared wound healing gene set (P2-WHGs, 146 genes) is constitutively expressed in uninjured oral mucosa but not in uninjured skin finding
  • Keratinocytes, and oral keratinocytes in particular, are the major cell type expressing P2-WHGs among all sequenced skin/mucosa cell types finding
  • A keratinocyte subcluster, SPRR1B+ KC4, shows the highest P2-WHGs score and is more enriched in normal oral mucosa than in skin finding
  • STAT3 transcriptionally regulates SPRR1B by binding its promoter region, and JAK/STAT3 inhibition reduces SPRR1B protein levels mechanism
  • SPRR1B knockdown significantly inhibits mucosal keratinocyte migration finding
  • SPRR1B+ keratinocytes are induced during wound healing in both skin and oral mucosa in a murine model finding
  • Oral KC4 shows increased expression of inflammation (S100A8, S100A9, IL36G) and EMT (VIM, LUM, COL1A1) genes compared to skin KC4 finding
  • Oral mucosa contains multiple layers of KRT14+ basal keratinocytes compared to a single basal layer in skin finding
Experimental setups
Assay System Perturbation Readout Platform
bulk mRNA-seq / DEG analysis human injured/uninjured oral mucosa and skin (GSE97615) wound (injury) vs uninjured differentially expressed genes, gene panels P1/P2
single-cell RNA-seq human normal skin (PRJCA006797, 5 samples) and oral mucosa (GSE164241, 21 samples: gingiva/buccal) none cell type clustering, P2-WHGs enrichment score per cell/cluster
gene ontology / GSEA enrichment analysis P2-WHGs gene list and KC4 marker genes (bioinformatic) none enriched biological processes/pathways
SCENIC transcription factor regulon analysis oral KC4 keratinocytes (scRNA-seq derived) none TF regulon activity scores SCENIC
protein-protein interaction network analysis TFs identified in KC4 (bioinformatic) none hub TF identification (STAT3) STRING; CytoHubba MCC algorithm
ChIP-seq re-analysis / ChIP-PCR epithelial cells / HOK (human oral keratinocytes) none STAT3 binding to SPRR1B promoter (0-500 bp upstream of TSS)
Western blot HOK cells JAK1/2 inhibitor ruxolitinib STAT3, p-STAT3, JAK1, p-JAK1, SPRR1B protein levels
immunofluorescence staining human normal skin and oral mucosa biopsies none KRT14, KRT10, SPRR1B protein localization/expression
Key results
  • 485 DEGs identified as shared wound healing-associated gene sets between injured/uninjured skin and oral mucosa adjusted P<0.05, log2FC>1
  • 146 P2-WHGs genes constitutively expressed in uninjured oral mucosa but not skin
  • Keratinocytes show significantly higher P2 gene set scores than other cell types; oral keratinocytes score higher than skin keratinocytes despite lower abundance p<2.2e-16
  • SPRR1B+ KC4 shows the highest P2-WHGs score among 5 keratinocyte subclusters, higher in oral than skin KC4 p<2.2e-16
  • STAT3 identified as hub transcription factor in KC4 PPI network, connecting 6 nodes 6 nodes
  • ChIP-PCR confirms STAT3 binds directly to SPRR1B promoter region 0-500 bp upstream of TSS
  • Ruxolitinib (JAK1/2 inhibitor) decreases p-STAT3 levels and reduces SPRR1B protein in HOK cells
  • SPRR1B knockdown inhibits mucosal keratinocyte migration
Key statistics
  • count 485 (differentially expressed genes defining shared wound healing-associated gene sets)
  • count 146 (P2 wound healing genes (P2-WHGs) constitutively expressed in uninjured oral mucosa)
  • fold_change log2 fold change > 1 (DEG threshold for injured vs uninjured comparison)
  • pvalue adjusted P < 0.05 (DEG and GO enrichment significance threshold)
  • count 103,758 (cells retained after QC for scRNA-seq clustering of skin and oral mucosa)
  • pvalue p < 2.2e-16 (Wilcoxon rank-sum test, P2-WHGs score comparison between oral and skin keratinocytes)
  • count 9 (distinct cell types identified by scRNA-seq clustering)
  • count 5 (keratinocyte subtypes (KC1-KC5) identified by subclustering)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study integrated publicly available bulk RNA-seq (GSE97615) and single-cell RNA-seq data (103,758 cells from 5 skin and 21 oral mucosal samples) to characterize wound-healing gene programs across tissue types and keratinocyte subtypes. Bulk DEGs were called with fold-change and adjusted-P thresholds; single-cell group differences were assessed with two-sided Wilcoxon rank-sum tests; gene module enrichment was visualized on UMAP and tested with GSEA. Transcription-factor regulon activity was estimated computationally (SCENIC), and key findings were validated by ChIP-PCR and western blot in human oral keratinocyte (HOK) cells.

Replicationbiological Sample sizescRNA-seq: 103,758 cells after QC from 5 skin samples and 21 oral mucosal samples (13 gingiva, 8 buccal); bulk RNA-seq: sample sizes within GSE97615 not stated in text Groupsoral mucosa vs skin; injured vs uninjured tissue; KC subclusters KC1–KC5; oral KC4 vs skin KC4; HOK cells ± ruxolitinib (JAK1/2 inhibitor) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionCorrection method not named; 'adjusted P' threshold of 0.05 stated for bulk DEG calling and GO enrichment; no multiplicity correction stated for the repeated Wilcoxon comparisons across KC subclusters
Statistical tests used
Test Applied to n Assumptions
Differential expression analysis, bulk RNA-seq (specific test not named; 'adjusted P' implies a count-model-based test) Injured vs uninjured oral mucosa and skin (GSE97615) to define 485 shared wound-healing DEGs (P1 and P2 gene panels) Not stated; derived from public dataset GSE97615 not stated
Two-sided Wilcoxon rank-sum test P2-WHG enrichment score comparison across all cell types (oral mucosa vs skin) — Fig. 2e 103,758 cells from 5 skin and 21 oral mucosal samples not stated
Two-sided Wilcoxon rank-sum test P2-WHG enrichment scores across KC1–KC5 subclusters (oral vs skin) — Fig. 3d Keratinocyte subset of 103,758 cells; per-subcluster n not stated not stated
Gene Set Enrichment Analysis (GSEA) P2-WHG signaling in oral KC4 vs skin KC4 — Fig. 3e KC4 cells from 5 skin and 21 oral mucosal samples; exact cell count not stated not stated
Two-sided Wilcoxon rank-sum test Individual gene expression (S100A8, S100A9, IL36G, VIM, LUM, COL1A1) in oral vs skin KC4 — Fig. 4h KC4 subset; exact n not stated not stated
Two-sided Wilcoxon rank-sum test on gene-signature module scores EMT, wound healing, and inflammatory response scores in oral vs skin KC4 — Fig. 4i KC4 subset; exact n not stated not stated
Gene Ontology (GO) overrepresentation enrichment analysis P2-WHG gene list (Fig. 1c); cell-type marker genes (Fig. 2c); KC4 marker genes (Fig. 4g); all with adjusted P < 0.05 Gene lists derived from scRNA-seq and bulk RNA-seq; correction algorithm not named not stated
Approaches that could also have been used
  • Cell-level gene module scores were compared with Wilcoxon rank-sum tests, treating each cell as an independent observation across tissue types
    Could also: A pseudo-bulk approach — aggregating expression or scores per donor, then applying a donor-level t-test or linear model — would also account for within-donor correlation among cells — Cells from the same donor share biological and technical co-variation; pseudo-bulk methods propagate donor-level variance into the test statistic, reducing inflation of effective sample size that can arise from treating thousands of cells as independent units
  • Multiple pairwise Wilcoxon tests were applied across KC1–KC5 subclusters and across tissue types without a stated multiplicity correction
    Could also: A Kruskal-Wallis omnibus test followed by Dunn post-hoc correction, or Benjamini-Hochberg FDR applied across the family of pairwise Wilcoxon comparisons, would also control the family-wise error rate across the full set of subcluster comparisons — Applying a global test or explicit FDR correction across the full comparison family conveys control over false discovery rate when multiple subclusters are being contrasted simultaneously
  • Wilcoxon test results were reported only as p < 2.2e-16 without exact p values or effect size measures
    Could also: Reporting the rank-biserial correlation alongside the p value would also quantify the magnitude of the group difference for each Wilcoxon test — With tens of thousands of cells, Wilcoxon tests approach their numerical precision floor (p ≈ 2.2e-16) regardless of practical effect size; an explicit effect size metric conveys the biological magnitude of the difference independently of sample size
  • Bulk RNA-seq differential expression was called with an adjusted-P threshold without naming the underlying statistical model or correction algorithm
    Could also: DESeq2 (negative-binomial Wald test with Benjamini-Hochberg FDR) or edgeR (quasi-likelihood F-test) are standard, named methods for count-based bulk RNA-seq DEG calling — Naming the specific model and correction method improves reproducibility and allows readers to assess whether distributional assumptions (e.g., negative-binomial dispersion estimation) are appropriate for the data
  • GO enrichment was performed as overrepresentation analysis (ORA) on a threshold-defined gene list (|log2FC| > 1, adjusted P < 0.05)
    Could also: Ranked GSEA on the full fold-change-ranked gene list would also test pathway enrichment without requiring a binary inclusion threshold — Ranked methods use the continuous fold-change signal across all genes and are less sensitive to the choice of fold-change/FDR cutoff; the paper already applies GSEA in one context (Fig. 3e), so extension to GO pathways would be consistent with the existing analytical repertoire
  • No dispersion measure (SD, SEM, or CI) is reported for any quantitative comparison, including those involving a small number of biological donors (5 skin samples)
    Could also: Reporting inter-sample SD or 95% CI at the donor level (especially for the n = 5 skin group) would also convey between-donor variability alongside the cell-level statistics — Donor-level dispersion metrics allow readers to assess consistency across individuals and distinguish biological heterogeneity from cell-level technical noise, which is particularly informative when the number of donors is small relative to the number of cells
Software: SCENIC · STRING · CytoHubba (MCC algorithm, Cytoscape plugin)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
8
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

Addgene RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE164241 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE97615 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRP070705 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRR3184126 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRR3184127 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRR3184130 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39300285

Paper: Xuanyuan et al. 2024, Commun Biol — "SPRR1B+ keratinocytes prime oral mucosa for rapid wound healing via STAT3 activation." PMID 39300285 / PMC11413210 / DOI 10.1038/s42003-024-06864-5.

Listed code: https://github.com/YuLab-SMU/clusterProfiler (a third-party GO/KEGG enrichment tool — P16: applying an existing tool to the paper's data is equally valid). Listed data: GEO GSE97615.

What GSE97615 actually is

Bulk RNA-seq (Illumina HiSeq 2000), Homo sapiens, 24 samples = 12 skin (S1–S12)

  • 12 oral mucosa (O1–O12), each tissue with biopsies at day 1 (n=4), day 3 (n=4), day 6 (n=4) of wound healing. This is the Iglesias-Bartolome et al. wound-healing dataset. Processed values shipped as GSE97615_..._RPKM.xlsx. Note: GSE97615 contains only post-wound timepoints — there is NO "uninjured" baseline arm in this accession, so the paper's "uninjured vs injured" wording cannot be the contrast within GSE97615. The only well-defined within-dataset contrast that yields genes "constitutively expressed in oral mucosa" (their P2 panel) is oral mucosa (12) vs skin (12). We reproduce that contrast.

IN SCOPE (pipeline-derived, low-hanging, uses listed code+data)

# Reported result Paper loc Pipeline
C1 485 DEGs between tissues, limma, |log2FC|>1 & adj.p<0.05 Results/Methods (bulk RNA-seq) limma on GSE97615 RPKM
C2 146 P2-WHGs (oral-high module, "constitutively expressed in oral mucosa") Results, Fig.1 direction-split of C1 DEGs
C3 GO enrichment of P2-WHGs, clusterProfiler, P<0.05 & FDR<0.05, top 15 GO terms Methods + Fig clusterProfiler::enrichGO (org.Hs.eg.db, BP)

OUT OF SCOPE (not the listed code; heavy; not attempted)

  • All single-cell analysis: GSE164241 + PRJCA006797, 103,758 cells, Seurat + Harmony, 9 cell types, KC1–KC5, AddModuleScore (different data, heavy).
  • SCENIC TF regulons (STAT3/KLF5/PRDM1/GRHL1), STRING/Cytoscape PPI + CytoHubba hub genes, Monocle2 pseudotrajectory — separate tools, not clusterProfiler.
  • ChIP-seq (SRP070705, bowtie2/MACS2/annotatr) — separate raw-data pipeline.
  • All wet-lab (IHC, organoids, scratch assays) — out of scope by definition.

Reproduction strategy

One «our HPC» SLURM R job (conda prefix env: r-base, bioconductor-limma, bioconductor-clusterprofiler, bioconductor-org.hs.eg.db, r-readxl, r-openxlsx). Download RPKM xlsx inside the job (compute node has internet), run limma oral-vs- skin on log2(RPKM+1), count DEGs, split by direction, run clusterProfiler GO on the oral-high set, dump small result JSON/TSV. All data stays on «infra»; only small result values come back to «host».

Figures / tables: Fig.1
C1
Reported
485 DEGs
Reproduced
479 oral-upregulated genes (limma oral-vs-skin on log2(RPKM+1), |log2FC|>1 & adj.p<0.05; 471/473 under alt low-expression filters)
within tolerance
C2
Reported
146 P2 wound-healing genes
Reproduced
307/172 module split (ward.D2 k=2 of the 479 oral-high genes); no partition equals 146 — clustering method under-specified
partial
C3
Reported
GO of P2 genes: cornification/keratinization, immune, antimicrobial
Reproduced
clusterProfiler enrichGO BP: keratinization (q=3e-12), epidermis/skin development, keratinocyte differentiation, plus defense response to bacterium (q=0.031), establishment of skin barrier
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

The bulk-RNA-seq branch of Fig.1 reproduces well: 485 DEGs ≈ 479 (-1.2%) and the GO themes (keratinization q=3e-12, defense response to bacterium q=0.031) match 1:1 from the public GSE97615 data. The deviations are on the authors'/methods side but non-fabrication: the paper mis-describes the contrast as injured-vs-uninjured (no such samples in GSE97615) and under-specifies the clustering, so the reported 146 P2 genes is not derivable (ward.D2 k=2 → 307/172). Severity is moderate — direction and themes hold, exact sub-module does not — and most of the paper's mechanistic claims were out of scope, so the core conclusion is only limitedly confirmed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

136.7 k
tokens (I/O) · 6.8 M incl. cache
18 min
runtime · 0.04 CPU-h
2.3 GB
peak RAM
3
HPC jobs
hummel
machine