Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Creation of a Single Cell RNASeq Meta-Atlas to Define Human Liver Immune Homeostasis.

Front Immunol · 2021
L1 63/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
63/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 24% of all assessed papers rank 875 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce: the paper combines 3 public liver scRNA-seq deposits (GSE115469/124395/125188) into a CD45+ immune meta-atlas (Seurat v3 + Harmony validation). Code link is the third-party Harmony tool (no authors' analysis repo); per P16 we reproduced by running an independent scanpy 1.10.4 + harmonypy-equivalent pipeline on the same data. STRONG 1:1 on structure and concordance: all 4 immune lineages recovered; GSE115469 and GSE124395 lineage proportions fall in the reported ranges and match the original authors' own labels; all 3 inter-dataset pseudobulk Pearson correlations reproduce within 0.02 (0.944/0.82/0.807 vs 0.95/0.81/0.79). TWO reported figures are contradicted by the deposited data and flagged possible-discrepancy: (C4) the paper's B-cell range of 2-7% is incompatible with GSE125188, which is B-rich (~19% B per its own deposited author labels, ~22% in our re-run); (C5) the ~32,000-cell atlas is smaller than the GSE125188 CD45+ matrix alone (~60,672 cells), implying undocumented subsampling. NOT attempted (hard last-20% / out of scope): combined Harmony merge, Seurat-vs-Harmony 96.5-99.9% agreement meta-claim, between-dataset DE-gene counts, IPA pathway analysis (commercial).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 63
    assessed: 2026-06-15 ⛓ 3a5aa75e25a3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a meta-analysis of existing normal human liver scRNA-seq datasets be performed—accounting for inter-study heterogeneity—to generate a comprehensive human liver immune meta-atlas that defines the dominant phenotypes of hepatic immune cell subpopulations and serves as a reference for healthy liver immune homeostasis?

Core claims
  • Independent human liver immune scRNA-seq datasets can be combined into an integrated meta-atlas in which all datasets co-cluster, despite differing cell-type proportions between studies. finding
  • A reproducible method (Seurat v3 integration with anchor-based batch correction) for combining heterogeneous human liver scRNA-seq datasets into a re-usable immune meta-atlas. method
  • Canonical pathways differing between datasets relate to cell stress and oxidative phosphorylation rather than immune-related function, supporting that integration is biologically meaningful. finding
  • Hepatic immune homeostasis is characterized by decreased expression across immunologic pathways and enhancement of pathways involved with cell death, relative to PBMC. mechanism
  • An online interactive human liver immune meta-atlas resource defining gene signatures for myeloid, NK/T, B, and plasma cell subpopulations. resource
  • Clustering results are robust to algorithm choice, with 96.5–99.9% agreement between Seurat and Harmony clusters. finding
  • NK and T cells are the most abundant hepatic immune subpopulation, while B cells and plasma cells are the rarest, consistent across datasets. finding
Experimental setups
Assay System Perturbation Readout Platform
scRNA-seq (mCEL-Seq2) Human liver, 9 liver resection patients (mCRC/ICC), CD45+ leukocytes (LACEe / Aizarani et al.) none single-cell UMI gene expression counts mCEL-Seq2; GSE124395
scRNA-seq (10x Genomics) Human liver, 5 caudate lobes of DBD transplant donors, CD45+ leukocytes (Lnb / MacParland et al.) none single-cell UMI gene expression counts 10x Genomics; GSE115469
scRNA-seq (10x Genomics) Human liver (CD45-enriched), 3 adult transplant donors (blood, spleen, liver barcoded; only liver used) (LCD45e / Zhao et al.) none single-cell UMI gene expression counts 10x Genomics; GSE125188
Integrated scRNA-seq meta-analysis (clustering/UMAP, differential expression) Combined human liver immune meta-atlas, 17 patients / ~32,000 CD45+ cells none integrated clustering, immune subpopulation proportions, differential gene expression Seurat v3.0 in R (RRID:SCR_007322); Harmony R package
Reference scRNA-seq comparison (differential expression) Normal human PBMC, 67,221 cells none differential gene expression vs liver immune subpopulations GSE171555
Ingenuity Pathway Analysis Differential expression gene lists from liver immune subpopulations none up/downregulated canonical pathways, gene function heatmaps IPA, Qiagen (RRID:SCR_008653)
Gene expression correlation analysis (linear/quadratic regression, RRHO) Pairwise between three liver datasets none correlation coefficient R / rank-rank hypergeometric overlap RRHO package in R
Key results
  • All three datasets co-clustered on shared UMAP coordinates with homogeneous interdigitation of clusters
  • Lnb and LCD45e (both 10x platform) showed high gene expression correlation R=0.95
  • LACEe had lower concordance with LCD45e R=0.79
  • LACEe had lower concordance with Lnb R=0.81
  • Immune cell subpopulation proportions differed significantly across all three datasets for each cell type p<0.01 each
  • Seurat vs Harmony clustering agreement across cell types 96.5%-99.9%
  • Hepatic immune homeostasis shows decreased expression across immunologic pathways and enhanced cell-death pathways vs PBMC
  • LCD45e comprised the majority of cells (~24,000) vs ~4,000 in each of the other two studies ~24,000 vs ~4,000
Key statistics
  • correlation R=0.95 (Gene expression correlation Lnb vs LCD45e (both 10x))
  • correlation R=0.81 (Gene expression correlation LACEe vs Lnb)
  • correlation R=0.79 (Gene expression correlation LACEe vs LCD45e)
  • pvalue p<0.01 (Chi-squared test, differences in each immune cell subpopulation proportion across datasets)
  • pvalue p=0.16 (Chi-squared, plasma cells LCD45e vs Lnb (not significant))
  • pvalue p=0.08 (Chi-squared, B cells Lnb vs LACEe (not significant))
  • count ~32,000 hepatic CD45+ cells (Combined meta-atlas across 17 normal human liver samples)
  • count 67,221 cells (Normal human PBMC reference dataset (GSE171555))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study performed a meta-analysis of three publicly available scRNA-seq datasets of normal human liver immune cells (~32,000 CD45+ cells from 17 donors), integrating them using Seurat v3 anchor-based batch correction. Dataset comparability was assessed via Chi-squared tests on cell-type proportions and Pearson/Spearman gene expression correlation. Differential gene expression between immune subpopulations and between the liver meta-atlas and a PBMC reference was conducted using the Wilcoxon rank-sum test, with Bonferroni-adjusted p-values and log fold-change thresholds used to define significant genes for downstream pathway analysis in IPA.

Replicationbiological Sample size17 donors drawn from three published datasets (9, 5, and 3 donors respectively); ~32,000 CD45+ cells total; no formal a priori power calculation described GroupsThree liver scRNA-seq datasets (LACEe, Lnb, LCD45e) compared against each other and against a published PBMC reference (67,221 cells) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionBonferroni correction (P-adj Bonfcorr)
Statistical tests used
Test Applied to n Assumptions
Chi-squared test Comparison of immune cell subpopulation proportions between the three datasets (Figure 2I) and pairwise between datasets ~32,000 cells across 17 donors (cell counts per subpopulation per dataset; exact per-group counts not stated) not stated
Wilcoxon rank-sum test (via Seurat FindMarkers) Differential gene expression between each immune cell subpopulation vs. all other subpopulations within and across datasets; also liver meta-atlas vs. PBMC reference Varies by subpopulation and dataset; exact per-group cell counts not stated not stated
Two-sample t-test Difference-of-differences analysis comparing differential expression between datasets and between cell subpopulations (volcano plots) Not stated not stated
Linear and quadratic regression (Pearson R) Pairwise gene expression correlation between datasets (Figures 2E–G) Averaged expression per gene across cells; number of genes not stated not stated
Spearman rank correlation Correlation of differential-expression rankings between datasets, stratified by cell type Not stated not stated
Rank-Rank Hypergeometric Overlap (RRHO) Assessment of overlap in ranked DE gene lists between datasets by cell type Not stated not stated
Approaches that could also have been used
  • Cell-type proportions were compared between datasets using Chi-squared tests, treating cell counts as independent observations
    Could also: Compositional data analysis methods such as Dirichlet regression or scCODA (a Bayesian compositional model for single-cell data) could also be applied — Cell-type proportions from scRNA-seq sum to one and are compositional by nature; methods designed for compositional data explicitly model this constraint and the correlation structure among cell types, and scCODA additionally accounts for donor-level variability
  • Differential gene expression was performed at the single-cell level using the Wilcoxon rank-sum test via Seurat FindMarkers, treating each cell as an independent observation
    Could also: Pseudobulk approaches — aggregating counts to the donor level and then applying DESeq2 or edgeR — could also be used — Cells from the same donor are not statistically independent; pseudobulk methods propagate donor-level variability into the test, reducing inflated type I error rates that can arise when many cells per donor are treated as independent replicates
  • Bonferroni correction was applied to control for multiple testing across genes in the volcano plot analyses
    Could also: Benjamini-Hochberg false discovery rate (FDR) correction could also be applied across the same family of gene-level tests — Bonferroni controls the family-wise error rate and is conservative when tests are correlated (as co-expressed genes tend to be); BH-FDR controls the expected proportion of false discoveries and is widely used in genomic contexts where thousands of genes are tested simultaneously, typically yielding more discoveries at the same nominal error threshold
  • Multiple pairwise Chi-squared tests were conducted across cell types and dataset pairs without a stated correction for the family of comparisons
    Could also: A single omnibus test (e.g., a log-linear model or permutation-based test) followed by post-hoc pairwise comparisons with a correction such as Bonferroni or Holm could also be used — Running multiple pairwise Chi-squared tests increases the probability of at least one spurious significant result; an omnibus-then-post-hoc strategy makes the error-rate control for the full family of comparisons explicit
  • Seurat v3 anchor-based canonical correlation analysis (CCA) integration was used for batch correction across the three datasets
    Could also: Alternative integration methods such as scVI (a deep generative model) or BBKNN (batch-balanced k-nearest neighbors) could also be applied to the same datasets — Different integration methods make different assumptions about batch structure; benchmarking studies show that no single method dominates across all data configurations, and trying a second method provides additional evidence that conclusions are not integration-algorithm-dependent (the authors did partially address this by comparing Seurat to Harmony)
  • Gene expression similarity between datasets was summarized using Pearson correlation on per-gene average expression values
    Could also: Mutual nearest neighbors (MNN) distance or a cosine similarity on the full single-cell embedding could also quantify dataset concordance — Pearson correlation on averaged expression collapses within-cell-type heterogeneity and can be dominated by highly expressed genes; cell-embedding-level similarity metrics preserve single-cell resolution and are less sensitive to a small number of outlier genes
Software: R/Seurat 3.0 · R/Harmony · R/RRHO · R/Circlize · Ingenuity Pathway Analysis (IPA, Qiagen) · RStudio

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
18
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE115469 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE124395 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE125188 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE171555 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
RRID:SCR_007322 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_008653 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

Downstream reach in the literature

94 downstream papers · 4 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34335581

Paper: Rocque B, Barbetta A, Singh P, et al. "Creation of a Single Cell RNASeq Meta-Atlas to Define Human Liver Immune Homeostasis." Front Immunol 2021;12:679521. PMCID PMC8322955. DOI 10.3389/fimmu.2021.679521.

What the paper does (pipeline)

Integrates THREE published human-liver scRNA-seq datasets into one CD45+ immune "meta-atlas", clusters, annotates immune lineages, compares composition.

paper alias dataset GEO platform cells sequenced
LACEe Aizarani et al. liver atlas GSE124395 mCEL-Seq2 10,372
Lnb MacParland et al. GSE115469 10x 8,444
LCD45e Zhao et al. CD45+ GSE125188 10x 70,706

Tools: Seurat v3.0 (integration FindIntegrationAnchors + clustering FindClusters, UMAP), Harmony R package as clustering validation, IPA (pathway, commercial). Normalization = library-size scaling, factor 1e4, log1p. 30 PCs.

In scope (pipeline-derived → attempted)

  • S1 Recover the 4 major immune lineages: NK&T (CD3D/KLRF1/FCGR3A), Myeloid (CD14/FCGR3A), B (CD19), Plasma (SDC1/CD138). [Results, Fig 1-2]
  • S2 Per-dataset immune-lineage proportions vs reported ranges: NK&T 51-69%, Myeloid 18-32%, Plasma 4-14%, B 2-7%.
  • S3 Total hepatic CD45+ cells ≈ 32,000 across the 3 datasets.
  • S4 (stretch) Pseudobulk Pearson R between datasets: LnbLCD45e 0.95, LACEeLnb 0.81, LACEe~LCD45e 0.79.

Out of scope (not pipeline / not attempted)

  • IPA pathway analysis (commercial Qiagen software, no key) — env_unresolvable.
  • Seurat-vs-Harmony 96.5-99.9% cluster-agreement: a meta-claim about the authors' own two clusterings on their own integrated object; not independently re-derivable without their exact Seurat object. Noted, not graded 1:1.
  • DE-gene counts between datasets (3,526 / 767 / 924 …) and signature sizes (16/22/45/54): highly sensitive to exact integration + DE test params; skipped as the hard last-20%.
  • Shiny web resource (interactive, not a numeric pipeline output).

Approach

Code link is immunogenomics/harmony (third-party tool) → per brief P16 this is a valid reproduction by applying the documented pipeline (Seurat-equivalent scanpy + Harmony) to the paper's own data. We run an INDEPENDENT scanpy 1.10 + harmonypy pipeline (standard equivalent of Seurat v3 + Harmony) on the three GEO matrices. Author-provided cluster labels (shipped by GSE125188 / GSE115469) are used only as an orthogonal cross-check, not as the reproduced value.

Figures / tables: FigsFig 2
C1
Reported
4 immune lineages (NK&T, Myeloid, B, Plasma)
Reproduced
all 4 recovered in all 3 datasets
exact
C2
Reported
GSE115469 in ranges NK&T 51-69/Mye 18-32/Plasma 4-14/B 2-7
Reproduced
54.1/30.4/12.5/3.0 (all in range; matches author labels)
within tolerance
C3
Reported
GSE124395 same ranges
Reproduced
NK&T 69.5/Mye 24.4/Plasma 3.9/B 2.1 (boundary)
within tolerance
C4
Reported
GSE125188 B 2-7%, NK&T 51-69%
Reproduced
NK&T 43.0/Mye 20.3/B 22.3/Plasma 14.4 -- matches Zhao author subsets (B 19.1%), NOT paper range
did not match
C5
Reported
~32,000 total CD45+ cells
Reproduced
~72,009 immune; GSE125188 alone ships ~60,672 CD45+
did not match
C6
Reported
Pearson R 0.95/0.81/0.79
Reproduced
0.944/0.82/0.807 (within 0.02)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 63/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

Structure and the central quantitative claim reproduce strongly: all 4 immune lineages recover, 2/3 dataset compositions fall in range, and all three pseudobulk Pearson R reproduce within 0.02 (0.944/0.82/0.807 vs 0.95/0.81/0.79). However, two reported figures — the 2-7% B-cell range (C4) and the ~32,000-cell atlas (C5) — are contradicted by GSE125188's own deposited data (~19% B, ~60,672 CD45+ cells), and our independent re-run matches the original deposit authors rather than this paper. The discrepancy sits on the authors' side (undocumented subsampling/denominator and an overgeneralized composition range), not our methodology. Net: solid reproduction with two unsupported descriptive figures flagged for admin review — not core-conclusion fabrication.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

171.5 k
tokens (I/O) · 10.2 M incl. cache
28 min
runtime · 0.1 CPU-h
18.1 GB
peak RAM
2
HPC jobs
hummel
machine