Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

PanglaoDB: a web server for exploration of mouse and human single-cell RNA sequencing data.

Database (Oxford) · 2019
L1 80/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 57% of all assessed papers rank 485 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PanglaoDB sample SRS2781556 (mouse microglia, 10x). Described WELL ENOUGH to reproduce the documented pipeline outputs from PanglaoDB's own shipped raw count matrix. 1:1 EXACT on all three per-sample QC summary statistics: cells=948, expressed genes=21771, median genes/cell=2590 (recomputed from the matrix after the documented >=1000-counts cell filter — the shipped matrix is the raw 70862-droplet matrix). The assigned tool Scrublet (0.2.3) was run on the 948-cell matrix and the paper's 'remove top 5%' policy reproduced deterministically (47 removed, 901 kept); noted that the fixed 5% policy is ~4x more aggressive than Scrublet's own automatic threshold (12 cells) here. NOT attempted: (a) full from-FASTQ count-matrix rebuild (HISAT2/featureCounts/UMI-tools on SRR6410935, ~43GB) because the 10x cell-calling/barcode whitelist is unspecified per sample, so the exact cell set is not faithfully recoverable; (b) the C4 removed-cell IDENTITY (PanglaoDB publishes no per-cell scores and pins no Scrublet version/params); (c) clustering (7 clusters) as resolution-dependent. No fabrication signal: reported QC metadata is bit-exactly derivable from the distributed data.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 80
    assessed: 2026-06-16 ⛓ ee2c01c8922f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Published single-cell RNA-seq data, though increasingly abundant in public archives, remains difficult for biological researchers to access due to raw-data formats and complex processing pipelines; the authors sought to build a unified, pre-processed, web-accessible database (PanglaoDB) to unlock exploration of this data and enable automated cell-type annotation.

Core claims
  • PanglaoDB is a web server providing pre-processed and pre-computed analyses of published mouse and human scRNA-seq experiments through a user-friendly interface. resource
  • PanglaoDB integrates more than 1054 single-cell samples (845 mouse, 209 human) from major platforms, representing more than 4 million cells. resource
  • A community-curated cell-type marker compendium of more than 6000 gene-cell-type associations was established as a resource for automatic cell-type annotation. resource
  • A cell-type activity (CTA) scoring method, using frequency-downweighted marker genes and a Fisher's exact test with FDR correction, accurately infers the cell type of a cluster. method
  • Regulons (transcription factor-centered co-expression networks) can be identified per cell cluster using GRNBoost2 co-expression networks combined with PWM motif enrichment (fimo) at transcription start sites. method
  • Gene set activity (GSA), cell cycle phase, differential expression, and disease association analyses are pre-computed for cell clusters using MSigDB, scran, and Seurat. method
  • Remapping and recounting raw SRA sequencing reads through a homogenized pipeline is preferable to using inconsistent GEO-deposited processed data. method
Experimental setups
Assay System Perturbation Readout Platform
scRNA-seq (10X Chromium, Drop-seq, SMART-seq2) data processing/clustering mouse and human tissues (multiple, e.g. cerebellum) none gene expression per cell, cell clusters, cell-type identity hisat2, Seurat v2.3.2, Scrublet, samtools/sambamba, subread, UMI-tools
Cell-type prediction validation against known biological cell type 17 human/mouse datasets with homogeneous, FACS-sorted cell populations none match between predicted and reported cell type
Orthogonal cell-type prediction validation on whole tissue samples 13 randomly selected tissue samples across various organs none consistency between expected/prominent tissue cell type and predicted cell type
Regulon/transcription factor network inference single-cell cluster expression data (mouse/human) none co-expressed genes under putative TF control, motif enrichment significance GRNBoost2, fimo (MEME v4.12.0), PWMs from JASPAR/HOCOMOCO/SwissRegulon/UniPROBE/CIS-BP
Gene set activity (GSA) analysis single-cell cluster expression data (mouse/human) none enrichment of MSigDB gene sets MSigDB v6.1, one-sided Fisher's exact test
Cell cycle phase assignment single cells within clusters none G1, G2M, S phase classification cyclone function, scran R package
Differential expression analysis single-cell clusters none differentially expressed genes between clusters Seurat FindMarkers
Key results
  • In all 17 validation datasets with known homogeneous cell populations, the biologically reported cell type matched the cell type predicted by the CTA method. 17/17
  • In all 13 randomly selected tissue samples, the expected/prominent cell type for the tissue was among the types predicted by the method. 13/13
  • PanglaoDB integrates 845 mouse and 209 human single-cell samples, totaling more than 4 million cells. 1054 samples total
  • The cell-type marker compendium contains more than 6000 gene-cell-type associations covering more than 150 possible cell types. >6000 associations
  • In one sample (SRS2781556) annotated as microglia, the method correctly identified the bulk as microglia but found a small subcluster predicted as neutrophils, attributed to shared myeloid lineage. 1 discordant subcluster
  • The regulon pipeline used a curated collection of 4104 PWMs representing 560 transcription factors. 4104 PWMs / 560 TFs
Key statistics
  • count >1054 single-cell experiments/samples (total integrated scRNA-seq samples in PanglaoDB (845 mouse + 209 human))
  • count >4 million cells (total cells represented across all integrated samples)
  • count >6000 gene-cell-type associations (size of the manually curated cell-type marker compendium)
  • count 17/17 datasets matched (validation of predicted vs reported cell type in homogeneous FACS-sorted samples)
  • count 13/13 tissue samples consistent (orthogonal validation of predicted vs expected cell type in whole-tissue samples)
  • pvalue adjusted P > 0.05 (Benjamini-Hochberg FDR) (threshold for labeling top-ranking predicted cell type as 'Unknown')
  • other AUC < 0.05 (significance threshold for transcription factor motif enrichment in regulon analysis)
  • count 4104 PWMs representing 560 transcription factors (positional weight matrix collection used for regulon/motif analysis)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a database/resource paper describing PanglaoDB, a web server aggregating mouse and human single-cell RNA-seq data through a unified bioinformatics pipeline; rather than hypothesis-testing experiments, its statistical components are embedded in computational analyses (cell-type inference, regulon/motif enrichment and gene-set activity). Cell-type calls used a cell-type activity score followed by a one-sided Fisher's exact (hypergeometric) test with Benjamini–Hochberg FDR control, and gene-set activity used a one-sided Fisher's exact test with Bonferroni correction. Method validation was descriptive, reporting concordance counts (17/17 sorted samples and 13/13 tissue samples) between predicted and expected/known cell types rather than formal inferential statistics.

Replicationunclear Sample sizevalidation used 17 independent sorted/homogeneous datasets and 13 randomly selected tissue samples; overall resource based on 1054 samples (845 mouse, 209 human) and >4 million cells; no power/sample-size calculation described Groupspredicted cell type vs known/expected cell type (method validation); enrichment vs background (motif/gene-set tests) Pairingna Randomization/blindingnot stated (13 tissue samples described as randomly selected) Dispersionnone Multiplicity correctionBenjamini–Hochberg FDR (cell-type inference) and Bonferroni correction (gene set activity)
Statistical tests used
Test Applied to n Assumptions
one-sided Fisher's exact test (hypergeometric test) cell-type inference: significance of the top-ranking cell type for each cell cluster, on genes expressed vs not expressed (expression >0) not stated
one-sided Fisher's exact test gene set activity (GSA) enrichment against MSigDB signatures gene sets restricted to between 10 and 500 genes not stated
kernel density estimate with integration over the curve (AUC-based significance), used in place of a Z-score regulon analysis: significance of a transcription-factor motif's enrichment in top-ranking co-expressed genes vs genomic background (threshold AUC<0.05) limited to top 200 ranking genes per motif stated (noted the distribution was often non-Gaussian, motivating the KDE approach)
Approaches that could also have been used
  • Cell-type significance was assessed with a one-sided Fisher's exact (hypergeometric) test on genes called expressed (>0) vs not.
    Could also: A permutation/empirical null built by shuffling gene-cell-type labels, or a rank-based enrichment such as a Mann-Whitney/AUCell-style test on continuous expression, could also be used. — A permutation or rank-based approach can incorporate expression magnitude rather than a binary expressed/not threshold, which may be informative when dropout makes the >0 cutoff sensitive to depth.
  • Multiplicity was handled with Benjamini–Hochberg FDR for cell-type inference and Bonferroni for gene set activity.
    Could also: A single consistent framework (e.g. BH FDR throughout, or Holm for family-wise control) could also be applied across both analyses. — Using one correction family throughout can make error-rate guarantees uniform across the resource; BH controls FDR while Bonferroni/Holm control family-wise error, and stating which target applies where can aid interpretation.
  • Validation was reported as exact concordance counts (17/17 and 13/13) between predicted and expected cell types.
    Could also: Accompanying these with a 95% confidence interval for the proportion (e.g. Wilson interval) could also convey the precision of the concordance estimate. — For small validation sets, an interval communicates the uncertainty around a perfect or near-perfect agreement rate that a raw count alone does not.
  • Motif enrichment significance was derived from a kernel density estimate integrated to an AUC, replacing the Z-score used in the source method because the distribution was often non-Gaussian.
    Could also: An empirical permutation P value against a shuffled-background null could also quantify significance without assuming a parametric form. — A permutation null is distribution-free like the KDE approach and additionally yields a directly interpretable P value and a natural route to multiple-testing correction across motifs.
  • Clustering resolution (Seurat resolution = 0.8) was selected by manual evaluation to produce fewer, larger clusters.
    Could also: A data-driven stability or silhouette-based criterion (e.g. scanning resolutions and selecting by cluster stability) could also guide the choice. — An automated stability metric can make the resolution choice reproducible and comparable across the many heterogeneous samples processed by the same pipeline.
Software: Seurat (FindClusters, FindMarkers) 2.3.2 · Scrublet (doublet detection) · scran (cyclone, cell cycle) · R packages sfsmisc and DescTools (regulon enrichment test) · GRNBoost2/arboreto (co-expression networks); also explored GENIE3 and WGCNA · fimo (MEME package, motif scanning) 4.12.0

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1,667
Impact: very high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (6)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Reproduction scope — pmid-30951143 (PanglaoDB)

Paper: Franzén, Gan & Björkegren (2019). PanglaoDB: a web server for exploration of mouse and human single-cell RNA sequencing data. Database (Oxford), baz046. PMID 30951143 · PMCID PMC6450036.

Assigned reproduction unit: apply the third-party tool Scrublet (github.com/AllonKleinLab/scrublet — doublet detection) to one PanglaoDB sample, SRS2781556 (PanglaoDB internal id SRA641134; SRA run SRR6410935; Mus musculus brain microglia, 10x chromium, "Repop D5"). Per BRIEF rule P16, using an existing third-party tool on the paper's own data is a valid reproduction.

What PanglaoDB's pipeline does (Methods, PMC6450036)

  1. Map reads with HISAT2 v2.1.0 (MAPQ ≥ 60) to GRCm38/GRCh38.
  2. GENCODE v27 annotation, exons collapsed to "meta genes", id = [symbol]_[ENSEMBL].
  3. featureCounts (subread v1.6.2) assigns gene to BAM; UMI-tools v0.5.3 dedup + count (3' UMI protocols).
  4. Keep cells with ≥ 1000 uniquely mapped reads after UMI dedup.
  5. Doublet detection: "Cell doublets were predicted with the Scrublet software and the top 5% cells with the highest scores were removed."
  6. Cluster retained cells; keep clusters with > 10 cells.

In scope (pipeline-derived, attempted)

# Result (reported for SRS2781556) Pipeline step How reproduced
C1 Number of cells = 948 (≥1000 reads) QC step 4 count columns of PanglaoDB's shipped raw count matrix
C2 Number of expressed genes = 21,771 quantification count genes (rows) with non-zero total
C3 Median expressed genes/cell = 2,590 quantification median of per-cell non-zero gene counts
C4 Doublet removal = top 5% by Scrublet score step 5 (assigned tool) run Scrublet on the shipped raw matrix; rank scores; top 5% = 47 of 948 cells; also report Scrublet's own auto-threshold doublet rate

C1–C3 are recomputed directly from PanglaoDB's own shipped per-sample raw count matrix (SRA641134_SRS2781556.sparse.RData; "Counts are not normalized, columns are cells, rows are genes"). This is an auditability check: do the numbers PanglaoDB reports match the data they distribute? C4 re-runs the assigned tool (Scrublet) exactly as the Methods describe.

Out of scope / hard 20% (not attempted, with reason)

  • Full from-FASTQ rebuild of the count matrix (HISAT2 → featureCounts → UMI-tools from SRR6410935, ~43 GB). Skipped as the costly, low-yield 20%: the 10x cell-calling / barcode whitelist PanglaoDB used is not specified per sample in the Methods, so the exact 948-cell set is not faithfully recoverable from the paper alone — many degrees of freedom. Using PanglaoDB's own shipped matrix removes that ambiguity and isolates the documented Scrublet step. Recorded as env/docs-limited for this sub-step, not a whole-RU drop.
  • Number of clusters = 7 (step 6): clustering method/resolution not pinned to an exact reproducible parameterisation; cluster count is resolution- dependent. Treated as optional; may be probed but not graded as exact.
  • Per-cell doublet scores: PanglaoDB publishes no per-cell Scrublet scores, so C4 is graded on the procedure (top-5% → 47 cells) and the data the step consumes, not on a published per-cell ground truth.

Drop assessment

NOT dropped. Code (Scrublet) resolves and is public; data (SRS2781556) resolves on SRA/ENA and PanglaoDB ships the exact raw matrix; identifiable reported results exist (948 / 21,771 / 2,590 / top-5%). → eligible, attempted.

C1
Reported
948 cells (>=1000 counted reads)
Reproduced
948
exact
C2
Reported
21771 expressed genes
Reproduced
21771
exact
C3
Reported
2590 median genes/cell
Reproduced
2590
exact
C4
Reported
Scrublet doublets: top 5% removed
Reproduced
47 of 948 removed (5%), 901 retained; Scrublet auto-threshold flags 12 (1.3%)
partial
C5
Reported
7 clusters (>10 cells)
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

The three directly comparable QC claims (948 / 21771 / 2590) reproduce bit-exactly from PanglaoDB's own distributed raw matrix, so the reported metadata is fully data-derived with no fabrication signal. The only gaps are on data-availability / underspecification: PanglaoDB ships no per-cell Scrublet scores (C4 verifiable at procedure level only — 47/948 removed) and pins no clustering resolution (C5 not attempted). These limits sit on the authors'/database side, not in our computation, and do not contradict any reported value — hence a solid but partial reproduction rather than a clean 1:1 across all five claims.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

209.1 k
tokens (I/O) · 15 M incl. cache
32 min
runtime · 0.09 CPU-h
8.4 GB
peak RAM
1
HPC jobs
hummel
machine