Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Allele-specific immune gene quantification and expression analysis in single-cell RNA-seq data.

NAR Genom Bioinform · 2025
L1 75/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
75/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 45% of all assessed papers rank 612 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for the DETERMINISTIC downstream, reproduced via the authors' own packages on the authors' own shipped data. scaeData ships the actual scIGD pipeline outputs (10x PBMC 5k/10k/20k); SingleCellAlleleExperiment::read_allele_counts(filter_mode='yes') runs the documented empty-droplet (knee-plot inflection) filtering deterministically. C2 (the multi-layer allele/gene/functional-class SCAE structure) reproduced EXACTLY: 17 alleles typed, aggregated into immune genes + 2 functional classes (HLA class I/II) for the 20k set. C1 (headline '19765 cells retained' for the 20k PBMC behind Fig 4C-F) does NOT match 1:1: the documented empty-droplet step alone retains 24025 cells; the paper's lower 19765 implies additional QC (mito/nGene and/or doublet removal) whose thresholds are not in the public Methods/packages (figure code is in private repo AGImkeller/scIGD_manuscript_2024 + a 2 GB Zenodo zip, 10.5281/zenodo.16919082). Direction (fewer after extra QC) is consistent and plausibly derivable -> under-specified preprocessing, NOT a fabrication signal. NOT attempted (hard 20%): the full Snakemake genotyping/quantification from raw FASTQ (arcasHLA + kallisto), and therefore the Merkel-cell GSE117988 cell counts (1794/3439/2118/4896), the scIGD-vs-kallisto Pearson~1.0 (Fig 3C), and AIDA HLA-A*24:02 frequencies; the multiple-myeloma data is non-public (on request).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 75
    assessed: 2026-06-14 ⛓ de14bc45ea64
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can allele-specific expression of hyperpolymorphic immune genes (notably HLA) be accurately typed and quantified from single-cell RNA-seq data through an integrated, user-friendly computational workflow that also enables multilayer (allele/gene/functional-class) representation and interactive exploration?

Core claims
  • scIGD is a Snakemake-based workflow that automates HLA allele-typing and allele-specific expression quantification from scRNA-seq data. method
  • SingleCellAlleleExperiment (SCAE) is a novel R/Bioconductor data structure extending SingleCellExperiment that represents immune gene expression across allele, gene, and functional-class layers. resource
  • The methodology accurately types and quantifies HLA alleles across diverse single-cell sequencing datasets, platforms, and experimental setups (WTA and amplicon-based). finding
  • The tools detect loss of HLA expression in tumor cells and discover differential HLA allele expression in specific immune cell subtypes. finding
  • SCAE aggregates allele-level counts into gene-level and functional-class-level counts via a lookup table while ensuring normalization operates only on original (non-aggregated) features. method
  • scIGD integrates arcasHLA for comprehensive WTA HLA typing and kallisto-bustools (with -mm multi-mapping) for alignment-free allele-specific quantification. method
  • scaeData is an R/ExperimentHub package providing curated example datasets for testing the workflow. resource
  • Allele-typing of the AIDA cohort revealed common HLA-A and HLA-C alleles shared across Korean and Japanese populations consistent with shared ancestry. finding
Experimental setups
Assay System Perturbation Readout Platform
WTA scRNA-seq (HLA allele-typing and allele-specific quantification) AIDA cohort PBMCs, 14 donors (8 Korean, 6 Japanese) none HLA allele typing across ethnic backgrounds 10x 5' v2
WTA scRNA-seq (allele-specific quantification, downstream analysis) Merkel-cell carcinoma, PBMC and tumor samples (pre- and post-treatment), 1 donor longitudinal treatment (longitudinal); tumor vs PBMC HLA loss in tumor cells; allele-specific expression 10x 3' v2
WTA scRNA-seq (quantification and clustering) 20k PBMC dataset, 1 donor none differential HLA allele expression across immune cell subtypes; Louvain clustering 10x 3' v3
Amplicon-based scRNA-seq (multiplexed, demultiplexing + minimal allele-typing) Multiple myeloma dataset, 8 donors none allelic variant typing and quantification BD Rhapsody
Key results
  • Three of eight Korean and four of six Japanese individuals carried the HLA A*24:02:01:01 allele, the overall most common allele. 3/8 KR and 4/6 JP
  • scIGD successfully typed and quantified HLA alleles across four distinct datasets spanning different platforms and experimental setups.
  • Full scIGD workflow processed a 50M-reads WTA dataset in ~3 h using 32 cores and 40 GB RAM. ~3 h, 50M reads, 32 cores, 40 GB RAM
  • An amplicon-based dataset of 280M reads was analyzed in ~80 min using 8 cores and 8 GB RAM. ~80 min, 280M reads, 8 cores, 8 GB RAM
  • Common HLA-A and HLA-C alleles were shared across Korean and Japanese populations, consistent with shared ancestry.
Key statistics
  • count ~28 000 HLA class I alleles (HLA class I alleles identified to date)
  • count ~12 400 HLA class II alleles (HLA class II alleles identified to date)
  • count >200 genes (genes in the HLA region on chromosome 6)
  • other at least 75% (read-count threshold for assigning a sample tag ID to a cell barcode during demultiplexing)
  • count 14 (AIDA cohort donors analyzed (8 Korean, 6 Japanese))
  • count 8 (donors in the multiplexed Multiple myeloma amplicon dataset)
  • other k = 50 (Louvain clustering parameter used for the 20k PBMC dataset)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational methods paper introducing scIGD (a Snakemake workflow) and SCAE (an R/Bioconductor package) for allele-specific immune gene quantification in scRNA-seq data. Validation was performed across four diverse scRNA-seq datasets using HLA allele typing concordance with population-level expectations and allele-specific expression patterns. Downstream analyses followed standard scRNA-seq preprocessing conventions (normalization, batch correction, dimensionality reduction, graph-based clustering) rather than formal hypothesis-testing frameworks. Biological illustrations (HLA loss in tumor, differential allele expression across cell types) are presented primarily via visual inspection of expression patterns in the available text.

Replicationmixed Sample sizeFour datasets used for validation: AIDA cohort (n=14 donors, 8 Korean / 6 Japanese, selected for highest sequencing depth from the full cohort); Merkel-cell carcinoma (n=1 donor, 4 longitudinal time points, PBMC and tumor); 20k PBMC (n=1 donor); Multiple myeloma (n=8 donors, multiplexed). No power calculation or formal sample-size justification described. GroupsKorean vs Japanese HLA allele distributions (AIDA); tumor vs PBMC and pre- vs post-treatment (Merkel-cell carcinoma); cell-type clusters defined by Louvain clustering (20k PBMC); allele-specific expression across immune cell subtypes Pairingmixed Randomization/blindingnot stated Dispersionnone
Statistical tests used
Test Applied to n Assumptions
Expectation-maximization (EM) model for HLA allele probability estimation (via arcasHLA) HLA allele typing stage: determines the combination of alleles that best explains measured sequencing reads mapped to chromosome 6 HLA loci 14 donors from AIDA cohort (8 Korean, 6 Japanese); additional single-donor and 8-donor datasets not stated
Louvain community detection on shared nearest neighbor (SNN) graph (via igraph) Cell-type clustering in the 20k PBMC dataset; k=50 selected based on returning 'a reasonable number of clusters' ~20 000 cells (implied by dataset name; exact post-QC n not stated in available text) not stated
Mutual nearest neighbors (MNN) batch correction (multiBatchNorm, reference [23]) Integration of four longitudinal Merkel-cell carcinoma samples into a single SCAE object prior to PCA and t-SNE 4 samples from 1 donor; per-sample cell counts not stated in available text not stated
Principal component analysis (PCA) on highly variable genes Dimensionality reduction step in both Merkel-cell carcinoma (post-MNN) and 20k PBMC (post-logNormCounts) analyses not stated
Approaches that could also have been used
  • Validation of HLA allele typing accuracy was demonstrated by concordance with population-level allele frequencies and shared alleles across ethnic groups (e.g. A*24:02:01:01 prevalence in Korean and Japanese donors)
    Could also: Formal quantitative accuracy metrics — field-level concordance rates and allele-level accuracy — compared against a clinical gold-standard HLA typing (e.g. Sanger or long-read NGS-based typing) could also be reported, together with confidence intervals — Quantitative concordance metrics with uncertainty estimates would allow direct numerical benchmarking against other HLA typing tools and make performance claims reproducible across laboratories
  • Batch correction across four Merkel-cell carcinoma samples was performed with MNN (mutual nearest neighbors)
    Could also: Harmony, Seurat RPCA, or scVI (deep generative model-based integration) are widely used alternatives for scRNA-seq batch correction; a benchmarking comparison using established metrics (e.g. scib) could also accompany the chosen method — Different integration methods carry different assumptions about batch structure and vary in how much biological signal they preserve; showing robustness across at least two methods, or citing a benchmark, helps readers gauge the sensitivity of downstream biological conclusions to the integration choice
  • Louvain clustering was applied with k=50, selected because it returned 'a reasonable number of clusters'
    Could also: The Leiden algorithm (a refinement of Louvain that guarantees well-connected communities) could also be used; additionally, cluster stability across a range of resolution parameters (e.g. assessed by adjusted Rand index) could guide the final choice more formally — Leiden is increasingly preferred in the field as it avoids the poorly connected community problem; quantitative stability metrics make the clustering choice more transparent and reproducible
  • Differential HLA allele expression across immune cell subtypes is a stated biological application of the method, with cell types identified by clustering
    Could also: Pseudobulk differential expression methods (DESeq2 or edgeR on donor-level aggregate counts per cell type) could also be applied when multiple donors are available, complementing or replacing single-cell-level tests — Pseudobulk approaches account for within-donor correlation and have been shown to have better-calibrated type I error rates than tests applied directly to individual-cell observations; they are particularly relevant when the goal is population-level inference across donors rather than within a single individual
  • Normalization was performed with logNormCounts (library-size scaling followed by log1p transformation)
    Could also: Pooling-based size factor estimation (scran computeSumFactors) or Pearson-residual normalization (SCTransform / glmGamPoi) could also be considered as normalization strategies — Library-size normalization assumes equal total RNA content across cells, an assumption that may not hold when comparing highly heterogeneous cell types; pooling- or regression-based methods relax this assumption and may better preserve biological variation, which could be relevant given the goal of comparing allele expression across distinct immune cell populations
  • Donor selection from the AIDA cohort was based on highest sequencing depth rather than random sampling
    Could also: A random or stratified sample from the full cohort could also be used for validation; alternatively, a sensitivity analysis showing that performance metrics are stable across a range of sequencing depths could accompany the depth-selected subset — Selecting high-depth donors may yield a more favorable performance estimate; a depth-stratified analysis or sensitivity check would clarify how much workflow accuracy depends on sequencing depth, which is useful practical guidance for prospective users
Software: R/Bioconductor — SingleCellAlleleExperiment (SCAE) · Snakemake (scIGD workflow) · arcasHLA · kallisto-bustools · STAR aligner · igraph (Louvain clustering) · scran — logNormCounts / batchelor — multiBatchNorm / MNN (referenced as [23][24])

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41229397 (scIGD)

Paper: Al Ajami, Schuck, Marini, Imkeller. Allele-specific immune gene quantification and expression analysis in single-cell RNA-seq data. NAR Genom Bioinform 2025. PMID 41229397 / PMC12604667 / DOI 10.1093/nargab/lqaf149.

Software (authors' own, MIT):

  • scIGD — Snakemake workflow: demultiplex → HLA allele-typing (arcasHLA) → allele-specific quantification (kallisto|bustools -mm). Repo github.com/AGImkeller/scIGD @ fed35f64568a6644cbd94f883a45f15fe12b75bc.
  • SingleCellAlleleExperiment (SCAE) — Bioconductor R class extending SingleCellExperiment with a 3-layer ontology (allele → gene → functional class) built from a lookup table. v1.7.0 @ 51ef84b1c6de24c554a37312964f7b30f509654e.
  • scaeData — ExperimentHub data package shipping scIGD pipeline outputs for three 10x Genomics PBMC datasets (5k / 10k / 20k, all 3' v3). @ 0c83291a09b1a32101099da2a6c7e123bd12508e.

Pipeline-derived results in the paper

# Reported result Pipeline In scope? Why
C1 20k PBMC: 19,765 cells retained after preprocessing scIGD output → SCAE empty-droplet filtering (knee-plot inflection point) YES scaeData ships the exact scIGD output (pbmc_20k); SCAE read_allele_counts(filter_mode="yes") runs DropletUtils::barcodeRanks() inflection filtering — fully deterministic, 1:1 checkable.
C2 SCAE multi-layer structure: alleles aggregate into immune genes (HLA-A/B/C, DR/DQ/DP…) and into functional classes (HLA class I / II) SCAE object construction from scIGD output + lookup table YES (structural) Reproduce the number of allele rows, added immune-gene rows, added functional-class rows for pbmc_20k; verifies the documented data-structure behaviour.
Merkel-cell carcinoma (GSE117988) retained cells 1794/3439/2118/4896; complete loss of both HLA-B alleles post-treatment full scIGD Snakemake from raw FASTQ NO Raw-read pipeline (STAR+arcasHLA+kallisto, ~3 h/donor, 32 cores) and these matrices are NOT in scaeData. Hard 20%.
scIGD vs standard kallisto bustools gene-level Pearson ≈ 1.0 (Fig 3C) run both quantifiers on raw FASTQ NO
AIDA cohort HLA-A*24:02:01:01 in 3/8 Korean + 4/6 Japanese donors full pipeline on HCA AIDA data NO Data on HCA DCP (controlled-ish), per-donor pipeline runs. Out of scope.
Multiple myeloma (BD Rhapsody, 8 donors) amplicon results full pipeline NO Data "not publicly available; contact authors" → data_restricted.

Reproduction strategy (80/20)

Reproduce C1 + C2 by running the authors' own downstream package (SCAE) on the authors' own shipped pipeline output (scaeData pbmc_20k) on «our HPC». This is a faithful 1:1 test of the deterministic, clearly-specified part of the pipeline. The hard 20% (running the Snakemake genotyping/quantification from raw FASTQ on GSE117988 / AIDA) is explicitly not attempted — both the compute (multi-hour STAR+arcasHLA+kallisto per donor) and, for the headline cohorts, the raw data gating put it outside the low-hanging-fruit budget. The 5k & 10k datasets are also processed for context (no paper number attached to their retained-cell counts).

Figures / tables: Fig 4CFig 2Table
C1
Reported
19765 cells retained (20k PBMC, after preprocessing; Fig 4C-F)
Reproduced
24025 cells retained (empty-droplet knee-inflection filtering, inflection=433.79 UMIs)
partial
C2
Reported
SCAE 3-layer structure: allele -> immune gene -> functional class (HLA class I/II)
Reproduced
pbmc_20k SCAE object: 17 allele rows, 2 functional-class rows, 62753 gene rows (Quant_type A/F/G)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

The deterministic downstream reproduced cleanly on the authors' own shipped scaeData: C2's allele→immune-gene→functional-class structure (17 alleles, 2 functional classes, 62753 genes for pbmc_20k) matched exactly. C1's headline count differs — 24025 reproduced vs 19765 reported (+21.6%) — because the paper applies additional standard QC beyond the documented empty-droplet step whose thresholds live only in private figure code. This is under-specified preprocessing on the authors' side, with the correct direction (fewer cells after more QC) and no fabrication signal; the method's central claim holds while the specific Fig 4C-F number cannot be matched 1:1 from public materials.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

122.3 k
tokens (I/O) · 11.3 M incl. cache
20 min
runtime · 0.06 CPU-h
12.3 GB
peak RAM
1
HPC jobs
hummel
machine