Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Meta-analysis of COVID-19 single-cell studies confirms eight key immune responses.

Sci Rep · 2021
L1 69/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
69/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 35% of all assessed papers rank 745 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL (1:1 on what is reproducible from the open dataset; the cross-dataset headline is a structural DROP). Repo Manikgarg/COVID-19 is downstream-only (its processed Harmony/SCCAF AnnData is undeposited), so we rebuilt the Methods-described scanpy pipeline on the room's open dataset GSE149689 and applied the repo's VERBATIM ISG (99-gene) and HLA-II (^HLA-D.*) gene lists with sc.tl.score_genes(ctrl_size=100). RESULTS: (1) Cell census reproduces EXACTLY — 20 samples and 58022 cells == Table 1 at the post-basic-QC stage (raw 85144), strong evidence the reported N is genuine and the stated QC thresholds are exactly as applied. (2) Response #5 (HLA class II downregulation in severe-COVID classical monocytes) reproduces strongly and per-sample-consistently (Healthy 0.80 -> Severe 0.24). (3) Response #2/#8 (type I IFN/ISG) reproduces only PARTIALLY: an ISG response is present and significantly elevated in COVID monocytes overall (driven by mild/asymptomatic), but it is IMPAIRED — not strongest — in severe disease (severe < mild, p=1e-149; opposite to the reported severity direction). This is biologically coherent (interferon-low dysfunctional monocytes in severe COVID) and the reproduced values are fully derivable from the shipped GEO data, so it is a genuine analytical discrepancy for human review, NOT a fabrication signal. (4) Major cell-type proportions: T cells 31.9% (in the reported 30-40%); monocytes 30.3% (vs ~20%, somewhat high; major-type annotation only, no SCCAF subtypes). NOT ATTEMPTED: the 80%/40% cross-dataset reproducibility headline (requires all 9 datasets incl. EGA-controlled Chua and the undeposited processed objects), SCCAF 37-subtype labels, and TCR/BCR clonal responses (no VDJ deposited). Dataset profiled: GSE149689 delivers its PBMC GEX exactly as promised (cell N matches to the cell) but omits the VDJ data; quality grade B. ALL GRADES ARE PROVISIONAL and must be independently checked by a human.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ 8b30fffb3ecc
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Given that most published scRNA-seq studies of COVID-19 immune response have small sample sizes, the paper tests whether the immunological conclusions from these individual studies are reproducible when re-analyzed in a standardized manner across multiple independent datasets.

Core claims
  • Only 8 of 20 previously published COVID-19 scRNA-seq findings were reproducible across all relevant datasets in a standardized meta-analysis finding
  • 16 of 20 (80%) published results were reproducible when reanalyzed within the original study's own dataset finding
  • T-cell proportion consistently decreases with increasing COVID-19 severity across datasets finding
  • Type I interferon signaling pathway (IFIT, IFI, ISG, OAS genes) is consistently upregulated across many immune cell types in COVID-19 patients finding
  • B-cell clonal expansion is consistently increased in COVID-19 patients compared to healthy controls across studies finding
  • No consistent cross-study trend of increased T-cell clonal expansion was found, contradicting some individual prior publications finding
  • Harmony batch-effect correction combined with SCCAF-refined, marker/logistic-regression-based cell annotation enables standardized integration and comparison of heterogeneous scRNA-seq COVID-19 datasets method
  • 3 of 10 findings from Ren et al. were reproducible across the compiled datasets finding
Experimental setups
Assay System Perturbation Readout Platform
scRNA-seq (10x Genomics, 5' and 3') human PBMC (Wen, Zhang, Lee, Yu, Jiang, 10x healthy reference) SARS-CoV-2 infection (varying severity stages) cell subpopulation proportions and differential gene expression 10x Chromium, Cell Ranger
scRNA-seq (Seq-Well) human PBMC (Wilk dataset) SARS-CoV-2 infection cell subpopulation proportions and pathway/gene expression changes Seq-Well
scRNA-seq (10x, 3') human bronchoalveolar lavage fluid (Liao, He datasets) SARS-CoV-2 infection (moderate/severe) cell subpopulation proportions, gene expression regulation 10x Chromium, Cell Ranger
scRNA-seq (10x, 3') human nasopharyngeal/bronchial tissue (Chua dataset) SARS-CoV-2 infection (severe) macrophage activation and inflammatory chemokine expression 10x Chromium
TCR repertoire sequencing human PBMC and BALF T/NK cells (Wen, Zhang, Liao) SARS-CoV-2 infection stage (healthy to severe/convalescent) T-cell clonal expansion 10x Chromium VDJ, Cell Ranger
BCR repertoire sequencing human PBMC B cells (Wen, Zhang) SARS-CoV-2 infection stage B-cell clonal expansion 10x Chromium VDJ, Cell Ranger
Batch integration / dimensionality reduction (Harmony, UMAP) combined 9 scRNA-seq datasets (159 samples, 862,354 cells) none (computational integration) cross-study clustering and cell-type mixing Harmony
Cell subtype classification (logistic regression, SCCAF) combined immune cell populations across all datasets none (computational annotation) self-projection accuracy of cell subpopulation labels SCCAF
Key results
  • 8 of 20 (40%) published results reproducible across all relevant datasets vs 16 of 20 (80%) in the original dataset alone 40% vs 80%
  • T-cell proportion decreases from healthy to severe COVID-19 stages across multiple datasets
  • 153 genes consistently upregulated in CD14+ monocytes across three studies, enriched for immune response, response to virus, and type I interferon signaling 153 genes
  • Type I interferon signaling upregulation observed consistently across T cells, NK cells, DC, pDC, B cells, Plasma cells and Neutrophils
  • B-cell clonal expansion increased in COVID-19 patients vs healthy controls, confirmed across studies
  • No consistent cross-study trend of increased T-cell clonal expansion in COVID-19 patients
  • Monocyte and neutrophil proportions increase from healthy to severe stages across datasets ~20% to ~80% (monocytes, Liao dataset)
  • Plasma cell proportion increases from healthy to severe patients
Key statistics
  • count 159 samples, 862,354 cells, 9 disease conditions (total scale of integrated meta-analysis dataset)
  • fold_change 8/20 (40%) reproduced across datasets (cross-dataset reproducibility of published findings)
  • fold_change 16/20 (80%) reproduced in original dataset (within-study reproducibility of published findings)
  • count 37 cell subpopulations identified (final cell-type annotation resolution, excluding doublets)
  • other self-projection accuracy above 92% (SCCAF validation of 5 major cell population annotations in Harmony latent space)
  • count 153 consistently upregulated genes (overlap of upregulated genes in CD14+ monocytes across Zhang, Wilk, Liao studies)
  • pvalue FDR < 0.01 (threshold for differentially expressed genes between severe/moderate and healthy samples)
  • fold_change 3/10 Ren et al. observations reproducible (validation of Ren et al. findings against the 9 compiled datasets)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This meta-analysis reprocessed 9 COVID-19 scRNA-seq datasets (159 samples, 862,354 cells) through a standardized pipeline including Harmony batch correction and reference-based logistic regression cell-type annotation. Differential expression was filtered at FDR < 0.01; pathway overrepresentation was assessed on sets of genes consistently upregulated across studies. Cross-dataset reproducibility of 20 published findings was evaluated qualitatively (binary yes/no per dataset), supplemented by visual inspection of cell-type proportions and TCR/BCR clonal expansion across disease stages.

Replicationbiological Sample size9 datasets comprising 159 samples and 862,354 cells described in total; no formal power analysis for the meta-analysis reported GroupsHealthy vs. Mild, Moderate, Severe, Convalescent, Asymptomatic, Post-Mild, Late Recovery, Influenza across 9 scRNA-seq studies from PBMC, BALF, and nasopharyngeal/bronchial tissue Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionFDR (specific algorithm not stated; 'False Detection Rate' terminology is consistent with Benjamini-Hochberg)
Statistical tests used
Test Applied to n Assumptions
Differential expression analysis (specific test not stated; FDR < 0.01 threshold applied) Gene expression comparison between disease stages (e.g., severe vs. healthy, moderate vs. healthy) per cell type across studies (Supplementary Figs. S10, S11) not stated
Gene set / pathway overrepresentation analysis (method not stated) 153 consistently upregulated genes in CD14+ monocytes across three studies tested for pathway enrichment (Supplementary Fig. S10e) 153 overlapping genes not stated
SCCAF self-projection accuracy (internal cluster validation metric, not a formal hypothesis test) Validation of 5 major cell-type clusters in Harmony latent space; threshold >92% (Supplementary Table S3) na
Visual/descriptive comparison of cell-type proportion changes across disease stages T cells, monocytes, neutrophils, plasma cells, B cells across healthy/mild/moderate/severe (Fig. 3, Supplementary Figs. S8–S9) na
Visual/descriptive comparison of TCR and BCR clonal expansion across disease stages T-cell and B-cell clonal expansion in Wen, Zhang, and Liao datasets (Supplementary Figs. S15–S18) na
Approaches that could also have been used
  • Cell-type proportion changes across disease stages were assessed by visual inspection of bar plots without formal statistical testing
    Could also: A compositional data analysis framework such as scCODA, Dirichlet regression, or beta regression could also be applied to test whether proportions differ significantly across stages — Compositional models account for the sum-to-one constraint of cell-type proportions and yield p-values and uncertainty estimates, complementing the visual summaries and allowing quantification of the consistency noted across studies
  • Cross-dataset reproducibility of 20 published findings was evaluated qualitatively as binary yes/no across datasets
    Could also: A formal quantitative meta-analysis (e.g., random-effects model pooling log fold changes or standardized mean differences per finding across studies) could also be applied — A random-effects meta-analysis would yield pooled effect estimates with confidence intervals and heterogeneity statistics (I², τ²), enabling more precise statements about the magnitude and consistency of each finding across the 9 studies
  • The specific differential expression test and software were not stated in the available text; only the FDR threshold is reported
    Could also: Explicitly documented pseudobulk approaches (e.g., aggregating counts per donor before testing with DESeq2 Wald test or edgeR quasi-likelihood F-test) are also widely used for scRNA-seq DE analysis — Pseudobulk methods account for within-donor correlation across cells, which cell-level tests may not, and full reporting of the DE method and software version allows readers to assess assumptions and reproduce the analysis
  • Disease-stage mapping across studies was performed manually and assumed comparability of similarly labeled stages
    Could also: Anchoring to a validated quantitative severity scale (e.g., WHO ordinal scale) or reporting inter-rater agreement (e.g., Cohen's κ) for stage assignments could also be used — Quantitative severity anchoring reduces ambiguity in stage equivalence and enables sensitivity analyses stratified by severity definition, which would be informative given the heterogeneity the authors explicitly note in cross-dataset comparisons
  • Pathway overrepresentation analysis was applied to the discrete intersection of genes upregulated across all three studies
    Could also: A rank-based gene set enrichment approach (GSEA) or AUCell applied to ranked gene lists from each study separately, with results then aggregated, could also be used — Ranked methods use the full distribution of expression changes rather than a binary overlap threshold, potentially recovering pathway signals from genes that are consistently but modestly upregulated across studies and providing enrichment scores with uncertainty estimates
  • Results are reported without dispersion measures, confidence intervals, or effect sizes for the cell-type proportion and DE findings
    Could also: Reporting 95% confidence intervals around proportion differences or log fold changes, alongside between-study heterogeneity estimates, would also convey the precision and variability of each observed effect — Effect sizes and confidence intervals allow readers to assess both the magnitude and uncertainty of findings, which is especially relevant given the range of cohort sizes (3–36 samples per study) and the study's own emphasis on the caution needed when interpreting small-cohort scRNA-seq results
Software: Cell Ranger 3.1.0 · Harmony · SCCAF · 10× Chromium reference pipeline (GRCh38 v3.0.0; VDJ reference v4.0.0)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34675242

Paper: Garg et al. 2021, Meta-analysis of COVID-19 single-cell studies confirms eight key immune responses. Sci Rep. PMCID PMC8531356. DOI 10.1038/s41598-021-00121-z. Code: https://github.com/Manikgarg/COVID-19 (master, pushed 2022-02-02). Jupyter/Python. Room target dataset: GEO GSE149689 (Lee et al., PBMC scRNA-seq).

What the paper is

A meta-analysis re-processing 9 published scRNA-seq datasets (159 samples, 862,354 cells) through one common pipeline, then checking which of 20 previously-reported immune findings reproduce. Headline: 16/20 (80%) reproduce within their original dataset; 8/20 (40%) reproduce across datasets → the "eight key immune responses."

Datasets the paper relies on (Table 1)

study accession tissue samples cells access
Wen PRJCA002413 PBMC 15 119,448 GSA (open)
Zhang PRJCA002564 PBMC 22 140,588 GSA (open)
Lee (this room) GSE149689 PBMC 20 58,022 GEO open
Yu PRJCA002579 PBMC 10 91,742 GSA (open)
Jiang (not stated) PBMC 12 93,804 unclear
Wilk GSE150728 PBMC 13 67,923 GEO open
Liao GSE145926 BALF 12 63,103 GEO open
He GSE147143 BALF 3 10,927 GEO open
Chua EGAS00001004481 nasoph./bronch. 36 156,025 EGA controlled
10x ref (10x Genomics) PBMC 16 60,772 open

CRITICAL repo finding (determines scope)

The repo is downstream-only. Its 5 scripts/notebooks (AnalyzeTCR_BCR.py, PlotTCR_BCR.ipynb, DE_and_pathway_analysis_one_vs_All.ipynb, HLA_classII_and_ISG_gene_module_score.ipynb, gsea.py, gprofiler_plotting.py) all read a fully pre-processed AnnData objectall11sets_hm.h5 (Harmony-integrated, SCCAF-annotated, with cell_type/stage/severity/sample/study) or per-study <Study>.h5 with VDJ in .uns. That processed object is NOT in the repo and NOT deposited anywhere. The upstream pipeline that BUILDS it (QC → normalize → HVG → Harmony integration → Louvain → SCCAF logistic- regression annotation → MAST DE) is not committed. Only the Methods text describes it.

Consequence: no repo figure can be run end-to-end as-shipped (its input is missing). To reproduce anything we must rebuild the processed object from raw GEO data using the Methods-described pipeline, then apply the repo's downstream logic (which IS shipped, incl. exact gene lists).

IN SCOPE (pipeline-derived, reproducible on the OPEN room dataset GSE149689)

  1. Dataset profile / cell census — download GSE149689, count samples (expect 20) & cells; run the paper's QC (>15% mito; <500 UMI or <200 genes; genes in <3 cells; Scrublet doublets) and report post-QC cell count vs the paper's reported 58,022. Pipeline: scanpy QC.
  2. Major cell-type annotation + proportions (Fig 3) — normalize/log/HVG/PCA/neighbors/Louvain, annotate T/B/NK/Mono(CD14,CD16)/DC/Plasma/Platelet by canonical markers; compare healthy-PBMC T-cell fraction to the paper's reported 30–40% (Lee) and monocyte ~20%. Pipeline: scanpy.
  3. Response #2/#8 — Type I IFN / ISG upregulation — compute the repo's exact ISG module score (sc.tl.score_genes, 99-gene ISG list verbatim from the notebook, ctrl_size=100) per cell; test ISG score higher in COVID vs Normal and severe vs mild, esp. in classical (CD14+) monocytes. This is the paper's central confirmed finding AND GSE149689's own headline (Lee: "type I IFN response in classical monocytes of severe COVID-19"). Method: scanpy + Wilcoxon.
  4. Response #5 — HLA-class II downregulation in CD14+ monocytes — repo's HLA_classII module score (all HLA-D* genes), severe COVID vs healthy in monocytes. Method: scanpy + Wilcoxon.

OUT OF SCOPE (not attempted, with reason)

  • Cross-dataset meta-analysis / the 80% & 40% reproducibility rates — require all 9 datasets, incl. EGA-controlled Chua (EG
Figures / tables: TableFig 3
GSE149689_samples
Reported
20 samples (Table 1, Lee et al.)
Reproduced
20 samples observed (barcode suffix -1..-20)
exact
GSE149689_cells
Reported
58022 cells (Table 1, Lee et al.)
Reproduced
58022 cells after mito<15%/>=500UMI/>=200genes/genes-in->=3-cells QC (pre-Scrublet); 57039 after Scrublet 0.2.1; 85144 raw
exact
response5_HLAII_down_monocytes
Reported
HLA class II downregulation in CD14+ monocytes in severe COVID (response #5, one of 8 key responses)
Reproduced
CD14 monocyte HLA-II module score Healthy 0.80 -> Severe 0.24 (Mann-Whitney p<<1e-300); down in all 6 severe samples; mild ~ healthy
within tolerance
response2_8_typeI_IFN_ISG_up
Reported
Type I IFN/ISG upregulation in COVID, strongest in classical (CD14+) monocytes (responses #2/#8)
Reproduced
CD14 mono ISG: pooled COVID 0.299 vs Healthy 0.286 (p=0.02 UP); mild/asymptomatic elevated; but SEVERE suppressed (median 0.19 vs 0.27, p=2e-98). IFN response present/elevated overall but NOT strongest in severe.
partial
response2_8_ISG_severe_vs_mild
Reported
type I IFN response associated with severe COVID (severity gradient)
Reproduced
severe 0.270 < mild 0.318 (p=1e-149) — OPPOSITE to reported direction; coherent with impaired-IFN-in-severe literature; value derivable from data (no fabrication concern)
did not match
celltype_T_normal
Reported
T cells ~30-40% of healthy PBMC
Reproduced
31.94%
within tolerance
celltype_mono_normal
Reported
monocytes ~20% of healthy PBMC
Reproduced
30.3% (CD14 23.6 + CD16 6.7)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 69/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

76.4 k
tokens (I/O) · 4.1 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.