Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Single-cell RNA-sequencing highlights a curtailed NK cell function in convalescent COVID-19 pregnant women.

Front Immunol · 2025
L1 87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough, and reproduced 1:1 on the headline + structural results. The Zenodo deposit ships a single processed Seurat v5 object (obj.Rds, md5 8bfdd79f... = matches the Zenodo record). The only 'code' link is upstream CellRanger (no FASTQ shipped) so the in-scope reproduction is the downstream pipeline output from obj.Rds, run with Seurat 5.5.0 on «our HPC» SLURM. EXACT: 30,394 cells, 8 sample libraries + 2 groups, and 25 clusters all match the shipped object exactly. HEADLINE biological claim ('curtailed NK function' = cytotoxic-gene downregulation): I identified NK clusters independently by canonical markers (CD3D<1 & CD3E<2 & NKG7>10 -> cluster 7 = CD16+ NK, cluster 18 = CD56bright NK) since no cell-type annotation is shipped, then ran Seurat DE PregRCOVID vs PregHC: all 7 named genes are downregulated and 6/7 are significant (padj<0.05); GZMH is downregulated but not significant. Result is robust to the NK-set choice (adding the CD3-intermediate cluster 22 barely changes it). NOT ATTEMPTED / NOT REPRODUCIBLE: (a) the exact NK I/II/III subtype naming and their individual up/down proportion directions - the cluster->cell-type mapping is NOT in the object (hard ~20%); overall NK fraction did decrease but I could not reproduce a NK subset increasing; (b) CellRanger FASTQ->counts (raw reads not deposited); (c) flow-cytometry validation (wet-lab); (d) the 55,588/46,594 pre-final-filter counts (object is already post-filter). No fabrication concerns: every reported structural number is exactly recoverable from the shipped data, and the headline DE direction reproduces under an independent NK definition.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.14066080

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 87
    assessed: 2026-06-14 ⛓ d39bec0de3f7
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Whether SARS-CoV-2 infection 'rewires' the maternal immune response during active infection and whether it 'resets' back to normal during the convalescent/recovery phase in pregnant women.

Core claims
  • NK cells show a sustained reduction during active SARS-CoV-2 infection and persisting after recovery in pregnant women finding
  • scRNA-seq shows SARS-CoV-2 infection rewires gene expression profiles of NK cells, monocytes, CD4+ and CD8+ effector T cells, and antibody-producing B cells in convalescent pregnant women finding
  • Gene pathways for cytotoxic function, type I & II interferon signalling, and pro-/anti-inflammatory responses in NK and CD8+ cytotoxic T cells are attenuated in recovered pregnant women compared with healthy pregnancies finding
  • NK cells from convalescent pregnant women show diminished levels of the cytotoxic proteins perforin, CD122 and granzyme B, validating the scRNA-seq findings finding
  • Two independent geographical cohorts (Malaysia discovery, Germany validation) were immunophenotyped using a 14-colour flow cytometry panel method
  • scRNA-seq was performed using a Chromium Single Cell 3' Gel Bead Chip and Library Kit (10x Genomics, Drop-seq method) method
  • SARS-CoV-2 infection deranges the adaptive immune response in pregnant women even after recovery and may contribute to post-COVID-19 sequelae mechanism
Experimental setups
Assay System Perturbation Readout Platform
14-color flow cytometry immunophenotyping PBMCs from pregnant women (Preg-HC, Preg-INF, Preg-R; Malaysia and Germany cohorts) SARS-CoV-2 infection / convalescence proportions of immune subsets (NK, CD4+, CD8+, monocytes, B cells, NKT, Tregs) and CD8+ T cell differentiation subsets BD LSRFortessa Cell Analyzer; FlowJo analysis
single-cell RNA sequencing (scRNA-seq) PBMCs from convalescent pregnant women vs healthy pregnancies SARS-CoV-2 infection (convalescent phase) gene expression profiles and pathway activity in NK, monocytes, CD4+, CD8+ effector T cells, and B cells 10x Genomics Chromium Single Cell 3' Gel Bead Chip and Library Kit
13-plex bead-based cytokine/chemokine assay serum from pregnant women (validation cohort) SARS-CoV-2 infection/recovery concentrations (pg/mL) of IL-1β, IFN-α2, IFN-γ, TNF-α, MCP-1, IL-6, IL-8, IL-10, IL-12p70, IL-17A, IL-18, IL-23, IL-33 LEGENDplex Human Inflammation Panel-1; BD LSRFortessa Cell Analyzer
3-plex bead-based serological IgG assay serum from Preg-R and Preg-HC women (validation cohort) SARS-CoV-2 infection/recovery IgG antibody levels against Spike S1, Nucleocapsid, and Spike RBD LEGENDplex SARS-CoV-2 Serological IgG Panel
flow cytometry (surface and intracellular staining) NK cells from PBMCs of convalescent pregnant women SARS-CoV-2 infection/recovery cytotoxic protein levels (perforin, CD122, granzyme B) BD LSRFortessa Cell Analyzer
Key results
  • NK cell frequency is sustainedly reduced during active infection and after recovery from SARS-CoV-2 in pregnant women
  • scRNA-seq reveals rewired gene expression profiles in NK, monocytes, CD4+, CD8+ effector T cells and B cells of convalescent pregnant women
  • Cytotoxic function, type I/II interferon signalling, and pro-/anti-inflammatory pathways are attenuated in NK and CD8+ T cells of recovered vs healthy pregnant women
  • Perforin, CD122 and granzyme B protein levels are diminished in NK cells of convalescent pregnant women
  • Proportions of naïve, intermediate effector memory (EM II), and late effector memory (EM III) CD8+ T cell subsets differ significantly among Preg-HC, Preg-INF, and Preg-R groups p ≤0.05
Key statistics
  • count N = 19 (Preg-HC=11, Preg-INF=4, Preg-R=4) (Cohort 1 (Malaysia, discovery) composition)
  • count Healthy controls N=20; recovered N=34 (cytokine/antibody); recovered N=14 (immunophenotyping) (Cohort 2 (Germany, validation) composition)
  • pvalue p ≤0.05 (significance threshold for Kruskal-Wallis test with Dunn's multiple comparisons across Preg-HC, Preg-INF, Preg-R groups)
  • other >7.1 million deaths; >778 million infected (as of April 2025) (global SARS-CoV-2 pandemic burden, background statistic)
  • other 10%-30% (proportion of individuals developing post-COVID-19/Long COVID condition following infection during pregnancy)
  • other almost 6 in 100 individuals (estimated proportion suffering Long COVID/post-COVID-19 syndrome among ~778 million reported cases)
  • count 1.0 × 10^6 live cells for 14-color FACS panel; 0.5 × 10^6 cells for scRNA-seq (cell input amounts per assay from thawed PBMCs)
  • count 100,000-200,000 cells acquired per sample (flow cytometry acquisition per PBMC sample)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used two independent geographical cohorts (discovery: Malaysia, N=19; validation: Germany, N up to 54) of pregnant women stratified into healthy controls, SARS-CoV-2 infected, and recovered groups. Immune profiling combined 14-colour flow cytometry with UMAP-based unsupervised clustering, scRNA-seq (10x Genomics Chromium), multiplexed cytokine bead assays, and serological IgG quantification. Group comparisons for flow cytometry data were performed with the Kruskal–Wallis test followed by Dunn's post-hoc correction, and results were displayed as violin plots with individual data points; statistical methods for scRNA-seq differential expression are not described in the provided text excerpt.

Replicationbiological Sample sizeGroup sizes stated per cohort: Cohort 1 Preg-HC N=11, Preg-INF N=4, Preg-R N=4; Cohort 2 HC N=20, Preg-R N=34 (cytokine/antibody) or N=14 (immunophenotyping); no formal power calculation described GroupsHealthy pregnant controls (Preg-HC) vs. SARS-CoV-2 infected pregnant (Preg-INF) vs. recovered pregnant (Preg-R); in Cohort 2 HC vs. recovered only Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionDunn's multiple comparisons test (post-hoc following Kruskal–Wallis)
Statistical tests used
Test Applied to n Assumptions
Kruskal–Wallis test with Dunn's multiple comparisons correction Percentages of CD8+ T cell subsets (Naïve, intermediate EM II, late EM III) across Preg-HC, Preg-INF, and Preg-R groups (Figure 1h) Preg-HC N=11, Preg-INF N=4, Preg-R N=4 (Cohort 1 discovery cohort) not stated
UMAP dimensionality reduction with unsupervised clustering (FlowJo) 14-colour flow cytometry data for immune cell subset identification (Figures 1b–1f) na
scRNA-seq differential expression analysis (specific test not stated in provided text) Gene expression comparisons across NK cells, monocytes, CD4+, CD8+ T cells, B cells between groups not stated
Approaches that could also have been used
  • Group comparisons used the Kruskal–Wallis test with Dunn's post-hoc correction across three groups
    Could also: A one-way ANOVA with Tukey's HSD post-hoc test could also be used if data distribution and variance homogeneity were verified — When normality holds, the parametric ANOVA framework has greater statistical power; Tukey's HSD provides explicit control of the family-wise error rate across all pairwise contrasts, making the correction scope transparent
  • Discovery cohort group sizes are very small (Preg-INF and Preg-R each N=4), yet the same parametric-style threshold (p ≤ 0.05) is applied
    Could also: Permutation-based or exact tests (e.g., exact Kruskal–Wallis, or pairwise Mann–Whitney U with Bonferroni adjustment) could also be used at such small n — Exact methods do not rely on large-sample asymptotic approximations, which can be unreliable when group sizes fall below ~5; reporting this explicitly helps readers gauge the reliability of the p-values
  • Violin plots are used to display flow cytometry percentage data, and individual dots represent samples
    Could also: Dot plots overlaid with a median-and-IQR bar, or strip plots with a box plot, could also be used for groups of n=4–11 — Violin plots estimate a density kernel, which can suggest smoother distributions than warranted at n=4; a box-and-dot overlay directly shows each value and the median without distributional assumptions, which is often preferred for very small n
  • Demographic continuous variables (age, gestational age) are summarized as mean ± SD
    Could also: Median with IQR or range could also be reported, particularly for small, potentially non-normally distributed samples — With group sizes of 4–14, the mean and SD can be heavily influenced by a single outlier; median and IQR are robust summaries and align with the non-parametric inferential tests used for the primary outcomes
  • The specific differential expression method for scRNA-seq is not described in the provided text
    Could also: Standard scRNA-seq workflows commonly use Seurat with Wilcoxon rank-sum tests or DESeq2 (pseudo-bulk) with negative-binomial Wald tests, with Benjamini–Hochberg FDR correction — Pseudo-bulk DESeq2 aggregates within-donor counts before testing, which better accounts for within-donor correlation and reduces inflation of significant genes compared to cell-level testing; explicitly stating the method and FDR threshold aids reproducibility
  • The two-cohort replication design (Malaysia discovery, Germany validation) was used to support generalizability
    Could also: A formal meta-analytic or pooled mixed-effects model with cohort as a random or fixed effect could also be applied to jointly estimate effect sizes across cohorts — Pooling cohorts in a mixed model would increase power, allow direct estimation of between-cohort heterogeneity, and yield a single effect-size estimate with a confidence interval rather than requiring qualitative agreement between two separate analyses
Software: FlowJo · 10x Genomics Chromium (scRNA-seq platform/pipeline) · BioLegend LEGENDplex Data Analysis Software · BD LSRFortessa Cell Analyzer (acquisition)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

10.5281/zenodo.14066080 DOI in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40661959

Paper: Salker et al. 2025, Front Immunol — "Single-cell RNA-sequencing highlights a curtailed NK cell function in convalescent COVID-19 pregnant women." DOI 10.3389/fimmu.2025.1560391.

Code link: github.com/10XGenomics/cellranger (third-party upstream aligner — NOT authors' own analysis code). Data: Zenodo 10.5281/zenodo.14066080 = a single file obj.Rds (406 MB, md5 8bfdd79ff35de4a01eec37377ccfb2a1) = the processed Seurat v5 object (30,394 cells, normalized, integrated, clustered).

What the shipped artifact actually contains

obj.Rds is a post-QC, CCA-integrated, clustered Seurat object. meta.data columns: orig.ident, nCount_RNA, nFeature_RNA, Patient (PregHC/PregRCOVID), Type (a1..a4,b5..b8), mitoPercent, percent.ribo, RNA_snn_res.0.8 (20), seurat_clusters (25), cca_clusters (25). There is NO cell-type-name column — clusters are numeric only; the "NK I/II/III", "T cell subset", etc. labels from the figures are NOT shipped.

In scope (pipeline-derived, reproducible from obj.Rds)

  • Final cell count (30,394 high-quality cells) — directly counted.
  • Sample / group structure (8 libraries; 2 groups HC vs post-COVID) — from metadata.
  • Number of clusters / immune subtypes (25) — from cca_clusters.
  • NK cytotoxic-gene downregulation in Preg-R (PRF1, GZMA, GZMB, GZMH, KLRD1, NKG7, IRF1) — NK clusters identified by canonical markers, then Seurat DE PregRCOVID vs PregHC. This is the paper's HEADLINE biological claim ("curtailed NK function").
  • NK proportion shift — fraction of NK cells per group (partial: see below).

Out of scope (not attempted, with reason)

  • CellRanger FASTQ→counts — raw FASTQ are NOT in the Zenodo deposit (only the processed object). The upstream alignment cannot be re-run; it is also not the pipeline step the biological claims rest on.
  • Exact NK I / NK II / NK III subtype naming + their individual proportion directions — the cluster→cell-type annotation is NOT shipped in obj.Rds, so the authors' exact NK-subset mapping cannot be reproduced 1:1 (the hard ~20%).
  • Flow-cytometry validation (perforin, IL-15RB/CD122) — wet-lab, non-pipeline.
  • Intermediate counts 55,588 (recovered) and 46,594 (after first filter) — these are pre-final-filter; the shipped object is already post-filter, so they are not derivable from it.

Pipeline used to reproduce

Seurat 5.5.0 / SeuratObject 5.4.0 / R 4.4.3 (conda, on «our HPC» SLURM). Authors used CellRanger v3.0.1 (GRCh38 v3.0.0) + Seurat 5.2.0, standard workflow, CCA integration.

Figures / tables: figure
C1
Reported
30,394 high-quality single cells
Reproduced
30394 (ncol obj.Rds)
exact
C2
Reported
8 samples (N=4 Preg-HC + N=2 Preg-R run in duplicate)
Reproduced
8 libraries (a1-a4,b5-b8); 2 groups PregHC=16159/PregRCOVID=14235
exact
C3
Reported
25 immune cell subtypes
Reproduced
25 cca_clusters (0-24)
exact
C4
Reported
PRF1,GZMA,GZMB,GZMH,KLRD1,NKG7,IRF1 significantly downregulated in Preg-R NK cells
Reproduced
7/7 negative log2FC; 6/7 padj<0.05 (PRF1 -0.70,GZMB -0.50,GZMA -0.46,KLRD1 -0.33,NKG7 -0.29,IRF1 -0.36); GZMH -0.25 down but ns
within tolerance
C5
Reported
NK I cells decreased, NK II cell subtypes increased in Preg-R
Reproduced
overall NK fraction decreased 6.89%->5.63%; subset increase not observed; NK I/II/III labels not shipped
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

Structural claims (30,394 cells, 8 libraries/2 groups, 25 clusters) reproduce exactly from the md5-verified Zenodo object, and the headline biological claim of curtailed NK cytotoxic function reproduces strongly (7/7 genes down, 6/7 padj<0.05) under an independent NK definition — so the central conclusion holds. The deviations are on the data-availability side, not a defect in the reported values: the authors shipped no cell-type annotation (so the NK I/II/III subset-shift direction, C5, can't be reproduced 1:1) and no pre-filter matrices (so the 55,588/46,594 counts, C6, are uncheckable). Severity is moderate — the secondary NK-subset 'increase' was not observed and GZMH was down but ns (padj=0.97), but magnitude and direction of the core result hold. No fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

121 k
tokens (I/O) · 7.3 M incl. cache
15 min
runtime · 0.02 CPU-h
5.5 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine