Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Severe COVID-19 Shares a Common Neutrophil Activation Signature with Other Acute Inflammatory States.

Cells · 2022
L2 88/100 PQI 97
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
88/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 75% of all assessed papers rank 277 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and reproduces 1:1. Target = the tow-sample-tests analysis in github.com/lschimke/COVID19-and-HLH-paper@3b2fbb9 (non-parametric two-sample Satterthwaite t-approx, ICU vs Non-ICU COVID-19). The repo ships input data (my_data2.csv, 503 vars x 102 samples), the unmodified script.R (set.seed(2021)), and the expected output tables. Running the shipped script on the shipped data in a clean conda env (R 4.0.5 = paper's exact version; nparcomp 3.0, MVN 5.9, ggplot2 3.3.6, npmv 2.4.1 from CRAN) on «our HPC» SLURM regenerated Supplementary Tables S1 and S2 BYTE-IDENTICAL (same SHA256, max abs diff 0.0 across all 503 rows of both tables) and the headline claim of 212 significant variables (p<0.05) exactly. FigS1/FigS2 regenerated (visual, not byte-compared). NOT attempted (out of 80/20 scope, recorded in scope.md): the single-cell Seurat pipeline on GSE152418 (partly stochastic, no shipped numeric reference) and the PCA/CCA/MLR/randomForest/heatmap visualisations (ship example data + plotting code but no pinnable expected value; randomForest also stochastic). Minor wording note: paper says '504 observed variables' but 503 are tested (the 504th column is the Group label) — wording off-by-one, not a result discrepancy. No fabrication signal: all compared values fully derivable from the deposited data+code.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 88
    assessed: 2026-06-15 ⛓ 77c0fb2b067d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
👤 1 human curator(s) · Level L2 2026-06-15
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Severe COVID-19 shares an underlying immunopathologic mechanism with other acute hyperinflammatory conditions (e.g., hemophagocytic lymphohistiocytosis, HLH), and this common molecular signature—centered on neutrophil activation—can be identified across transcriptomic and proteomic data and used to stratify COVID-19 severity.

Core claims
  • COVID-19 and HLH share a transcriptional overlap of 239 unique common differentially expressed genes (DEGs), 237 of which are up-regulated finding
  • The shared up-regulated DEGs are enriched for cytokine/chemokine signaling and neutrophil-mediated immune response pathways (activation, degranulation, migration) finding
  • The neutrophil-associated gene cluster shared by COVID-19 and HLH is also dysregulated in Kawasaki disease and bacterial pneumonia, indicating a generalized hyperinflammatory signature rather than one specific to COVID-19/HLH finding
  • These neutrophil/cytokine genes are dysregulated at the protein level across several COVID-19 proteomic studies and form an interconnected PPI network pointing to neutrophil hyperactivation in ICU COVID-19 patients finding
  • scRNAseq analysis shows the signature genes are specifically upregulated across multiple leukocyte populations, including lymphocyte subsets and immature neutrophils finding
  • A random forest/AI model confirms strong association of the identified gene signature with COVID-19 severity (ICU vs non-ICU discrimination) finding
  • The two shared down-regulated genes (GNLY, MYOM2) may reflect a defect in cytotoxic activity typical of HLH, warranting future investigation mechanism
  • Seven fHLH-causing genes (AP3B1, LYST, PRF1, RAB27A, STX11, STXBP2, UNC13D), though not among the common DEGs, interact with the shared DEG network and enrich neutrophil degranulation/NET-related processes mechanism
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (blood, PBL/PBMC) human patients (COVID-19, HLH, healthy controls) none (disease state comparison) differentially expressed genes (DEGs) NetworkAnalyst 3.0 / limma-voom
bulk RNA-seq (nasopharyngeal swab) human patients (COVID-19, HLH) none (disease state comparison) differentially expressed genes (DEGs) NetworkAnalyst 3.0 / limma-voom
single-cell RNA-seq (scRNAseq) human PBMCs (COVID-19 patients, Schulte-Schrepping dataset) none (disease state) cell-type-specific gene expression (UMAP, dotplot, heatmap) Seurat pipeline
quantitative plasma proteomics (LFQ mass spectrometry) human plasma (COVID-19_ICU vs COVID-19_nonICU) none (disease severity comparison) differentially expressed proteins (DEP) LC-MS/MS (Overmyer et al. dataset)
protein-protein interaction (interactome) analysis in silico, derived from human DEGs/DEPs none PPI network connectivity of shared DEGs IID database v2021-05; NAViGaTOR 3.0.14
canonical correlation analysis (CCA) of transcript sets human PBL (GSE157103, COVID-19 and SARS-CoV-2-negative respiratory controls) none correlation between cytokine/chemotaxis and neutrophil gene signatures R packages CCA, whitening
random forest classification / machine learning human transcriptome data (COVID-19_ICU vs COVID-19_nonICU) none (classification task) variable importance (gini decrease, mean minimum depth) for ICU admission prediction; OOB error, ROC R package randomForest v4.6.14
Key results
  • 21,583 total DEGs identified from seven COVID-19 cohorts and one HLH cohort
  • 239 unique common DEGs shared between HLH and all COVID-19 datasets; 237 up-regulated, 2 down-regulated (GNLY, MYOM2)
  • Shared up-regulated DEGs include neutrophil activation genes (CD177, MPO, ELANE) and cytokine/chemokine genes (IL1R2, CXCR2, CCL4, CCL2)
  • Of the common DEGs, 25 assigned to cytokine, 34 to chemotaxis, and 58 to neutrophil-mediated immunity gene signatures
  • Top enriched biological processes among common DEGs include neutrophil activation, degranulation, migration, and IL-1 response
  • The 7 fHLH-causing genes (not themselves DEGs) enrich cellular components (secretory/azurophilic granules) and biological processes (neutrophil degranulation)
Key statistics
  • count 1596 individuals from 11 datasets (total cohort size across meta-analysis)
  • count 21,583 DEGs (total DEGs across seven COVID-19 cohorts and HLH cohort)
  • count 239 unique common DEGs (237 up, 2 down) (shared DEGs between HLH and COVID-19 datasets)
  • count 25 cytokine, 34 chemotaxis, 58 neutrophil DEGs (functional categorization of common DEGs)
  • count 100 COVID-19 patients and 26 SARS-CoV-2-negative respiratory controls (GSE157103 PBL dataset used for CCA analysis)
  • fold_change log2 fold-change > 1 (up) or < -1 (down), adjusted p-value < 0.05 (DEG statistical cut-offs (limma-voom))
  • count 5000 trees, 3 variables resampled per split (random forest model parameters)
  • count 75%/25% train/test split (cross-validation for random forest classifier)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is an integrative bioinformatics/meta-analysis of publicly available transcriptome (bulk and single-cell RNA-seq) and plasma proteome data from 1596 individuals across 11 datasets, comparing COVID-19 with HLH, other inflammatory syndromes, and respiratory infections. Differential expression was computed per dataset with the limma-voom pipeline using fixed fold-change and adjusted p-value cut-offs, and shared genes were intersected and explored through enrichment, protein-protein interaction networks, multivariate correlation (PCA, CCA, multilinear regression), nonparametric multivariate testing of proteins, and a random forest classifier for ICU prediction. Results were reported largely as DEG/DEP lists, enrichment dot plots, heatmaps, network diagrams, and classifier performance (OOB error, ROC) rather than as conventional per-comparison significance tables with dispersion statistics.

Replicationbiological Sample sizeSample size driven by aggregating publicly available datasets; authors followed meta-analysis guidance recommending at least three to four studies and a minimum of 1000 participants to increase statistical power/signal-to-noise; final n = 1596 from 11 datasets GroupsCOVID-19 (ICU vs nonICU and vs controls), HLH, MIS-C, Kawasaki disease, seasonal coronavirus, influenza, bacterial pneumonia, and healthy controls Pairingunpaired Randomization/blindingna Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionAdjusted p-value < 0.05 applied to differential expression (limma-voom default is Benjamini-Hochberg FDR, though the specific method is not explicitly named); no multiplicity adjustment described for the multivariate/correlation analyses
Statistical tests used
Test Applied to n Assumptions
limma-voom differential expression (moderated linear model on log2 CPM) DEGs between groups within each transcriptome dataset (e.g., COVID-19 vs control, HLH vs control) per-dataset sample sizes from 11 datasets totaling 1596 individuals; e.g., GSE157103 with 100 COVID-19 and 26 controls not stated
Seurat differential expression for single-cell RNA-seq scRNAseq dataset EGAS00001004571 (Schulte-Schrepping et al.) across leukocyte populations not stated
Nonparametric MANOVA (multivariate analysis of variance) differences in plasma protein expression between COVID-19_ICU and COVID-19_nonICU (Overmyer et al. proteomics) na
Nonparametric inference for multivariate data (npmv / nparcomp) follow-up comparison of differentially expressed proteins between COVID-19_ICU and nonICU na
Canonical Correlation Analysis (CCA) linear relationship between cytokine/chemotaxis and neutrophil-mediated immunity gene signatures (COVID-19_PBL, GSE157103) 100 COVID-19 and 26 control individuals not stated
Principal Component Analysis (PCA) and multilinear regression 25 transcripts associated with COVID-19 severity; combinations of genes not stated
Approaches that could also have been used
  • DEGs were defined using fixed cut-offs of |log2 fold-change| > 1 together with adjusted p-value < 0.05.
    Could also: A ranked or continuous approach (e.g., reporting the full statistic with FDR and using gene-set enrichment such as GSEA on ranked lists rather than a hard fold-change threshold) could also be used. — Continuous/ranked methods retain genes with moderate but coherent changes and reduce dependence on an arbitrary fold-change boundary, which can complement the threshold-based list.
  • Shared genes were identified by intersecting per-dataset DEG lists (Venn/Circos) across studies generated on different platforms.
    Could also: A formal meta-analysis model (e.g., random-effects effect-size combination across studies, or batch-aware joint modeling) could also be applied. — A pooled statistical model provides combined effect estimates with confidence intervals and explicitly accounts for between-study heterogeneity, complementing the overlap-counting approach.
  • Multiplicity was handled per dataset via an adjusted p-value threshold, without a stated correction across the multivariate, correlation, and CCA analyses.
    Could also: Reporting the specific adjustment method (e.g., Benjamini-Hochberg FDR) and extending multiplicity control to the correlation/multivariate family could also be done. — Naming the method and covering all test families makes the family-wise or false-discovery scope fully transparent across the analyses.
  • Plasma protein differences between ICU and nonICU groups were assessed with nonparametric MANOVA and nonparametric multivariate inference.
    Could also: Per-protein moderated tests (e.g., limma on proteomics) with FDR control, accompanied by effect sizes and confidence intervals, could also be reported. — Per-protein estimates with intervals convey the magnitude and precision for individual proteins alongside the global multivariate result.
  • The random forest ICU classifier was evaluated with a single 75/25 train/test split, OOB error, and ROC.
    Could also: Repeated k-fold or nested cross-validation with confidence intervals on AUC could also be used. — Resampling-based validation gives a more stable performance estimate and uncertainty bounds that are less sensitive to a single partition.
  • Group summaries were displayed via box plots and heatmaps without a specified dispersion statistic.
    Could also: Explicitly reporting SD, IQR, or 95% confidence intervals alongside the plots could also be done. — Stating the dispersion measure makes the spread and uncertainty of each group directly interpretable, which is especially helpful for smaller per-group n.
Software: NetworkAnalyst 3.0 (limma-voom pipeline) 3.0 · R 4.0.5 · RStudio 1.4.1106 · Seurat · NAViGaTOR 3.0.14 · Integrated Interactions Database (IID) 2021-05 · R/randomForest 4.6.14 · ClusterProfiler / Enrichr / Morpheus / circlize / ggpubr / ggplot2 / factoextra / CCA / whitening / npmv / nparcomp

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
51
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE152075 GEO in Table (http://semanticscience.org/resource/SIO_000419)
also used by 2 papers:
EGAS00001004571 EGA in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE152641 GEO in Table (http://semanticscience.org/resource/SIO_000419)
also used by 1 paper:
GSE157103 GEO in Table (http://semanticscience.org/resource/SIO_000419)
also used by 1 paper:
GSE163151 GEO in Table (http://semanticscience.org/resource/SIO_000419)
also used by 1 paper:
GSE152418 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
GSE156063 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
GSE161731 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
GSE178388 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
GSE26050 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
GSE73461 GEO in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Reproduction scope — pmid-35269470

Paper: Schimke LF et al. (2022) Severe COVID-19 Shares a Common Neutrophil Activation Signature with Other Acute Inflammatory States. Cells 11(5):847. DOI 10.3390/cells11050847 · PMCID PMC8909161.

Code: https://github.com/lschimke/COVID19-and-HLH-paper @ commit 3b2fbb9bc74324fb9af93a858d7de4afc4d67d8b (HEAD of main, pushed 2021-07-30). Public, no license file (re-use for reproduction only; not redistributed).

Data: the repo is organized as one folder per analysis method; each folder ships its own input data + scripts + (mostly) expected output figures/tables. The GEO accession GSE152418 is used only by the single-cell folder.

Repo structure (analysis units)

Folder Method Ships input Ships expected output Determinism
tow-sample-tests non-param. MANOVA + Satterthwaite two-sample t-approx (npmv, nparcomp) ✅ my_data2.csv (503 vars, ICU/Non-ICU) TabS1.csv, TabS2.csv, FigS1/S2 set.seed(2021), deterministic tests
PCA PCA + ggplot ✅ example datafile PCA.txt ✖ (figure only, no value) deterministic
random forest randomForest classifier ✅ my_data.csv ✖ (no shipped metric) ✖ stochastic (RF)
CCA canonical correlation (CCApackage) ✅ demo csv/xlsx deterministic
Multilinear regression MLR (2- & 3-group) ✅ example txt deterministic
Correlograms / Circular Heatmaps / DotPlots / BoxPlot ComplexHeatmap / ggplot visualisations ✅ example txt/xlsx ✖ (plots only) deterministic
single cell analysis Seurat 10x PBMC pipeline on GSE152418 ✖ (needs GEO download) partly stochastic (UMAP/clustering)

In scope (attempted)

Primary — tow-sample-tests. This is the cleanest reproducible unit: it ships the exact input data, the exact scripts (with a fixed seed), AND the exact expected output tables (TabS1.csv = full two-sample-test result for all 503 variables; TabS2.csv = log2 fold-change + p-value). The deterministic statistical tests (npmv::nonpartest, nparcomp::npar.t.test method t.app) should regenerate these tables 1:1. Pinnable claim from the repo README/TabS1: "From the 504 observed variables only 212 had significantly mean differences between ICU and Non-ICU patients." (Tested variables = 503; significant at p<0.05 = 212.)

Pipeline: R 4.0.5reshape2/dplyr reshape → MVN:::descriptivesnpmv::nonpartest (NP-MANOVA) → per-variable nparcomp::npar.t.test → subset p<0.05 → log2 fold change → ggplot2 FigS1/FigS2 + write.table TabS1/S2.

Out of scope (not attempted) — and why

  • Single-cell (GSE152418, Seurat): main figures of the paper, but the pipeline is partly stochastic (UMAP/Louvain seeds, package-version sensitive) and ships no expected numeric reference in the repo → not 1:1 verifiable in the 80/20 budget. Recorded as not-attempted, not a drop.
  • PCA / CCA / MLR / heatmaps / boxplots / dotplots / random forest: ship example data and plotting code but no shipped expected value/figure with a pinnable number to compare against (random forest is additionally stochastic). Visual-only reproductions are not auditable as exact/within-tol.
  • All wet-lab / clinical / proteomic measurements (the upstream generation of the 503-variable matrix): external, not pipeline-derived → out of scope.

Faithfulness notes

  • R version pinned to the paper's 4.0.5 in the conda env where it solves; numeric outputs of these tests are version-robust (3-decimal rounding).
  • The shipped TabS1/TabS2 are treated as the expected reference; we run the shipped script on the shipped input in a fresh copy and compare byte/numeric.
  • All heavy steps run on «our HPC» (SLURM); «host» holds results only.
Figures / tables: FigS1tablesFigS2figures
C1
Reported
212 significant variables (p<0.05) of ~504 observed (503 tested)
Reproduced
212 of 503
exact
C2
Reported
Supplementary Table S1 (503 rows: rel.effect, 95% CI, T, p)
Reproduced
503 rows byte-identical (sha256 53f3f065...2321)
exact
C3
Reported
Supplementary Table S2 (503 rows: log2 fold change, p)
Reproduced
503 rows byte-identical (sha256 92d91e71...dc65)
exact
C4
Reported
Supplementary Figures S1/S2 (relative-effect CIs; log2FC volcano)
Reproduced
FigS1/FigS2 regenerated (png+pdf) from shipped script
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 88/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

Re-running the authors' own shipped script.R on their deposited my_data2.csv in the paper's exact R 4.0.5 environment regenerated Supplementary Tables S1 and S2 byte-identically and the headline claim of 212 significant variables (p<0.05) exactly. There is no deviation on any compared value, no fabrication signal, and the data/code are on the authors' side and fully sufficient. The only blemish is cosmetic — the paper says '504 observed variables' while 503 are tested (the 504th column is the Group label) — an off-by-one in wording. Note the reproduction was scoped (80/20) to the two-sample-tests target; the stochastic single-cell Seurat and PCA/RF visualizations were not attempted, but the scoped claims reproduce 1:1.

👤 Schlein Lab (curation team) L2 100/100
🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

136.6 k
tokens (I/O) · 11.8 M incl. cache
16 min
runtime · 0.02 CPU-h
3.3 GB
peak RAM
1
HPC jobs
hummel
machine