In vivo antiviral host transcriptional response to SARS-CoV-2 by viral load, sex, and age.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce 1:1. GSE152075 deposits the raw count matrix + per-sample metadata, and the authors' R repo (greninger-lab/COVID19-RNAseq @9f0be99) runs DESeq2 directly on it; no fastq re-alignment needed (raw reads withheld for privacy, as stated). Five reported pipeline-derived results reproduced via the authors' exact design formulas/prefilters/contrasts on «our HPC»: pos-vs-neg DEGs 84 vs 83; ACE2 viral-load medians (Mid 3.46/High 7.82 exact); high-vs-low viral-load DEGs 369 vs 363; age-effect CXCL11 29.7-fold (paper 30) and PCGF6 15.9-fold (paper 17); sex x infection 20 genes vs 19 with all three named genes (IGHG1, MS4A1, SLAMF6) down-regulated. All within tolerance or exact. Deviations explained by DESeq2 1.50.2 (paper used 1.28.1; exact pin did not solve) and the GEO age field carrying 445 numeric ages vs the 467 implied by the repo prefilter. No fabrication concern: every headline value is derivable from the shipped counts+code and reproduces independently. NOT attempted (out of scope): upstream kallisto alignment (no raw fastq), GSEA/GO/CIBERSORTx (external desktop/web tools + intermediate csvs), and the separate HAE (GSE154768) and longitudinal (GSE154769) accessions outside the brief's named GSE152075.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 89assessed: 2026-06-18 ⛓ e27e5e7b7734
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhat host transcriptional responses underlie the wide range of SARS-CoV-2 clinical manifestations, and how do these antiviral responses vary by viral load, infection time course, sex, and age?
- ★ SARS-CoV-2 infection induces a strong interferon-mediated antiviral response in the nasopharynx, up-regulating antiviral factors (OAS1-3, IFIT1-3, MX2, RSAD2, HERC5) and Th1 chemokines CXCL9/10/11, while down-regulating ribosomal protein transcripts. finding
- ★ Expression of interferon-responsive genes, including ACE2, increases as a function of viral load, while B cell–specific transcripts and neutrophil chemokines are elevated in low viral load samples. finding
- ★ Older individuals have reduced expression of Th1 chemokines CXCL9/10/11, their receptor CXCR3, CD8A, and granzyme B, suggesting deficiencies in cytotoxic T cell and NK cell trafficking/function. finding
- ★ Males have reduced B cell– and NK cell–specific transcripts and increased inhibitors of NF-κB signaling relative to females, possibly throttling antiviral responses. finding
- ★ Human airway epithelial (HAE) cultures recapitulate the cell-intrinsic antiviral response seen in NP swabs at 7 days post infection, with no interferon induction at 3 days, defining a consensus set of 19 antiviral genes. finding
- ★ Longitudinal patient-matched samples show reduced viral load over time accompanied by reduced interferon-induced transcription, recovery of ribosomal protein expression, and initiation of wound healing and humoral immune responses. finding
- Shotgun RNA sequencing of diagnostic NP swabs enables simultaneous recovery of viral genomes and in situ host response profiling. method
- SARS-CoV-2 may exploit host antiviral responses by inducing ACE2 (an interferon-regulated gene) upon interferon exposure. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| shotgun/bulk RNA-seq (metagenomic) | nasopharyngeal swabs from 430 SARS-CoV-2 PCR-positive and 54 negative human individuals | natural SARS-CoV-2 infection (none experimental) | host gene expression / differentially expressed genes | — |
| bulk RNA-seq | male human airway epithelial (HAE; tracheal/bronchial) cultures | SARS-CoV-2 infection (3 and 7 days post infection) | host gene expression changes and percent viral reads | — |
| bulk RNA-seq (longitudinal) | patient-matched nasopharyngeal swabs (n=3), mean 6.3 days apart | natural infection over time course | changes in viral load (Ct) and host gene expression | — |
| RT-PCR (diagnostic) | nasopharyngeal swabs | none | SARS-CoV-2 N1 cycle threshold (viral load) | — |
| in silico cell-type deconvolution (CIBERSORTx) | NP swab RNA-seq (high vs low viral load) | none | immune cell type proportions | CIBERSORTx |
| Gene Set Enrichment Analysis (GSEA) / GO / DisGeNET | NP swab and HAE RNA-seq data | none | enriched pathways and disease ontology terms | MSigDB Hallmark Gene Sets, Gene Ontology, DisGeNET |
- – 83 DE genes between SARS-CoV-2 positive and negative samples (41 up, 42 down) 83 genes (padj<0.1, |log2FC|>1)
- ▲ ACE2 expression increases with viral load (median counts negative/low/medium/high = 0, 1.93, 3.45, 7.82) 0→7.82 median counts; p=7.46×10^-13
- – 363 DE genes between high vs low viral load; high enriched for CXCL9/10, IDO1, CD80; low enriched for neutrophil (CXCL8, S100A9) and B cell transcripts (IGHG1, IGHM, CD22) 363 genes (padj<0.1)
- – Low viral load samples have higher proportions of naïve B/T cells, neutrophils, M2 macrophages; high viral load has more M1 macrophages, activated NK cells, activated DCs 3.5-, 2.2-, 1.6-, 1.8-fold (low); 2.5-, 1.6-, 1.6-fold (high)
- ▲ Consensus set of 19 up-regulated antiviral genes shared across positive-vs-negative, high-vs-low, and infected HAE analyses 19 genes; p=4.81×10^-69
- ▲ HAE show no interferon response at 3 days post infection (virus 0.3% of reads) despite 10-fold higher dose; antiviral response by 7 days (virus 5.3% of reads) 0.3%→5.3% viral reads
- ▼ Longitudinal samples show 39-fold reduction in viral load and increased humoral/wound-healing genes (C1QA/B/C, HLA-DQB1, APOE, CD36, RHOC) with recovery of ribosomal proteins mean Ct increase 5.29 = 39-fold reduction
- ▼ TMPRSS2 reduced upon infection but not modulated by viral dose; ribosomal proteins (RPL4, RPS6) not modulated by viral load
- pvalue 7.46 × 10^-13 (ACE2 expression across viral load groups, Kruskal-Wallis one-way ANOVA)
- pvalue 4.81 × 10^-69 (SuperExact Test for 19-gene consensus antiviral set overlap)
- count 83 (DE genes positive vs negative (41 up, 42 down))
- count 363 (total DE genes high vs low viral load (padj<0.1))
- fold_change 39-fold reduction (viral load reduction over time in longitudinal samples (mean Ct increase 5.29))
- fold_change 3.5-, 2.2-, 1.6-, 1.8-fold (increased naïve B/T cells, neutrophils, M2 macrophages in low viral load (CIBERSORTx))
- count 430 positive, 54 negative (NP swab samples analyzed; median 1.61×10^6 human reads pseudoaligned)
- other 0.3% and 5.3% (SARS-CoV-2 reads in HAE at day 3 and day 7 post infection)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This cross-sectional observational RNA-seq study compared nasopharyngeal host transcriptomes from 430 RT-PCR-confirmed SARS-CoV-2-positive individuals and 54 negative controls, with stratified analyses by viral load (Ct-defined low/medium/high), sex, and age. Transcriptome-wide differential expression was assessed using FDR-adjusted p-values (padj < 0.1) and log2 fold-change thresholds, and pathway enrichment was evaluated with GSEA (FDR < 0.05). Individual gene-level comparisons across viral load groups used the Kruskal-Wallis test and Mann-Whitney U test; immune cell-type proportion estimates (CIBERSORTx) were compared by t-test. A small paired longitudinal sub-analysis (n = 3) and an in vitro HAE cell experiment supplemented the main cross-sectional analysis.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| FDR-adjusted differential expression analysis (padj < 0.1, |log2FC| > 1) | SARS-CoV-2 positive vs. negative (Figs 1A–1B) and high vs. low viral load (Fig 2B) | 430 positive vs. 54 negative; 108 high vs. 99 low viral load | not stated |
| Kruskal-Wallis one-way ANOVA | ACE2 and select interferon-stimulated gene expression across negative, low, medium, and high viral load groups (Fig 2A) | 54 negative, 99 low, 206 medium, 108 high | not stated |
| Mann-Whitney U test | Individual gene expression comparisons between low and high viral load groups (Figs 2A, 2D) | n = 99 low, n = 108 high | not stated |
| t-test (type not further specified) | CIBERSORTx-estimated immune cell-type proportions between viral load groups (Fig 2C) | n = 99 low, n = 108 high | not stated |
| Gene Set Enrichment Analysis (GSEA) with FDR correction | Hallmark Gene Sets and GO Biological Processes enrichment in SARS-CoV-2 positive vs. negative (Fig 1C, S1 Figs) | 430 positive, 54 negative | not stated |
| SuperExact Test (multi-set intersection probability) | Overlap of DE gene lists across positive vs. negative, high vs. low viral load, and HAE day-7 analyses (Fig 3A) | null | not stated |
-
Multiple individual Mann-Whitney U tests were applied to selected genes for the high-vs-low viral load comparison without an explicitly stated multiplicity correction across those individual tests↳ Could also: A transcriptome-wide model (e.g., negative-binomial GLM as in DESeq2 or edgeR) with viral load as a covariate and Benjamini-Hochberg FDR correction across all genes would also address these comparisons — the paper already applies an analogous approach for the positive-vs-negative contrast — A unified model applied consistently across all genes simultaneously controls the false discovery rate across the full gene set and can yield more precise effect-size estimates than a subset of selected-gene tests
-
Viral load, age, and sex effects were characterized in separate stratified analyses rather than in a single multivariable model↳ Could also: A multivariable linear model or generalized linear model incorporating viral load, age, and sex simultaneously as covariates (e.g., within a DESeq2 design formula) would also estimate each factor's contribution to gene expression — Because viral load, age, and sex may be correlated in the sample, a joint model can partially separate their independent associations with the transcriptional response
-
CIBERSORTx-estimated immune cell-type proportions were compared between viral load groups using a t-test↳ Could also: A Wilcoxon rank-sum (Mann-Whitney U) test, permutation test, or Dirichlet regression would also be applicable for deconvolved cell-type proportions — Deconvolved proportions are bounded between 0 and 1 and are often non-normally distributed and compositionally constrained; non-parametric or compositional methods can better reflect the data structure
-
Viral load was categorized into discrete groups (low/medium/high defined by Ct cut-points) for most analyses↳ Could also: Modeling Ct value as a continuous variable using linear regression or penalized splines would also characterize the relationship between viral load and host gene expression — Continuous modeling retains all quantitative information in the Ct values and can reveal non-linear dose–response patterns that fixed cut-point categorization may obscure
-
Dispersion in violin plots is conveyed solely by the shape of the distribution, without overlaid numeric summary statistics↳ Could also: Overlaying median with interquartile range (IQR) markers, or reporting 95% confidence intervals alongside the plots, would also summarize central tendency and spread — Explicit numeric dispersion measures allow readers to assess variability and estimate uncertainty independently of visual interpretation of the plot shape
-
The longitudinal sub-analysis assessed gene expression changes across two timepoints in n = 3 patient-matched pairs↳ Could also: A paired Wilcoxon signed-rank test with explicit FDR correction across tested genes, or framing the analysis as purely descriptive or hypothesis-generating, would also be applicable at this sample size — With three pairs, statistical power to detect true longitudinal changes while controlling error rates is limited; explicitly scoping the analysis as exploratory aligns inferential claims with the available sample size
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32898168
Paper: Lieberman et al. 2020, In vivo antiviral host transcriptional response to SARS-CoV-2 by viral load, sex, and age, PLoS Biology. PMID 32898168. Code: github.com/greninger-lab/COVID19-RNAseq (commit 9f0be99, MIT, 100% R). Data: GEO GSE152075 — bulk RNA-seq of 484 nasopharyngeal swabs (430 SARS-CoV-2 positive, 54 negative). GEO deposits the raw gene-count matrix + per-sample metadata (covid status, N1 Ct, age, sex, sequencing batch), so the DESeq2 pipeline starts directly from deposited counts (no fastq re-alignment needed).
In scope (pipeline-derived, reproduced here)
The repo's DESeq2 scripts operate directly on the GEO count matrix. We reproduce:
| ID | Result | Reported | Source | Script |
|---|---|---|---|---|
| A1 | DE genes, SARS-CoV-2 pos vs neg (padj<0.1 & |log2FC|>1) | 83 (41 up, 42 down) | S1 Table / Results | Fig1_ABC_Deseq_pos_v_neg.R |
| A1b | ACE2 median norm counts by viral load (neg/low/mid/high) | 0 / 1.93 / 3.45 / 7.82 | Results, Fig2A | Fig2_ABCD_IFN_dose_response.R |
| A2 | DE genes, high vs low viral load (padj<0.1) | 363 | S2 Table | Fig2 (hi_v_lo DESeq) |
| A3 | age×covid interaction; CXCL11, PCGF6 reduction in older | 30-fold / 17-fold | Results, Fig5 | Fig5_DEseq_age.R |
| A4 | sex×covid interaction DE genes | 19 (incl. IGHG1, MS4A1, SLAMF6) | Results, Fig5 | Fig5_DEseq_sex.R |
Pipeline: DESeq2 v1.28.1 (Wald test, BH padj, batch in design formula). Exact design formulas, prefilters (rowSums≥484/207/467/431 = subset sizes) and contrasts taken verbatim from the repo scripts.
Out of scope (not pipeline-derived from GSE152075, not attempted)
- Upstream alignment (kallisto 0.46 pseudoalignment of fastq → counts): raw reads are NOT deposited (GEO: "Raw data not provided … patient privacy"). We start from the deposited counts, exactly as the repo scripts do.
- GSEA, GO plots, CIBERSORTx (Fig1C, FigS1, Fig2C): rely on external desktop software (GSEA app, CIBERSORTx web) + intermediate csvs not in GEO. Not a GSE152075-pipeline output; not attempted.
- HAE cultures (Fig3, GSE154768), longitudinal (Fig4, GSE154769): separate accessions outside the brief's named dataset (GSE152075). Not attempted.
- Wet-lab / clinical (Ct assays, patient ascertainment): non-computational.
Honesty notes
- The repo expects a hand-built metadata CSV (columns alt_name, seq_run, covid_status, Viral_Load, hi_v_lo, sixty_plus, sex). We reconstruct it from the GEO series-matrix characteristics; mapping is documented in build_meta.py.
- sixty_plus threshold: age ≥ 60 (variable name; paper says "≥60").
- A4 "19 genes": repo additionally runs biomaRt to drop patch-contig duplicates; we report the raw padj<0.1 interaction count and the named genes.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean, derivable reproduction: the deposited GSE152075 count matrix plus the authors' DESeq2 scripts regenerate every headline value within tolerance or exactly (ACE2 Mid/High medians exact, CXCL11 29.7≈30-fold, named sex genes all down). The only deviations — off-by-1/6 DEGs, PCGF6 15.9 vs 17-fold, one fewer low-VL sample — are small and explained by DESeq2 1.50.2 vs 1.28.1 version drift and a minor GEO deposit gap (445 vs 467 annotated ages), not by anything on the authors' analytic side. No fabrication concern; the central conclusions hold in magnitude and direction, so this is a solid yellow-overall reproduction driven by an input/version technicality rather than any output-logic problem.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.