Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

In vivo antiviral host transcriptional response to SARS-CoV-2 by viral load, sex, and age.

PLoS Biol · 2020
L1 89/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
89/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 77% of all assessed papers rank 246 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1. GSE152075 deposits the raw count matrix + per-sample metadata, and the authors' R repo (greninger-lab/COVID19-RNAseq @9f0be99) runs DESeq2 directly on it; no fastq re-alignment needed (raw reads withheld for privacy, as stated). Five reported pipeline-derived results reproduced via the authors' exact design formulas/prefilters/contrasts on «our HPC»: pos-vs-neg DEGs 84 vs 83; ACE2 viral-load medians (Mid 3.46/High 7.82 exact); high-vs-low viral-load DEGs 369 vs 363; age-effect CXCL11 29.7-fold (paper 30) and PCGF6 15.9-fold (paper 17); sex x infection 20 genes vs 19 with all three named genes (IGHG1, MS4A1, SLAMF6) down-regulated. All within tolerance or exact. Deviations explained by DESeq2 1.50.2 (paper used 1.28.1; exact pin did not solve) and the GEO age field carrying 445 numeric ages vs the 467 implied by the repo prefilter. No fabrication concern: every headline value is derivable from the shipped counts+code and reproduces independently. NOT attempted (out of scope): upstream kallisto alignment (no raw fastq), GSEA/GO/CIBERSORTx (external desktop/web tools + intermediate csvs), and the separate HAE (GSE154768) and longitudinal (GSE154769) accessions outside the brief's named GSE152075.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 89
    assessed: 2026-06-18 ⛓ e27e5e7b7734
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What host transcriptional responses underlie the wide range of SARS-CoV-2 clinical manifestations, and how do these antiviral responses vary by viral load, infection time course, sex, and age?

Core claims
  • SARS-CoV-2 infection induces a strong interferon-mediated antiviral response in the nasopharynx, up-regulating antiviral factors (OAS1-3, IFIT1-3, MX2, RSAD2, HERC5) and Th1 chemokines CXCL9/10/11, while down-regulating ribosomal protein transcripts. finding
  • Expression of interferon-responsive genes, including ACE2, increases as a function of viral load, while B cell–specific transcripts and neutrophil chemokines are elevated in low viral load samples. finding
  • Older individuals have reduced expression of Th1 chemokines CXCL9/10/11, their receptor CXCR3, CD8A, and granzyme B, suggesting deficiencies in cytotoxic T cell and NK cell trafficking/function. finding
  • Males have reduced B cell– and NK cell–specific transcripts and increased inhibitors of NF-κB signaling relative to females, possibly throttling antiviral responses. finding
  • Human airway epithelial (HAE) cultures recapitulate the cell-intrinsic antiviral response seen in NP swabs at 7 days post infection, with no interferon induction at 3 days, defining a consensus set of 19 antiviral genes. finding
  • Longitudinal patient-matched samples show reduced viral load over time accompanied by reduced interferon-induced transcription, recovery of ribosomal protein expression, and initiation of wound healing and humoral immune responses. finding
  • Shotgun RNA sequencing of diagnostic NP swabs enables simultaneous recovery of viral genomes and in situ host response profiling. method
  • SARS-CoV-2 may exploit host antiviral responses by inducing ACE2 (an interferon-regulated gene) upon interferon exposure. mechanism
Experimental setups
Assay System Perturbation Readout Platform
shotgun/bulk RNA-seq (metagenomic) nasopharyngeal swabs from 430 SARS-CoV-2 PCR-positive and 54 negative human individuals natural SARS-CoV-2 infection (none experimental) host gene expression / differentially expressed genes
bulk RNA-seq male human airway epithelial (HAE; tracheal/bronchial) cultures SARS-CoV-2 infection (3 and 7 days post infection) host gene expression changes and percent viral reads
bulk RNA-seq (longitudinal) patient-matched nasopharyngeal swabs (n=3), mean 6.3 days apart natural infection over time course changes in viral load (Ct) and host gene expression
RT-PCR (diagnostic) nasopharyngeal swabs none SARS-CoV-2 N1 cycle threshold (viral load)
in silico cell-type deconvolution (CIBERSORTx) NP swab RNA-seq (high vs low viral load) none immune cell type proportions CIBERSORTx
Gene Set Enrichment Analysis (GSEA) / GO / DisGeNET NP swab and HAE RNA-seq data none enriched pathways and disease ontology terms MSigDB Hallmark Gene Sets, Gene Ontology, DisGeNET
Key results
  • 83 DE genes between SARS-CoV-2 positive and negative samples (41 up, 42 down) 83 genes (padj<0.1, |log2FC|>1)
  • ACE2 expression increases with viral load (median counts negative/low/medium/high = 0, 1.93, 3.45, 7.82) 0→7.82 median counts; p=7.46×10^-13
  • 363 DE genes between high vs low viral load; high enriched for CXCL9/10, IDO1, CD80; low enriched for neutrophil (CXCL8, S100A9) and B cell transcripts (IGHG1, IGHM, CD22) 363 genes (padj<0.1)
  • Low viral load samples have higher proportions of naïve B/T cells, neutrophils, M2 macrophages; high viral load has more M1 macrophages, activated NK cells, activated DCs 3.5-, 2.2-, 1.6-, 1.8-fold (low); 2.5-, 1.6-, 1.6-fold (high)
  • Consensus set of 19 up-regulated antiviral genes shared across positive-vs-negative, high-vs-low, and infected HAE analyses 19 genes; p=4.81×10^-69
  • HAE show no interferon response at 3 days post infection (virus 0.3% of reads) despite 10-fold higher dose; antiviral response by 7 days (virus 5.3% of reads) 0.3%→5.3% viral reads
  • Longitudinal samples show 39-fold reduction in viral load and increased humoral/wound-healing genes (C1QA/B/C, HLA-DQB1, APOE, CD36, RHOC) with recovery of ribosomal proteins mean Ct increase 5.29 = 39-fold reduction
  • TMPRSS2 reduced upon infection but not modulated by viral dose; ribosomal proteins (RPL4, RPS6) not modulated by viral load
Key statistics
  • pvalue 7.46 × 10^-13 (ACE2 expression across viral load groups, Kruskal-Wallis one-way ANOVA)
  • pvalue 4.81 × 10^-69 (SuperExact Test for 19-gene consensus antiviral set overlap)
  • count 83 (DE genes positive vs negative (41 up, 42 down))
  • count 363 (total DE genes high vs low viral load (padj<0.1))
  • fold_change 39-fold reduction (viral load reduction over time in longitudinal samples (mean Ct increase 5.29))
  • fold_change 3.5-, 2.2-, 1.6-, 1.8-fold (increased naïve B/T cells, neutrophils, M2 macrophages in low viral load (CIBERSORTx))
  • count 430 positive, 54 negative (NP swab samples analyzed; median 1.61×10^6 human reads pseudoaligned)
  • other 0.3% and 5.3% (SARS-CoV-2 reads in HAE at day 3 and day 7 post infection)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This cross-sectional observational RNA-seq study compared nasopharyngeal host transcriptomes from 430 RT-PCR-confirmed SARS-CoV-2-positive individuals and 54 negative controls, with stratified analyses by viral load (Ct-defined low/medium/high), sex, and age. Transcriptome-wide differential expression was assessed using FDR-adjusted p-values (padj < 0.1) and log2 fold-change thresholds, and pathway enrichment was evaluated with GSEA (FDR < 0.05). Individual gene-level comparisons across viral load groups used the Kruskal-Wallis test and Mann-Whitney U test; immune cell-type proportion estimates (CIBERSORTx) were compared by t-test. A small paired longitudinal sub-analysis (n = 3) and an in vitro HAE cell experiment supplemented the main cross-sectional analysis.

Replicationbiological Sample size430 SARS-CoV-2 RT-PCR-confirmed positive and 54 negative NP swab samples selected from a larger sequencing repository based on a >500,000-read threshold; n = 3 matched longitudinal pairs (mean elapsed time 6.3 days); HAE cell replicate number not stated in provided text GroupsSARS-CoV-2 positive vs. negative; high vs. medium vs. low viral load (Ct-defined); male vs. female; age as a variable (median and range reported); longitudinal early vs. later timepoints (n = 3 pairs); in vitro HAE infected vs. uninfected at day 3 and day 7 Pairingmixed Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionFDR (padj used for transcriptome-wide DE gene calling; FDR < 0.05 for GSEA; correction method not explicitly named); no correction explicitly stated for individual Mann-Whitney U tests or t-tests on selected genes or cell-type proportions
Statistical tests used
Test Applied to n Assumptions
FDR-adjusted differential expression analysis (padj < 0.1, |log2FC| > 1) SARS-CoV-2 positive vs. negative (Figs 1A–1B) and high vs. low viral load (Fig 2B) 430 positive vs. 54 negative; 108 high vs. 99 low viral load not stated
Kruskal-Wallis one-way ANOVA ACE2 and select interferon-stimulated gene expression across negative, low, medium, and high viral load groups (Fig 2A) 54 negative, 99 low, 206 medium, 108 high not stated
Mann-Whitney U test Individual gene expression comparisons between low and high viral load groups (Figs 2A, 2D) n = 99 low, n = 108 high not stated
t-test (type not further specified) CIBERSORTx-estimated immune cell-type proportions between viral load groups (Fig 2C) n = 99 low, n = 108 high not stated
Gene Set Enrichment Analysis (GSEA) with FDR correction Hallmark Gene Sets and GO Biological Processes enrichment in SARS-CoV-2 positive vs. negative (Fig 1C, S1 Figs) 430 positive, 54 negative not stated
SuperExact Test (multi-set intersection probability) Overlap of DE gene lists across positive vs. negative, high vs. low viral load, and HAE day-7 analyses (Fig 3A) null not stated
Approaches that could also have been used
  • Multiple individual Mann-Whitney U tests were applied to selected genes for the high-vs-low viral load comparison without an explicitly stated multiplicity correction across those individual tests
    Could also: A transcriptome-wide model (e.g., negative-binomial GLM as in DESeq2 or edgeR) with viral load as a covariate and Benjamini-Hochberg FDR correction across all genes would also address these comparisons — the paper already applies an analogous approach for the positive-vs-negative contrast — A unified model applied consistently across all genes simultaneously controls the false discovery rate across the full gene set and can yield more precise effect-size estimates than a subset of selected-gene tests
  • Viral load, age, and sex effects were characterized in separate stratified analyses rather than in a single multivariable model
    Could also: A multivariable linear model or generalized linear model incorporating viral load, age, and sex simultaneously as covariates (e.g., within a DESeq2 design formula) would also estimate each factor's contribution to gene expression — Because viral load, age, and sex may be correlated in the sample, a joint model can partially separate their independent associations with the transcriptional response
  • CIBERSORTx-estimated immune cell-type proportions were compared between viral load groups using a t-test
    Could also: A Wilcoxon rank-sum (Mann-Whitney U) test, permutation test, or Dirichlet regression would also be applicable for deconvolved cell-type proportions — Deconvolved proportions are bounded between 0 and 1 and are often non-normally distributed and compositionally constrained; non-parametric or compositional methods can better reflect the data structure
  • Viral load was categorized into discrete groups (low/medium/high defined by Ct cut-points) for most analyses
    Could also: Modeling Ct value as a continuous variable using linear regression or penalized splines would also characterize the relationship between viral load and host gene expression — Continuous modeling retains all quantitative information in the Ct values and can reveal non-linear dose–response patterns that fixed cut-point categorization may obscure
  • Dispersion in violin plots is conveyed solely by the shape of the distribution, without overlaid numeric summary statistics
    Could also: Overlaying median with interquartile range (IQR) markers, or reporting 95% confidence intervals alongside the plots, would also summarize central tendency and spread — Explicit numeric dispersion measures allow readers to assess variability and estimate uncertainty independently of visual interpretation of the plot shape
  • The longitudinal sub-analysis assessed gene expression changes across two timepoints in n = 3 patient-matched pairs
    Could also: A paired Wilcoxon signed-rank test with explicit FDR correction across tested genes, or framing the analysis as purely descriptive or hypothesis-generating, would also be applicable at this sample size — With three pairs, statistical power to detect true longitudinal changes while controlling error rates is limited; explicitly scoping the analysis as exploratory aligns inferential claims with the available sample size
Software: GSEA (Gene Set Enrichment Analysis) · CIBERSORTx · SuperExact Test · DisGeNET

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-32898168

Paper: Lieberman et al. 2020, In vivo antiviral host transcriptional response to SARS-CoV-2 by viral load, sex, and age, PLoS Biology. PMID 32898168. Code: github.com/greninger-lab/COVID19-RNAseq (commit 9f0be99, MIT, 100% R). Data: GEO GSE152075 — bulk RNA-seq of 484 nasopharyngeal swabs (430 SARS-CoV-2 positive, 54 negative). GEO deposits the raw gene-count matrix + per-sample metadata (covid status, N1 Ct, age, sex, sequencing batch), so the DESeq2 pipeline starts directly from deposited counts (no fastq re-alignment needed).

In scope (pipeline-derived, reproduced here)

The repo's DESeq2 scripts operate directly on the GEO count matrix. We reproduce:

ID Result Reported Source Script
A1 DE genes, SARS-CoV-2 pos vs neg (padj<0.1 & |log2FC|>1) 83 (41 up, 42 down) S1 Table / Results Fig1_ABC_Deseq_pos_v_neg.R
A1b ACE2 median norm counts by viral load (neg/low/mid/high) 0 / 1.93 / 3.45 / 7.82 Results, Fig2A Fig2_ABCD_IFN_dose_response.R
A2 DE genes, high vs low viral load (padj<0.1) 363 S2 Table Fig2 (hi_v_lo DESeq)
A3 age×covid interaction; CXCL11, PCGF6 reduction in older 30-fold / 17-fold Results, Fig5 Fig5_DEseq_age.R
A4 sex×covid interaction DE genes 19 (incl. IGHG1, MS4A1, SLAMF6) Results, Fig5 Fig5_DEseq_sex.R

Pipeline: DESeq2 v1.28.1 (Wald test, BH padj, batch in design formula). Exact design formulas, prefilters (rowSums≥484/207/467/431 = subset sizes) and contrasts taken verbatim from the repo scripts.

Out of scope (not pipeline-derived from GSE152075, not attempted)

  • Upstream alignment (kallisto 0.46 pseudoalignment of fastq → counts): raw reads are NOT deposited (GEO: "Raw data not provided … patient privacy"). We start from the deposited counts, exactly as the repo scripts do.
  • GSEA, GO plots, CIBERSORTx (Fig1C, FigS1, Fig2C): rely on external desktop software (GSEA app, CIBERSORTx web) + intermediate csvs not in GEO. Not a GSE152075-pipeline output; not attempted.
  • HAE cultures (Fig3, GSE154768), longitudinal (Fig4, GSE154769): separate accessions outside the brief's named dataset (GSE152075). Not attempted.
  • Wet-lab / clinical (Ct assays, patient ascertainment): non-computational.

Honesty notes

  • The repo expects a hand-built metadata CSV (columns alt_name, seq_run, covid_status, Viral_Load, hi_v_lo, sixty_plus, sex). We reconstruct it from the GEO series-matrix characteristics; mapping is documented in build_meta.py.
  • sixty_plus threshold: age ≥ 60 (variable name; paper says "≥60").
  • A4 "19 genes": repo additionally runs biomaRt to drop patch-contig duplicates; we report the raw padj<0.1 interaction count and the named genes.
Figures / tables: S1 TableFig 2AS2 TableFig 5
A1_deg
Reported
83 (41 up, 42 down)
Reproduced
84 (42 up, 42 down)
within tolerance
A1b_ace2
Reported
ACE2 medians 0/1.93/3.45/7.82
Reproduced
0/1.82/3.46/7.82
within tolerance
A2_deg
Reported
363 DEGs high vs low viral load
Reproduced
369
within tolerance
A3_cxcl11
Reported
~30-fold CXCL11 reduction in older
Reproduced
29.7-fold (log2FC -4.893)
exact
A3_pcgf6
Reported
~17-fold PCGF6 reduction in older
Reproduced
15.9-fold (log2FC -3.992)
within tolerance
A4_sex
Reported
19 sex x covid genes (IGHG1, MS4A1, SLAMF6 down)
Reproduced
20; IGHG1/MS4A1/SLAMF6 all down
within tolerance
N_samples
Reported
484 (430 pos / 54 neg)
Reproduced
484 (430 pos / 54 neg)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 89/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

This is a clean, derivable reproduction: the deposited GSE152075 count matrix plus the authors' DESeq2 scripts regenerate every headline value within tolerance or exactly (ACE2 Mid/High medians exact, CXCL11 29.7≈30-fold, named sex genes all down). The only deviations — off-by-1/6 DEGs, PCGF6 15.9 vs 17-fold, one fewer low-VL sample — are small and explained by DESeq2 1.50.2 vs 1.28.1 version drift and a minor GEO deposit gap (445 vs 467 annotated ages), not by anything on the authors' analytic side. No fabrication concern; the central conclusions hold in magnitude and direction, so this is a solid yellow-overall reproduction driven by an input/version technicality rather than any output-logic problem.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

158.9 k
tokens (I/O) · 13.2 M incl. cache
36 min
runtime · 0.37 CPU-h
3.8 GB
peak RAM
1
HPC jobs
hummel
machine