Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

COVID-19 lung disease shares driver AT2 cytopathic features with Idiopathic pulmonary fibrosis.

EBioMedicine · 2022
L1 85/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

salvaged by watchdog from agreement.json (agent omitted ROOM_RESULT.json)

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-07-29

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper tests whether post-COVID-19 fibrotic lung disease (PCLD) shares fundamental molecular, cytopathic, and immunologic driver features with idiopathic pulmonary fibrosis (IPF), and seeks to identify the earliest cellular/molecular triggers of alveolar type II (AT2) dysfunction driving fibrosis.

Core claims
  • COVID-19 lung disease resembles IPF at a fundamental level, recapitulating ViP/IPF gene expression patterns, an IL15-centric cytokine storm, and AT2 cytopathic changes (injury, DNA damage, transient progenitor-state arrest, senescence/SASP). finding
  • ER stress is a shared early trigger of both COVID-19 and IPF that culminates in AT2 progenitor-state arrest and SASP. mechanism
  • AT2 immunocytopathic features induced by SARS-CoV-2 can be recapitulated in pre-clinical models (adult lung organoids and hamster) and reversed with effective anti-CoV-2 therapeutics in hamsters. finding
  • tg-mice with AT2-specific induced ER stress faithfully recapitulate the host immune response and alveolar cytopathic changes induced by SARS-CoV-2. finding
  • ViP signatures in monocytes may be key determinants of prognosis in these fibrotic lung diseases. finding
  • An AI/machine-learning-guided approach using ViP, sViP, and COVID-lung gene signatures plus PPI network analysis can identify shared disease drivers across >1000 lung transcriptomic datasets. method
  • PPI-network analysis pinpointed ER stress as a point of convergence among diverse AT2 quality-control failure phenotypes. mechanism
  • The disease models, gene signatures, and biomarkers identified are translational resources applicable to IPF and other fibrotic interstitial lung diseases. resource
Experimental setups
Assay System Perturbation Readout Platform
Composite gene signature / transcriptomic meta-analysis (BoNE, StepMiner) Human lung transcriptomic datasets (>1000) across lung conditions including COVID-19 and IPF none Composite signature scores (ViP, sViP, COVID-lung, IPF signatures), ROC-AUC classification Affymetrix microarray (RMA) and RNASeq (TPM) from NCBI GEO
Single-cell RNA-seq analysis (pseudo-bulk) Human lung/BAL cells (GSE145926, GSE159354, GSE132914, GSE146981, GSE149878) none Cell-type-resolved composite signature scores Seurat v3 / scanpy v1.5.1; SCINA cell typing
Kaplan-Meier survival / outcome analysis COVID-19 patients (GSE157103) and IPF patients (GSE28221) none Survival / hospital-free days stratified by signature score lifelines python v0.14.6; log-rank test
Protein-protein interaction network (PPIN) construction Human protein interaction data (STRING database) seeded with signature genes none Shortest-path connecting nodes identifying shared triggers (ER stress) STRING; NetworkX; Cytoscape
Immunofluorescence / immunohistochemistry Human COVID-19 autopsy/biopsy lung tissue and SARS-CoV-2-challenged hamster lungs SARS-CoV-2 infection Protein expression of ER stress (GRP78/BIP), senescence (p14ARF, p53, p21), AT2/progenitor markers (CK8, Claudin-4), SARS-CoV-2 nucleoprotein Leica DMI4000B microscope; ImmPRESS HRP detection kit
qPCR Human adult lung organoid (ALO) and hamster COVID-19 pre-clinical models SARS-CoV-2 infection +/- anti-CoV-2 therapeutics Validation of key transcriptomic findings (gene expression)
Cytokine quantification (MSD multiplex) Pre-clinical COVID models / lung samples SARS-CoV-2 infection Cytokine levels (e.g., IL15-centric storm) MESO QuickPlex SQ 120; MSD DISCOVERY WORKBENCH 4.0
Multivariate (OLS) regression analysis Single-cell (GSE132914) and bulk (GSE150910) IPF datasets none Healthy vs IPF modeled as linear combination of sViP, ViP, CoV-lung, and six IPF signature scores python statsmodels
Key results
  • More than a third of COVID-19 survivors develop fibrotic lung abnormalities (post-COVID-19 ILD). >1/3 of survivors
  • Fibrosis development scales with COVID-19 disease duration. ~4% (<1 wk), ~24% (1-3 wk), ~61% (>3 wk)
  • COVID-19 and IPF share gene expression patterns (ViP and IPF signatures) in lungs and blood.
  • AT2 cytopathic changes (injury, DNA damage, transient progenitor arrest, senescence/SASP) induced in ALO and hamster COVID models.
  • AT2 cytopathic features reversed with effective anti-CoV-2 therapeutics in hamsters.
  • ER stress validated by IHC in lungs of deceased COVID-19 subjects and SARS-CoV-2-challenged hamster lungs.
  • AT2-specific ER stress in tg-mice recapitulates SARS-CoV-2-induced host immune response and alveolar cytopathic changes.
Key statistics
  • count 166 genes (ViP signature gene set)
  • count 20 genes (severe-ViP (sViP) signature classifying disease severity)
  • count 52 gene (IPF signature used in Kaplan-Meier analysis)
  • count ~61% (fibrosis in patients with disease duration >3 weeks)
  • count n=438 (ViP training dataset GSE47963)
  • count n=118 (ViP training dataset GSE113211)
  • count n=159 (sViP severity cohort GSE101702)
  • count 114 patients (male n=87, female n=27) (IPF survival dataset GSE28221)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is an AI/bioinformatics-driven study that analyzes >1000 publicly available human (and model-system) lung transcriptomic datasets using composite gene signatures (ViP, sViP, IPF, COVID-lung) scored via the StepMiner/BoNE framework, with experimental validation in hamster and organoid models by IHC/qPCR. Sample categories were classified and the separation quantified by ROC-AUC, and signature scores between groups were compared with Welch's unpaired two-sample t-test; survival was assessed by Kaplan-Meier with log-rank tests, and OLS multivariate regression related signatures to disease state. Multiple-comparison p-values were adjusted by Benjamini-Hochberg FDR, and sample numbers were reported alongside each dataset/plot.

Replicationmixed Sample sizesample number for each analysis reported alongside the associated plot/GSE ID; survival cohorts give explicit n (e.g., IPF n=114) Groupsdisease vs control categories (e.g., COVID-19/IPF vs normal lung), high vs low signature groups Pairingunpaired Randomization/blindingnot stated Dispersionunclear Effect sizesyes Multiplicity correctionBenjamini-Hochberg FDR (statsmodels.stats.multitest.multipletests, fdr_bh)
Statistical tests used
Test Applied to n Assumptions
Welch's two-sample t-test (unpaired, unequal variance, unequal n) comparison of composite/gene-signature scores between sample categories (violin/swarm/bubble plots; Figs 2,4,6,8,S1) stated per plot beside each GSE ID/sample name (specific values not given in methods) stated
ROC-AUC performance of gene-signature-based (multi-class) sample classification na
Log-rank test (Kaplan-Meier) survival/hospital-free-days analysis for sViP, 52-gene IPF, and COVID-lung signatures (GSE157103 COVID; GSE28221 IPF) IPF GSE28221: 114 patients (male n=87, female n=27); COVID GSE157103 limited to <70 yr na
Ordinary least-squares (OLS) multivariate regression modeling healthy vs IPF as a linear combination of sViP/ViP/CoV-lung and six IPF signatures (GSE132914 single-cell; GSE150910 bulk; Fig S1) stated (null hypothesis: coefficient = 0)
StepMiner adaptive-regression step-fit (F-statistic) binarization of gene expression into high/low and threshold setting for signatures stated (regression test statistic defined)
Approaches that could also have been used
  • Group signature scores were compared with Welch's unpaired two-sample t-test.
    Could also: A nonparametric Mann-Whitney U (Wilcoxon rank-sum) test, or a permutation test on the score difference. — A rank-based or permutation approach makes no normality assumption and can be informative when signature-score distributions are skewed or when some groups have small n; it would complement the parametric Welch's result.
  • Multiple comparisons were adjusted using the Benjamini-Hochberg FDR procedure.
    Could also: Bonferroni or Holm-Bonferroni family-wise error control, or Benjamini-Yekutieli FDR for dependent tests. — Family-wise procedures control the probability of any false positive (more stringent for confirmatory claims), while BY-FDR is valid under arbitrary dependence; reporting which family each correction spans clarifies the inferential scope.
  • Survival differences were assessed with Kaplan-Meier curves and the log-rank test, with groups split at the StepMiner threshold.
    Could also: A Cox proportional-hazards model treating the signature score as a continuous covariate (optionally with age/sex). — A Cox model yields a hazard ratio with a confidence interval and avoids dichotomization, retaining information lost when a continuous score is split into high/low groups and allowing adjustment for covariates such as the gender variable already examined.
  • Separation between sample categories was summarized with ROC-AUC.
    Could also: Reporting an accompanying 95% confidence interval for the AUC (e.g., via DeLong's method or bootstrap), and/or precision-recall AUC. — An interval conveys the precision of the AUC estimate, and PR-AUC can be more informative under class imbalance; together they give a fuller picture of classification performance.
  • Healthy-vs-IPF state was modeled with OLS linear regression on composite signature scores.
    Could also: Logistic regression (or penalized/regularized logistic regression) for the binary outcome. — Logistic regression matches the dichotomous outcome and yields interpretable odds ratios with proper variance estimates, while regularization can stabilize coefficients when multiple correlated signatures are entered together.
  • Dispersion/uncertainty around plotted signature scores and estimates is conveyed via distribution plots and AUC values.
    Could also: Explicitly annotating SD, a 95% CI, or IQR for group summaries and reporting effect-size measures (e.g., Cohen's d, rank-biserial correlation). — Stating a specific spread measure and an effect size alongside p-values helps readers gauge the magnitude and precision of differences, which is especially useful when sample sizes vary across datasets.
Software: R 3.2.3 (2015-12-10) · Python scipy.stats (ttest_ind, Welch's) 0.19.0 · Python statsmodels (multipletests; OLS) · lifelines (Kaplan-Meier/log-rank) 0.14.6 · seaborn (plots) 0.10.1 · scanpy (single-cell) 1.5.1 · Seurat (single-cell) v3 · SCINA / StepMiner / BoNE · GraphPad Prism · ImageJ; NetworkX; Cytoscape; MSD DISCOVERY WORKBENCH 4.0 (MSD)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
42
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE122960 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
also used by 1 paper:
103390 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE132914 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE146981 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE149878 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE157057 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE159354 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
NCT04405570 NCT in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
NCT04653831 NCT in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet
NCT04856111 NCT in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35870428 (COVID-19 lung shares driver AT2 cytopathic features with IPF)

Paper: Sinha S, ..., Sahoo D, Ghosh P. EBioMedicine 2022. PMID 35870428 / PMC9297827. DOI 10.1016/j.ebiom.2022.104185. Tool/code: github.com/sahoo00/BoNE (Boolean Network Explorer, Sahoo lab), GPL-3.0, commit c950951ddaf1bd99f8bc6e1b2b13d170d31ff4d8 (master, default branch). Brief's named data: GEO GSE149878.

What this paper is

A BoNE / signature-scoring paper. Its central computational result: published gene signatures (the ViP = 166-gene viral-pandemic signature and sViP = 20-gene severe-ViP signature, from Sahoo et al. Nat Commun 2021; plus IPF / DATP / AT2-senescence / TERC signatures) produce a per-sample composite score whose ROC-AUC separates diseased lung (COVID-19, IPF/ILD) from healthy lung across many GEO cohorts. AUCs are displayed as bubble-plot radii (Fig 2A/2B/2I, 4A/4D/4E, S1, S2) — the repo's lung-fibrosis/Lung_fibrosis_bubble_plots.ipynb is the exact code that generates them.

IN SCOPE (pipeline-derived, attempted)

The composite-score → ROC-AUC pipeline for the ViP and sViP signatures on the named pooled cohort the authors call COV339 = GSE149878 + GSE122960 ("Xu 2020 CoV2 bulk, n=21" in the notebook; the binary contrast actually scored is 8 healthy donor lungs (GSE122960 Donor_01..08) vs 4 COVID-19 lungs (GSE149878 C166/C168/C170/C172) = 12 samples). This is the Fig 2A/2B target and the only figure-cell that uses the brief's named accession GSE149878.

  • Pipeline: 10x scRNA-seq filtered_*_bc_matrix.h5 per sample → pseudobulk (sum raw UMI counts over all cells per gene) → CPM → log2 → per-gene standardize over the 12 samples → cluster score = Σ z over signature genes → composite = Σ weightᵢ·clusterᵢ (single positive cluster, weight +1 for ViP and for sViP, matching the notebook bone.getViP()/getSViP() single-cluster output [164]/[20]) → ROC-AUC (sklearn roc_curve/auc, pos_label = diseased), exactly as bone.py getRanks2 + mergeRanks + getROCAUC.
  • Gene lists: taken verbatim from the repo: SMaRT/database/vip-signature.txt (166), SMaRT/database/svip-signature.txt (20).

Why this is a faithful 1:1 (per BRIEF rule P16)

Applying the authors' own published tool (BoNE) + their own published gene lists to the paper's own named GEO data, with the scoring algorithm ported verbatim from the public bone.py. The one unspecified step is the scRNA→pseudobulk normalization (the authors' Hegemon-internal preprocessing of these h5 files is not shipped); we use a standard sum→CPM→log2 pseudobulk and report the matched-gene count and AUC honestly.

OUT OF SCOPE (not attempted, with reason)

  • The other ~30 dataset bubbles across Fig 2/4/S1/S2 (GSE171524, GSE158127, GSE132914, GSE146981, GSE159354, GSE145926, mouse GSE161615/GSE167400, etc.): each needs the same pseudobulk reconstruction; the named accession is GSE149878, so we reproduce that cohort and treat the rest as the optional 80/20 tail.
  • IPF Kaplan–Meier / survival panels (Lung_fibrosis_KMplots.ipynb) and the multivariate signature panel (MULTIVARIATE_SIG.ipynb): need Hegemon survival objects / additional bulk cohorts; not the named accession.
  • All wet-lab results (autopsy IHC, hamster/mouse challenge, organoid GRP78-KO, serum cytokines): non-pipeline, out of scope.
  • Exact numeric AUC ground truth: the paper prints AUCs only as bubble radii, no numbers in text/tables → the "reported" value is read qualitatively from the figure (high, ~0.9–1.0, diseased > healthy) and flagged as not text-pinnable.

Reproduction unit

1 in-scope claim attempted: ViP & sViP composite-score ROC-AUC, COVID-19 vs healthy lung, COV339 (GSE149878 + GSE122960), Fig 2A.

Figures / tables: Fig 2A
ViP-COV339
Reported
high-AUC bubble (~0.9-1.0), COVID/diseased > healthy; no numeric value printed in paper (encoded as bubble radius)
Reproduced
AUC = 1.00 (perfect separation), COVID-high direction; 162/166 signature genes matched; COVID score mean +99.99 vs healthy -49.99
within tolerance
sViP-COV339
Reported
high-AUC bubble (~0.9-1.0), COVID/diseased > healthy; no numeric value printed
Reproduced
AUC = 1.00 (perfect separation), COVID-high direction; 19/20 signature genes matched; COVID score mean +10.40 vs healthy -5.20
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

The central Fig 2A claim reproduces cleanly: using the authors' own BoNE tool, published 166/20-gene lists, and their named public GEO cohort (GSE149878+GSE122960), both ViP and sViP signatures give AUC=1.00 in the COVID-high direction (162/166 and 19/20 genes matched), confirming the qualitative high-AUC bubble. The deviations are not on the authors' side: the paper simply prints no numeric AUC (bubble-radius encoding), so the endpoint is only qualitatively comparable, and the scRNA→pseudobulk normalization was a self-chosen step because the authors' preprocessing was not deposited. Severity is negligible — direction and high-AUC magnitude hold and nothing is fabrication-suspect (perfect separation is expected for 4 vs 8 samples on a strong viral signature). Graded yellow overall because the comparison is indirect and only 1 of ~30 bubbles was reproduced.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

145.6 k
tokens (I/O) · 10.7 M incl. cache
24 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.