Systematic identification of ACE2 expression modulators reveals cardiomyopathy as a risk factor for mortality in COVID-19 patients.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the paper's BIOLOGICAL PIVOT 1:1 on public data; its headline METHOD and CLINICAL CONCLUSION are not reproducible from shipped/public artifacts. Reproduced («our HPC»/«infra», GEO data fetched in-job): (C1) the flagship claim that ACE2 is upregulated in hypertrophic cardiomyopathy (GSE89714, Fig 2A/B) -> ACE2 RPKM 31.38(HCM,n=5) vs 9.20(normal,n=4), 3.41x up, Welch p=5.7e-4, MWU p=0.016 -- direction matches decisively. (C2) the cardiomyopathy meta-analysis claim 'ACE2 significantly elevated, p<0.001' on the obtainable public subset of the 7 datasets: GSE89714 (HCM) 3.41x up, GSE71613 (RCM/DCM) 2.27x up, GSE99321 (DCM) 2.17x up; combined mixed-effect model beta=1.42 p=2.0e-8, all 3/3 datasets up. NOT attempted and why: (a) the GENEVA screen over ARCHS4 ('27 significant datasets FDR<0.05'; FABP2-ACE2 r=0.72) -- the geneva-webtool repo is a Django app that does NOT ship its backend (exp.csv 286,650-sample ARCHS4 percentile matrix, df_embed.csv) and is tied to a 2020 ARCHS4 snapshot now superseded, so not version-reproducible (the hard ~80%); (b) the COVID-19 clinical mortality results (Cox PH, propensity matching, p=0.004, '438% increase') use the UCSF COVID-19 Data Mart EHR = controlled-access (data_restricted); (c) GSE63161 (hiPSC-CM) downloaded but has no ACE2 row (filtered in-vitro), GSE29819 (Affy CEL/RMA), GSE120838, GSE121893 (single-cell) skipped per 80/20; LVNC/ARVC subtypes not separable in the public subset. Fabrication check: no discrepancy -- every publicly-checkable directional claim holds and the pooled significance clears p<0.001; GENEVA/clinical numbers are un-auditable from public artifacts, not contradicted. Verdict provisional; human audit sheet in AUDIT.md.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 68assessed: 2026-06-15 ⛓ 9eaf4798063b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusACE2, the SARS-CoV-2 cell-entry receptor, may have its expression modulated by specific diseases, drugs, and genetic perturbations; systematically identifying these modulators across large-scale RNA-seq data could reveal risk factors for severe COVID-19, with the specific hypothesis that pre-existing cardiomyopathy (with elevated cardiac ACE2) leads to increased COVID-19 mortality.
- ★ GENEVA is a semi-automated, study-design-agnostic framework that mines large-scale public RNA-seq data to identify conditions modulating a gene of interest's expression method
- ★ Multiple drugs, genetic perturbations, and diseases are associated with ACE2 expression, including cardiomyopathy, HNF1A overexpression, RAD140, and itraconazole finding
- ★ ACE2 expression is significantly elevated in heart tissue across all major cardiomyopathy categories (DCM, HCM, RCM, LVNC, and trend in ARVC) finding
- ★ COVID-19 patients with pre-existing cardiomyopathy have an increased mortality risk relative to patients with other cardiovascular conditions and patients without cardiovascular disease finding
- ★ HNF1A regulates ACE2 expression, with overexpression increasing and knockdown reducing ACE2 in LNCaP prostate cancer cells mechanism
- ACE2 lacks a single conserved transcriptional network and is co-regulated with different gene sets depending on tissue/condition finding
- GENEVA is provided as a freely accessible web application applicable to any gene of interest at genevatool.org resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (meta-analysis/co-expression) | 286,650 human RNA-seq samples from 9124 GEO series (ARCHS4) | none/various (existing conditions) | ACE2 expression and Pearson correlation with other genes | ARCHS4 uniformly preprocessed RNA-seq |
| bulk RNA-seq (validation of tissue effect) | GTEx human tissues | none | ACE2 co-expression correlations across tissues | — |
| RNA-seq (genetic perturbation) | LNCaP prostate cancer cell line | HNF1A overexpression and knockdown; HNF4G overexpression | ACE2 expression change | — |
| RNA-seq (disease comparison, GSE89714) | human heart tissue (normal vs hypertrophic cardiomyopathy) | disease (HCM) | ACE2 expression | — |
| RNA-seq (drug treatment, GSE104177) | human breast cancer xenografts | RAD140 (selective androgen receptor modulator) | ACE2 expression | — |
| RNA-seq (drug treatment, GSE114013) | colorectal cancer cell lines HT55 and SW948 | itraconazole (antifungal) | ACE2 expression | — |
| meta-analysis (mixed-effect model) of RNA-seq and microarray | 7 cardiomyopathy heart-tissue datasets (5 RNA-seq + microarray) | cardiomyopathy (DCM, HCM, RCM, LVNC, ARVC) | ACE2 expression with publication-bias (Egger) assessment | — |
| clinical survival analysis (Cox proportional-hazards, Kaplan-Meier, propensity matching) | 3936 COVID-19 patients from UCSF EHR | pre-existing cardiomyopathy vs other/no cardiovascular disease | mortality / risk of death | UCSF electronic health records |
- ▲ ACE2 expression significantly elevated in heart tissue of cardiomyopathy patients across 7 datasets p < 0.001
- ▲ Cardiomyopathy significantly associated with risk of death vs patients without cardiovascular disease (Cox model) p = 0.004
- ▲ Cardiomyopathy significantly associated with risk of death vs patients with other cardiovascular diseases p = 0.038, 438% increase in observed death rate
- ▲ FABP2 is the gene most highly correlated with ACE2 across all samples r = 0.72
- – HNF1A overexpression increases and knockdown reduces ACE2 expression in LNCaP cells
- ▼ HNF4G overexpression does not significantly increase ACE2 (trend toward reduction) despite positive correlation
- ▲ RAD140 induces ACE2 expression in breast cancer xenografts; itraconazole upregulates ACE2 in HT55 and SW948
- – GENEVA identified 27 significant datasets (FDR < 0.05) as ACE2 modulators FDR < 0.05
- correlation 0.72 (highest correlation between FABP2 and ACE2 across all samples)
- correlation 0.42 (p value < 0.0001) (agreement of ACE2 co-expression between GTEx and GEO data)
- pvalue < 0.001 (elevated ACE2 in cardiomyopathy heart tissue, 7-dataset meta-analysis)
- pvalue 0.004 (Cox model: cardiomyopathy vs no cardiovascular disease, mortality)
- pvalue 0.038 (Cox model: cardiomyopathy vs other cardiovascular diseases, mortality)
- fold_change 438% increase in observed death rate (cardiomyopathy vs other cardiovascular disease patients)
- count 27 significant datasets (FDR < 0.05) (GENEVA-identified ACE2-modulating datasets)
- count 43 / 624 / 3269 (COVID-19 patients: cardiomyopathy / other cardiovascular / no cardiovascular disease)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study uses a multi-component statistical design: (1) large-scale Pearson correlation and mixed-effects modeling across 286,650 public RNA-seq samples to map ACE2 co-expression networks; (2) the GENEVA framework, which applies permutation-based variance scoring with Benjamini-Hochberg FDR correction across 9,124 GEO datasets to identify significant ACE2 modulators; (3) a mixed-effects meta-analysis of seven cardiomyopathy RNA-seq datasets supplemented by t-tests within individual cardiomyopathy subtypes; and (4) multivariable Cox proportional-hazards regression and log-rank tests on propensity score-matched EHR data from 3,936 COVID-19 patients to evaluate mortality risk associated with pre-existing cardiomyopathy.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson correlation | ACE2 co-expression with all human genes across 286,650 samples (global) and within each GEO dataset; also used to validate agreement between GTEx and GEO co-expression profiles | 286,650 samples from 9,124 GEO series (global); r=0.42 between GTEx and GEO correlation vectors | not stated |
| Mixed-effects model (random-effects meta-analytic regression) | Estimating conserved ACE2 gene co-expression across all GEO datasets, with dataset as random effect, to prioritize genes with stable correlations | 9,124 GEO series | not stated |
| Permutation test with Benjamini-Hochberg FDR correction | GENEVA scores across all GEO datasets to identify those with statistically significant ACE2 expression variance; FDR threshold 0.05, yielding 27 significant datasets | 286,650 samples across 9,124 datasets | not stated |
| Mixed-effects model (cardiomyopathy as fixed effect, dataset as random effect) | Meta-analysis of ACE2 expression across 7 cardiomyopathy datasets overall and for DCM specifically (Fig. 3A) | 7 datasets | not stated |
| t-test | ACE2 expression comparison between each cardiomyopathy subtype (HCM, RCM, LVNC, ARVC) and normal hearts (Fig. 3B–E) | not stated per subtype | not stated |
| Egger regression | Publication bias assessment across the 7 cardiomyopathy datasets (Additional file 1, Figure S2) | 7 datasets | na |
| Multivariable Cox proportional-hazards model | COVID-19 mortality risk comparing cardiomyopathy vs. other cardiovascular disease vs. no cardiovascular disease groups, controlling for age, gender, and pre-existing conditions (Fig. 4A) | 3,936 COVID-19 patients (cardiomyopathy N=43, other CVD N=624, no CVD N=3,269) | not stated |
| Log-rank test | Survival comparison in propensity score-matched cohorts: cardiomyopathy vs. matched other CVD in COVID-19-positive patients (Fig. 4B) and COVID-19-negative validation cohort (Fig. 4C) | 43 vs. 344 (COVID-19+); 2,250 vs. 18,000 (COVID-19-) | na |
| Propensity score matching | Constructing a matched comparison cohort of other cardiovascular disease patients to control for confounding when comparing mortality to cardiomyopathy patients | N=43 cardiomyopathy matched to N=344 other CVD (COVID-19+); N=2,250 matched to N=18,000 (COVID-19-) | not stated |
-
Individual cardiomyopathy subtype comparisons (HCM, RCM, LVNC, ARVC) were each assessed with a separate t-test without a stated multiplicity correction (Fig. 3B–E)↳ Could also: A single mixed-effects model including subtype as a categorical fixed effect, or application of FDR or Bonferroni correction across the five subtype comparisons, could also be used — When multiple comparisons share a common scientific question (ACE2 elevation across cardiomyopathy subtypes), a correction method would explicitly bound the family-wise error rate or FDR across those tests; a single model including subtype would also allow direct contrasts between subtypes
-
Pearson correlation was used to assess ACE2 co-expression both globally and within datasets↳ Could also: Spearman rank correlation could also be used — Spearman correlation makes no assumption of bivariate normality and is robust to outliers; with RNA-seq data that may retain skew even after normalization, a rank-based measure is also a standard and widely accepted alternative
-
Propensity score matching was used to construct a comparison cohort, discarding unmatched other-CVD patients (N=344 retained from N=624)↳ Could also: Inverse probability of treatment weighting (IPTW) could also be used — IPTW uses all available patients by reweighting rather than discarding unmatched controls, potentially retaining more statistical power and preserving representativeness of the full cardiovascular disease population, at the cost of dependence on correct propensity model specification
-
The Cox proportional-hazards model was used for multivariable survival analysis↳ Could also: A Fine-Gray subdistribution hazard model for competing risks could also be used — In a hospitalized COVID-19 cohort with multiple comorbidities, patients may die from causes other than COVID-19; the Fine-Gray model explicitly accounts for competing events when estimating the cumulative incidence of COVID-19-related death, and the two approaches can yield different estimates when competing events are common
-
ACE2 expression distributions across conditions were displayed as box plots↳ Could also: Overlaid individual data points (strip charts or jittered dot plots) alongside summary statistics could also be used — Showing individual observations is especially informative for smaller dataset-level group sizes (e.g., ARVC, LVNC), where the number of samples may be too small for a box plot to reliably convey the underlying distribution shape
-
Correlation coefficients were used as the primary measure of effect size for ACE2 co-expression, with p values deprioritized due to the very large sample size↳ Could also: Standardized mean differences or variance-explained metrics (R²) could also supplement correlation coefficients in the meta-analytic context — In a random-effects meta-analysis, the between-study heterogeneity statistic I² alongside the pooled effect size gives an explicit decomposition of how much variance is attributable to true effect heterogeneity vs. within-study sampling error, which is a standard complement to the pooled estimate
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Pre-existing cardiomyopathy is significantly associated with increased mortality risk in COVID-19 patients compared to those without cardiovascular diseaseother human-covid-19-patient up 2022×1papers★ This paper is the founder (earliest)
-
Meta-analysis across GEO datasets identified 27 experimental conditions significantly modulating ACE2 expression at FDR < 0.05other human 2022×1papers★ This paper is the founder (earliest)
-
ACE2 expression is upregulated by RAD140 treatment in breast cancer xenografts and by itraconazole treatment in colorectal cancer cell linesRNA-seq human-cancer-model up 2022×1papers★ This paper is the founder (earliest)
-
ACE2 expression is elevated in heart tissue of cardiomyopathy patients in a meta-analysis across seven independent datasetsRNA-seq human heart up 2022×1papers★ This paper is the founder (earliest)
-
FABP2 is the most strongly positively correlated gene with ACE2 across diverse human RNA-seq samplesRNA-seq human up 2022×1papers★ This paper is the founder (earliest)
-
HNF1A overexpression increases and knockdown decreases ACE2 expression in LNCaP cells, confirming HNF1A as a positive transcriptional regulator of ACE2RNA-seq lncap mixed 2022×1papers★ This paper is the founder (earliest)
-
HNF4G overexpression does not significantly change ACE2 expression in LNCaP cells despite a positive population-level correlationRNA-seq lncap none 2022×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35012625
Paper: Kaur N, Oskotsky B, Butte AJ, Hu Z (2022). Systematic identification of ACE2 expression modulators reveals cardiomyopathy as a risk factor for mortality in COVID-19 patients. Genome Biol 23:15. PMID 35012625 · PMCID PMC8743438 · DOI 10.1186/s13059-021-02589-4.
Code: https://github.com/NavchetanKaur/geneva-webtool — pinned commit
8686776b7fc0960a0044461139c492afb0 (master, not archived, no license file in repo; paper
cites GPL + Zenodo 10.5281/zenodo.5735451). It is a Django web app implementing the GENEVA
score; it does NOT ship the backend data (exp.csv, df_embed.csv, df_meta.csv,
df_var_mean.csv, gse_meta.csv, df_N.csv are read at runtime; db.sqlite3 is 0 bytes).
exp.csv = the ARCHS4-derived percentile-rank expression matrix (286,650 samples, uint8);
df_embed.csv = the precomputed Levenshtein+MDS 2-D metadata embedding. Neither is in the repo.
Data (brief's accession): GEO GSE89714 = "Differential gene expressions in the heart of hypertrophic cardiomyopathy patients", RNA-seq (Illumina HiSeq 2000), 5 HCM vs 4 normal-donor hearts. PUBLIC. This is the flagship example dataset highlighted in the paper (Fig 2A,B).
What the paper does (methods)
- GENEVA framework over ARCHS4 (286,650 GEO samples, 9124 series): percentile-rank
transform → ACE2 co-expression (Pearson + lme4 mixed model
Gene ~ ACE2 + (1|dataset)) → metadata embedding (concatenate sample metadata, pairwise Levenshtein, MDS to 2-D) → GENEVA score =VARg · R² / VARmper dataset, permutation p-values, FDR. Output: "27 significant datasets (FDR<0.05)" where ACE2 expression varies with a metadata axis. - ACE2 in cardiomyopathy: flagship dataset GSE89714 shows ACE2 ↑ in HCM (Fig 2). A
meta-analysis of 7 cardiomyopathy datasets (GSE89714, GSE120838, GSE121893, GSE63161,
GSE71613, GSE99321, GSE29819) with
ACE2 ~ cardiomyopathy_status + (1|study)→ ACE2 ↑, p<0.001; per-subtype (DCM/HCM/RCM/LVNC up; ARVC trend, n.s.) Fig 3. - Clinical: UCSF COVID-19 Data Mart (3,936 COVID+; 2,250 COVID−), ICD-10, Cox PH + propensity matching → cardiomyopathy associated with COVID-19 mortality (p=0.004; "438% increase in death rate" vs other CVD).
In scope (pipeline-derived, attempted on public data + «our HPC»)
| id | result | reported (paper) | pipeline | reproducibility |
|---|---|---|---|---|
| C1 | ACE2 upregulated in hypertrophic cardiomyopathy | "GSE89714 show upregulated expression of ACE2 in HCM" (Fig 2A,B); HCM (n=5) vs normal (n=4) | DE / group comparison of ACE2 in GSE89714 processed expression | YES — small public RNA-seq; direct 1:1 |
| C2 | ACE2 elevated in cardiomyopathy across multiple datasets | meta-analysis 7 datasets, ACE2 ↑ p<0.001; subtypes DCM/HCM/RCM/LVNC ↑ (Fig 3) | per-dataset ACE2 direction + mixed-effect `ACE2 ~ status + (1 | study)` |
Out of scope (not attempted — and why)
- Full GENEVA recomputation over ARCHS4 ("27 significant datasets FDR<0.05"; FABP2–ACE2
r=0.72 co-expression). The webtool's backend (
exp.csv= 286,650-sample percentile matrix,df_embed.csv) is NOT in the repo and is tied to a specific 2020 ARCHS4 snapshot; current ARCHS4 has >2× the samples → not version-reproducible. This is the hard ~80%; per the 80/20 rule we run the GENEVA score formula as documented but do not rebuild the backend. →docs_insufficient(backend data not shipped) for the full screen. - COVID-19 clinical mortality (Cox PH, propensity matching, "438% increase", p=0.004). Uses
the UCSF COVID-19 Data Mart EHR — controlled-access, not obtainable. →
data_restricted. - GSE29819 (Affymetrix CEL RAW.tar → needs affy/RMA normalization) and GSE121893 (single-nucleus UMI) — inc
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's biological pivot — ACE2 upregulated in cardiomyopathy heart tissue — reproduces decisively on public data (GSE89714 3.41× UP, Welch p=5.7e-4; pooled mixedlm p=2.0e-8 clears the reported p<0.001), with no factual deviation in any checkable value. The gaps are on the availability side, not the authors' computation: the flagship GENEVA/ARCHS4 method ships no backend in the repo, and the title's clinical mortality conclusion uses a restricted UCSF EHR — both un-auditable but not contradicted. The meta-analysis was also run on a self-defined 3-of-7 dataset subset. Net: a solid partial reproduction whose confirmed core is the biology, while the headline method and clinical conclusion remain untestable.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.