Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Systematic identification of ACE2 expression modulators reveals cardiomyopathy as a risk factor for mortality in COVID-19 patients.

Genome Biol · 2022
L1 68/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
68/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 32% of all assessed papers rank 765 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the paper's BIOLOGICAL PIVOT 1:1 on public data; its headline METHOD and CLINICAL CONCLUSION are not reproducible from shipped/public artifacts. Reproduced («our HPC»/«infra», GEO data fetched in-job): (C1) the flagship claim that ACE2 is upregulated in hypertrophic cardiomyopathy (GSE89714, Fig 2A/B) -> ACE2 RPKM 31.38(HCM,n=5) vs 9.20(normal,n=4), 3.41x up, Welch p=5.7e-4, MWU p=0.016 -- direction matches decisively. (C2) the cardiomyopathy meta-analysis claim 'ACE2 significantly elevated, p<0.001' on the obtainable public subset of the 7 datasets: GSE89714 (HCM) 3.41x up, GSE71613 (RCM/DCM) 2.27x up, GSE99321 (DCM) 2.17x up; combined mixed-effect model beta=1.42 p=2.0e-8, all 3/3 datasets up. NOT attempted and why: (a) the GENEVA screen over ARCHS4 ('27 significant datasets FDR<0.05'; FABP2-ACE2 r=0.72) -- the geneva-webtool repo is a Django app that does NOT ship its backend (exp.csv 286,650-sample ARCHS4 percentile matrix, df_embed.csv) and is tied to a 2020 ARCHS4 snapshot now superseded, so not version-reproducible (the hard ~80%); (b) the COVID-19 clinical mortality results (Cox PH, propensity matching, p=0.004, '438% increase') use the UCSF COVID-19 Data Mart EHR = controlled-access (data_restricted); (c) GSE63161 (hiPSC-CM) downloaded but has no ACE2 row (filtered in-vitro), GSE29819 (Affy CEL/RMA), GSE120838, GSE121893 (single-cell) skipped per 80/20; LVNC/ARVC subtypes not separable in the public subset. Fabrication check: no discrepancy -- every publicly-checkable directional claim holds and the pooled significance clears p<0.001; GENEVA/clinical numbers are un-auditable from public artifacts, not contradicted. Verdict provisional; human audit sheet in AUDIT.md.

💻 Code ↗ 🗄 Data: GSE89714

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 68
    assessed: 2026-06-15 ⛓ 9eaf4798063b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

ACE2, the SARS-CoV-2 cell-entry receptor, may have its expression modulated by specific diseases, drugs, and genetic perturbations; systematically identifying these modulators across large-scale RNA-seq data could reveal risk factors for severe COVID-19, with the specific hypothesis that pre-existing cardiomyopathy (with elevated cardiac ACE2) leads to increased COVID-19 mortality.

Core claims
  • GENEVA is a semi-automated, study-design-agnostic framework that mines large-scale public RNA-seq data to identify conditions modulating a gene of interest's expression method
  • Multiple drugs, genetic perturbations, and diseases are associated with ACE2 expression, including cardiomyopathy, HNF1A overexpression, RAD140, and itraconazole finding
  • ACE2 expression is significantly elevated in heart tissue across all major cardiomyopathy categories (DCM, HCM, RCM, LVNC, and trend in ARVC) finding
  • COVID-19 patients with pre-existing cardiomyopathy have an increased mortality risk relative to patients with other cardiovascular conditions and patients without cardiovascular disease finding
  • HNF1A regulates ACE2 expression, with overexpression increasing and knockdown reducing ACE2 in LNCaP prostate cancer cells mechanism
  • ACE2 lacks a single conserved transcriptional network and is co-regulated with different gene sets depending on tissue/condition finding
  • GENEVA is provided as a freely accessible web application applicable to any gene of interest at genevatool.org resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (meta-analysis/co-expression) 286,650 human RNA-seq samples from 9124 GEO series (ARCHS4) none/various (existing conditions) ACE2 expression and Pearson correlation with other genes ARCHS4 uniformly preprocessed RNA-seq
bulk RNA-seq (validation of tissue effect) GTEx human tissues none ACE2 co-expression correlations across tissues
RNA-seq (genetic perturbation) LNCaP prostate cancer cell line HNF1A overexpression and knockdown; HNF4G overexpression ACE2 expression change
RNA-seq (disease comparison, GSE89714) human heart tissue (normal vs hypertrophic cardiomyopathy) disease (HCM) ACE2 expression
RNA-seq (drug treatment, GSE104177) human breast cancer xenografts RAD140 (selective androgen receptor modulator) ACE2 expression
RNA-seq (drug treatment, GSE114013) colorectal cancer cell lines HT55 and SW948 itraconazole (antifungal) ACE2 expression
meta-analysis (mixed-effect model) of RNA-seq and microarray 7 cardiomyopathy heart-tissue datasets (5 RNA-seq + microarray) cardiomyopathy (DCM, HCM, RCM, LVNC, ARVC) ACE2 expression with publication-bias (Egger) assessment
clinical survival analysis (Cox proportional-hazards, Kaplan-Meier, propensity matching) 3936 COVID-19 patients from UCSF EHR pre-existing cardiomyopathy vs other/no cardiovascular disease mortality / risk of death UCSF electronic health records
Key results
  • ACE2 expression significantly elevated in heart tissue of cardiomyopathy patients across 7 datasets p < 0.001
  • Cardiomyopathy significantly associated with risk of death vs patients without cardiovascular disease (Cox model) p = 0.004
  • Cardiomyopathy significantly associated with risk of death vs patients with other cardiovascular diseases p = 0.038, 438% increase in observed death rate
  • FABP2 is the gene most highly correlated with ACE2 across all samples r = 0.72
  • HNF1A overexpression increases and knockdown reduces ACE2 expression in LNCaP cells
  • HNF4G overexpression does not significantly increase ACE2 (trend toward reduction) despite positive correlation
  • RAD140 induces ACE2 expression in breast cancer xenografts; itraconazole upregulates ACE2 in HT55 and SW948
  • GENEVA identified 27 significant datasets (FDR < 0.05) as ACE2 modulators FDR < 0.05
Key statistics
  • correlation 0.72 (highest correlation between FABP2 and ACE2 across all samples)
  • correlation 0.42 (p value < 0.0001) (agreement of ACE2 co-expression between GTEx and GEO data)
  • pvalue < 0.001 (elevated ACE2 in cardiomyopathy heart tissue, 7-dataset meta-analysis)
  • pvalue 0.004 (Cox model: cardiomyopathy vs no cardiovascular disease, mortality)
  • pvalue 0.038 (Cox model: cardiomyopathy vs other cardiovascular diseases, mortality)
  • fold_change 438% increase in observed death rate (cardiomyopathy vs other cardiovascular disease patients)
  • count 27 significant datasets (FDR < 0.05) (GENEVA-identified ACE2-modulating datasets)
  • count 43 / 624 / 3269 (COVID-19 patients: cardiomyopathy / other cardiovascular / no cardiovascular disease)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study uses a multi-component statistical design: (1) large-scale Pearson correlation and mixed-effects modeling across 286,650 public RNA-seq samples to map ACE2 co-expression networks; (2) the GENEVA framework, which applies permutation-based variance scoring with Benjamini-Hochberg FDR correction across 9,124 GEO datasets to identify significant ACE2 modulators; (3) a mixed-effects meta-analysis of seven cardiomyopathy RNA-seq datasets supplemented by t-tests within individual cardiomyopathy subtypes; and (4) multivariable Cox proportional-hazards regression and log-rank tests on propensity score-matched EHR data from 3,936 COVID-19 patients to evaluate mortality risk associated with pre-existing cardiomyopathy.

Replicationbiological Sample sizeSample sizes stated per group for EHR analysis (N=43, 624, 3,269); total RNA-seq samples stated (286,650 across 9,124 series); 7 datasets for cardiomyopathy meta-analysis; no formal power calculation described GroupsCardiomyopathy vs. other cardiovascular diseases vs. no cardiovascular disease in COVID-19 EHR cohort; cardiomyopathy subtypes (DCM, HCM, RCM, LVNC, ARVC) vs. normal heart in RNA-seq meta-analysis; experimental vs. control conditions across 27 significant GEO datasets Pairingunpaired Randomization/blindingnot stated Dispersionunclear Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
Pearson correlation ACE2 co-expression with all human genes across 286,650 samples (global) and within each GEO dataset; also used to validate agreement between GTEx and GEO co-expression profiles 286,650 samples from 9,124 GEO series (global); r=0.42 between GTEx and GEO correlation vectors not stated
Mixed-effects model (random-effects meta-analytic regression) Estimating conserved ACE2 gene co-expression across all GEO datasets, with dataset as random effect, to prioritize genes with stable correlations 9,124 GEO series not stated
Permutation test with Benjamini-Hochberg FDR correction GENEVA scores across all GEO datasets to identify those with statistically significant ACE2 expression variance; FDR threshold 0.05, yielding 27 significant datasets 286,650 samples across 9,124 datasets not stated
Mixed-effects model (cardiomyopathy as fixed effect, dataset as random effect) Meta-analysis of ACE2 expression across 7 cardiomyopathy datasets overall and for DCM specifically (Fig. 3A) 7 datasets not stated
t-test ACE2 expression comparison between each cardiomyopathy subtype (HCM, RCM, LVNC, ARVC) and normal hearts (Fig. 3B–E) not stated per subtype not stated
Egger regression Publication bias assessment across the 7 cardiomyopathy datasets (Additional file 1, Figure S2) 7 datasets na
Multivariable Cox proportional-hazards model COVID-19 mortality risk comparing cardiomyopathy vs. other cardiovascular disease vs. no cardiovascular disease groups, controlling for age, gender, and pre-existing conditions (Fig. 4A) 3,936 COVID-19 patients (cardiomyopathy N=43, other CVD N=624, no CVD N=3,269) not stated
Log-rank test Survival comparison in propensity score-matched cohorts: cardiomyopathy vs. matched other CVD in COVID-19-positive patients (Fig. 4B) and COVID-19-negative validation cohort (Fig. 4C) 43 vs. 344 (COVID-19+); 2,250 vs. 18,000 (COVID-19-) na
Propensity score matching Constructing a matched comparison cohort of other cardiovascular disease patients to control for confounding when comparing mortality to cardiomyopathy patients N=43 cardiomyopathy matched to N=344 other CVD (COVID-19+); N=2,250 matched to N=18,000 (COVID-19-) not stated
Approaches that could also have been used
  • Individual cardiomyopathy subtype comparisons (HCM, RCM, LVNC, ARVC) were each assessed with a separate t-test without a stated multiplicity correction (Fig. 3B–E)
    Could also: A single mixed-effects model including subtype as a categorical fixed effect, or application of FDR or Bonferroni correction across the five subtype comparisons, could also be used — When multiple comparisons share a common scientific question (ACE2 elevation across cardiomyopathy subtypes), a correction method would explicitly bound the family-wise error rate or FDR across those tests; a single model including subtype would also allow direct contrasts between subtypes
  • Pearson correlation was used to assess ACE2 co-expression both globally and within datasets
    Could also: Spearman rank correlation could also be used — Spearman correlation makes no assumption of bivariate normality and is robust to outliers; with RNA-seq data that may retain skew even after normalization, a rank-based measure is also a standard and widely accepted alternative
  • Propensity score matching was used to construct a comparison cohort, discarding unmatched other-CVD patients (N=344 retained from N=624)
    Could also: Inverse probability of treatment weighting (IPTW) could also be used — IPTW uses all available patients by reweighting rather than discarding unmatched controls, potentially retaining more statistical power and preserving representativeness of the full cardiovascular disease population, at the cost of dependence on correct propensity model specification
  • The Cox proportional-hazards model was used for multivariable survival analysis
    Could also: A Fine-Gray subdistribution hazard model for competing risks could also be used — In a hospitalized COVID-19 cohort with multiple comorbidities, patients may die from causes other than COVID-19; the Fine-Gray model explicitly accounts for competing events when estimating the cumulative incidence of COVID-19-related death, and the two approaches can yield different estimates when competing events are common
  • ACE2 expression distributions across conditions were displayed as box plots
    Could also: Overlaid individual data points (strip charts or jittered dot plots) alongside summary statistics could also be used — Showing individual observations is especially informative for smaller dataset-level group sizes (e.g., ARVC, LVNC), where the number of samples may be too small for a box plot to reliably convey the underlying distribution shape
  • Correlation coefficients were used as the primary measure of effect size for ACE2 co-expression, with p values deprioritized due to the very large sample size
    Could also: Standardized mean differences or variance-explained metrics (R²) could also supplement correlation coefficients in the meta-analytic context — In a random-effects meta-analysis, the between-study heterogeneity statistic I² alongside the pooled effect size gives an explicit decomposition of how much variance is attributable to true effect heterogeneity vs. within-study sampling error, which is a standard complement to the pooled estimate
Software: ARCHS4 · GENEVA (custom framework, genevatool.org)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Authors · 4
1Navchetan Kaur 2Boris Oskotsky 3Atul J. Butte 4Zicheng Hu
Citations
28
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (1)

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GO:0003700 Gene Ontology (GO) in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE104177 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE114013 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE89714 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35012625

Paper: Kaur N, Oskotsky B, Butte AJ, Hu Z (2022). Systematic identification of ACE2 expression modulators reveals cardiomyopathy as a risk factor for mortality in COVID-19 patients. Genome Biol 23:15. PMID 35012625 · PMCID PMC8743438 · DOI 10.1186/s13059-021-02589-4.

Code: https://github.com/NavchetanKaur/geneva-webtool — pinned commit 8686776b7fc0960a0044461139c492afb0 (master, not archived, no license file in repo; paper cites GPL + Zenodo 10.5281/zenodo.5735451). It is a Django web app implementing the GENEVA score; it does NOT ship the backend data (exp.csv, df_embed.csv, df_meta.csv, df_var_mean.csv, gse_meta.csv, df_N.csv are read at runtime; db.sqlite3 is 0 bytes). exp.csv = the ARCHS4-derived percentile-rank expression matrix (286,650 samples, uint8); df_embed.csv = the precomputed Levenshtein+MDS 2-D metadata embedding. Neither is in the repo.

Data (brief's accession): GEO GSE89714 = "Differential gene expressions in the heart of hypertrophic cardiomyopathy patients", RNA-seq (Illumina HiSeq 2000), 5 HCM vs 4 normal-donor hearts. PUBLIC. This is the flagship example dataset highlighted in the paper (Fig 2A,B).

What the paper does (methods)

  1. GENEVA framework over ARCHS4 (286,650 GEO samples, 9124 series): percentile-rank transform → ACE2 co-expression (Pearson + lme4 mixed model Gene ~ ACE2 + (1|dataset)) → metadata embedding (concatenate sample metadata, pairwise Levenshtein, MDS to 2-D) → GENEVA score = VARg · R² / VARm per dataset, permutation p-values, FDR. Output: "27 significant datasets (FDR<0.05)" where ACE2 expression varies with a metadata axis.
  2. ACE2 in cardiomyopathy: flagship dataset GSE89714 shows ACE2 ↑ in HCM (Fig 2). A meta-analysis of 7 cardiomyopathy datasets (GSE89714, GSE120838, GSE121893, GSE63161, GSE71613, GSE99321, GSE29819) with ACE2 ~ cardiomyopathy_status + (1|study) → ACE2 ↑, p<0.001; per-subtype (DCM/HCM/RCM/LVNC up; ARVC trend, n.s.) Fig 3.
  3. Clinical: UCSF COVID-19 Data Mart (3,936 COVID+; 2,250 COVID−), ICD-10, Cox PH + propensity matching → cardiomyopathy associated with COVID-19 mortality (p=0.004; "438% increase in death rate" vs other CVD).

In scope (pipeline-derived, attempted on public data + «our HPC»)

id result reported (paper) pipeline reproducibility
C1 ACE2 upregulated in hypertrophic cardiomyopathy "GSE89714 show upregulated expression of ACE2 in HCM" (Fig 2A,B); HCM (n=5) vs normal (n=4) DE / group comparison of ACE2 in GSE89714 processed expression YES — small public RNA-seq; direct 1:1
C2 ACE2 elevated in cardiomyopathy across multiple datasets meta-analysis 7 datasets, ACE2 ↑ p<0.001; subtypes DCM/HCM/RCM/LVNC ↑ (Fig 3) per-dataset ACE2 direction + mixed-effect `ACE2 ~ status + (1 study)`

Out of scope (not attempted — and why)

  • Full GENEVA recomputation over ARCHS4 ("27 significant datasets FDR<0.05"; FABP2–ACE2 r=0.72 co-expression). The webtool's backend (exp.csv = 286,650-sample percentile matrix, df_embed.csv) is NOT in the repo and is tied to a specific 2020 ARCHS4 snapshot; current ARCHS4 has >2× the samples → not version-reproducible. This is the hard ~80%; per the 80/20 rule we run the GENEVA score formula as documented but do not rebuild the backend. → docs_insufficient (backend data not shipped) for the full screen.
  • COVID-19 clinical mortality (Cox PH, propensity matching, "438% increase", p=0.004). Uses the UCSF COVID-19 Data Mart EHR — controlled-access, not obtainable. → data_restricted.
  • GSE29819 (Affymetrix CEL RAW.tar → needs affy/RMA normalization) and GSE121893 (single-nucleus UMI) — inc
Figures / tables: Fig 2AFig 3
C1
Reported
GSE89714: ACE2 upregulated in hypertrophic cardiomyopathy (Fig 2A,B; directional)
Reproduced
HCM mean 31.38 (n=5) vs normal 9.20 (n=4) RPKM; 3.41x up; Welch p=5.7e-4; MWU p=0.016
within tolerance
C2a
Reported
ACE2 elevated in cardiomyopathy meta-analysis (p<0.001) - RCM/DCM (GSE71613)
Reproduced
disease 37.90 (n=4) vs control 16.69 (n=4); 2.27x up; Welch p=0.139 n.s.
partial
C2b
Reported
ACE2 elevated in cardiomyopathy meta-analysis (p<0.001) - DCM (GSE99321)
Reproduced
failing 34.38 (n=7) vs control 15.83 (n=7); 2.17x up; Welch p=0.033
within tolerance
C2c
Reported
ACE2 significantly elevated in cardiomyopathy across datasets (p value < 0.001)
Reproduced
mixedlm z(log2 ACE2) ~ status + (1|study) on {GSE89714,GSE71613,GSE99321}: beta=1.42, p=2.0e-8; 3/3 datasets up
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 68/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

The paper's biological pivot — ACE2 upregulated in cardiomyopathy heart tissue — reproduces decisively on public data (GSE89714 3.41× UP, Welch p=5.7e-4; pooled mixedlm p=2.0e-8 clears the reported p<0.001), with no factual deviation in any checkable value. The gaps are on the availability side, not the authors' computation: the flagship GENEVA/ARCHS4 method ships no backend in the repo, and the title's clinical mortality conclusion uses a restricted UCSF EHR — both un-auditable but not contradicted. The meta-analysis was also run on a self-defined 3-of-7 dataset subset. Net: a solid partial reproduction whose confirmed core is the biology, while the headline method and clinical conclusion remain untestable.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

192.2 k
tokens (I/O) · 11 M incl. cache
19 min
runtime · 0 CPU-h
0.3 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine