Transcriptome and machine learning analysis of the impact of COVID-19 on mitochondria and multiorgan damage.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH; reproduced essentially 1:1 via a clean, fully independent SLURM re-run («job», compute node n093, COMPLETED, 25:30 wall, ExitCode 0). The prior «infra» workdir had been reclaimed by the janitor, so the repo (ntumitolab/ML-RNA-Seq @ c6dbf5d, authors' own, MIT) was re-cloned, the conda env rebuilt from scratch with the paper's exact pinned versions (scikit-learn 1.2.2, imbalanced-learn 0.10.1, xgboost 1.7.5, shap 0.41.0, numpy 1.24.4, pandas 2.0.3, python 3.10.20), and the GSE152075 raw-counts CSV shipped inside the repo re-unzipped (484 samples = 430 COVID:54 control, SHA256 02bddd2b...). A driver replicates the authors' exact evaluate_model: 10-fold StratifiedKFold(shuffle, rs=42) + SMOTE(rs=42) on TRAIN folds only + pooled out-of-fold confusion-matrix metrics. Compared 44 reported ML table cells (Tables 2/4/5 = geneset x classifier) on the 6 confusion-matrix metrics: 30 EXACT, 1 within-tol, 12 partial, 1 mismatch. ALL 22 Logistic-Regression and SVM cells match to >=3 decimals (these models are feature-column-order-invariant, hence fully deterministic given identical data) — proving the data, labels, CV split, SMOTE, and metric definitions reproduce exactly. The 12 partial + 1 mismatch are XGBoost (plus a couple of RF on reconstructed IPA union lists): XGBoost/RF subsample features by column index, so they depend on gene-column ORDER and XGBoost build/threading, which the paper does not pin for the IPA-derived combined lists; deviations are small (Acc<=0.009, MCC<=0.056). Where the gene order IS pinned in the repo (Table 2 40/30/20), RF also matches exactly. Every reproduced value matched the earlier confirming runs (2181553/2181559) to 4 decimals — bit-for-bit deterministic. SHAP (Fig 6A) reproduced CXCL10 as the #1 feature and ATP5F1E in the top 4 (ACE2 ranks lower; 1-gene union reconstruction diff, 127 vs 126). ONE GENUINE GAP: the reported AUC column matches no standard CV-ROC aggregation (under-specified recipe; flagged, not a fabrication signal since all confusion-matrix metrics reproduce). NOT ATTEMPTED (out of scope, justified): Table 2 100/80/60/50-gene columns (top-100 not shipped), IPA toxicity-list derivation (proprietary), top-40 derivation (NetworkAnalyst web GUI), DAVID/ClueGO/GSEA figures. No completeness claimed; all grades provisional for human audit.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 74assessed: 2026-06-16 ⛓ 9b1e9d922058
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetCOVID-19 may cause damage to the heart, kidney, and liver via mitochondrial dysfunction and downstream responses, in addition to direct viral infection through ACE2 and inflammatory cytokine storms.
- ★ Potential cardiac, hepatic, and renal impairments in COVID-19 are associated with ACE2, inflammatory cytokine storms, and mitochondrial pathways. finding
- ★ Machine learning can serve as a verification tool to test biological hypotheses by incorporating gene list adjustment alongside statistically and biologically ranked differentially expressed genes. method
- ★ SARS-CoV-2 may damage the heart, kidney, and liver via mitochondrial infection/dysfunction as a mechanism distinct from direct viral invasion or cytokine storm alone. mechanism
- ★ Five publicly available GEO RNA-seq datasets (GSE152075, GSE163151, GSE157103, GSE169241, GSE152641) were curated as a resource for COVID-19 multiorgan transcriptome analysis. resource
- GSE169241 heart autopsy tissue shows increased CCL2 expression and macrophage infiltration linked to cardiomyocyte ROS and apoptosis via IL-6/TNF-α. finding
- A two-step machine learning classifier on nasopharyngeal swab RNA profiles (GSE163151) distinguished COVID-19 from non-COVID-19 acute respiratory disease with 85.1%-86.5% accuracy. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq / differential expression analysis | Nasopharyngeal swab (human, GSE152075) | SARS-CoV-2 infection (COVID vs control) | Differentially expressed genes, immune cell signatures | Illumina NextSeq 500 (GPL18573) |
| bulk RNA-seq / differential expression analysis | Nasopharyngeal swab (human, GSE163151) | SARS-CoV-2 infection (COVID vs control) | Differentially expressed genes; ML diagnostic classifier accuracy | Illumina NovaSeq 6000 (GPL24676) |
| bulk RNA-seq / differential expression analysis | Human heart autopsy tissue (GSE169241) | SARS-CoV-2 infection (COVID vs control) | CCL2 expression, macrophage infiltration, cardiomyocyte ROS/apoptosis | Illumina NovaSeq 6000 (GPL24676) |
| bulk RNA-seq / multiomic analysis | Leukocytes (human, GSE157103) | SARS-CoV-2 infection, ICU vs non-ICU severity | Coagulation-related protein levels (cFN, prothrombin) | Illumina NovaSeq 6000 (GPL24676) |
| bulk RNA-seq / differential expression analysis | Whole blood (human, GSE152641) | SARS-CoV-2 infection vs other viral infections vs control | Differentially expressed genes, immune cell composition | Illumina NovaSeq 6000 (GPL24676) |
| Pathway/toxicity enrichment analysis (IPA Tox, DAVID, ClueGO, GSEA, cnetplot) | Computational analysis of the five transcriptome sample sets | none | Enriched KEGG/REACTOME/WIKIPATHWAYS/GO pathways and toxicity lists | IPA v84978992; DAVID; ClueGO (Cytoscape); GSEA; clusterProfiler 4.8.3 (R 4.3.1) |
| Machine learning classification (XGBoost and related models) | Nasopharyngeal swab RNA-seq features (GSE152075, primary; GSE163151, GSE157103, GSE152641, compared) | none (classification of COVID vs control using DEG-derived features) | F1-Score, MCC, AUC model performance | Python 3.10.11; sklearn 1.2.2, imblearn 0.10.1, XGBoost 1.7.5 |
- – SARS-CoV-2 triggered an interferon-driven antiviral response and decreased transcription of ribosomal proteins in nasopharyngeal swabs.
- – Two-step ML classifier segregated COVID-19 from non-COVID-19 acute respiratory disease using nasal swab RNA profiling. 85.1%-86.5% accuracy
- ▲ COVID-19 heart autopsy tissue showed significant increase in CCL2 expression and macrophage infiltration.
- ▲ Macrophages exposed to SARS-CoV-2 induced increased cardiomyocyte reactive oxygen species and apoptosis via IL-6 and TNF-α secretion.
- ▲ Coagulation-related proteins, including cellular fibronectin (cFN), were significantly increased in COVID-19 leukocyte samples.
- ▼ Prothrombin abundance was significantly reduced and correlated with COVID-19 severity.
- – ACO1 and ATL3 showed expression changes specifically different between SARS-CoV-2 and other viral infections in whole blood.
- ▲ CD56bright NK cells, M2 macrophages, and total NK cells were increased in COVID-19 whole blood samples.
- other adjusted p value threshold 0.05; log2-fold change threshold 1.5 (Thresholds for defining differentially expressed genes via limma)
- count 484 samples (COVID:control = 430:54) (GSE152075 nasopharyngeal swab sample set)
- count 149 samples (COVID:control = 138:11) (GSE163151 nasopharyngeal swab sample set)
- count 8 samples (COVID:control = 3:5) (GSE169241 human heart autopsy sample set)
- count 126 samples (COVID:control = 100:26; ICU:non-ICU = 50:50) (GSE157103 leukocyte sample set)
- count 86 samples (COVID:control = 62:24) (GSE152641 whole blood sample set)
- other 85.1%-86.5% accuracy (Two-step ML classifier for COVID-19 vs non-COVID-19 respiratory disease on GSE163151)
- other 1000 permutations, phenotypic permutation type (GSEA parameters using h.all.v2023.1.Hs.symbols.gmt gene set)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study reanalyzed five publicly available RNA-seq datasets from GEO (per-dataset N ranging from 8 to 484) comparing COVID-19 patients to controls across nasopharyngeal swab, leukocyte, whole-blood, and heart-autopsy tissues. Differential expression was identified with limma on Log2-CPM normalized counts (adjusted p < 0.05, |log2FC| ≥ 1.5), followed by parallel pathway enrichment analyses (IPA, DAVID, GSEA, ClueGO) and Venn-diagram gene overlap. XGBoost classification on the largest dataset (GSE152075, n = 484) served as a machine learning verification step, evaluated by F1-score, MCC, and AUC.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma linear model with empirical Bayes moderation (via NetworkAnalyst), applied to Log2-CPM normalized counts | Differential gene expression in all five GEO datasets (GSE152075, GSE163151, GSE157103, GSE169241, GSE152641) | 484, 149, 126, 86, and 8 respectively, as stated in Table 1 | not stated |
| Gene Set Enrichment Analysis (GSEA) with 1000 phenotype permutations, h.all.v2023.1.Hs.symbols.gmt gene set database | Assessment of predefined gene-set distributions across phenotype-ranked gene lists for each dataset | — | not stated |
| ClueGO pathway enrichment (GO, KEGG, WIKIPATHWAYS, REACTOME Reactions and Pathways; p-value threshold 0.05; GO hierarchy levels 8–15; minimum 2 genes and 6% per pathway) | Functional annotation and pathway grouping of DEGs across datasets | — | not stated |
| DAVID functional annotation and enrichment (KEGG, REACTOME, WIKIPATHWAYS) | Pathway enrichment of significantly expressed genes across all transcriptomes | — | not stated |
| XGBoost classification evaluated by F1-score, MCC, and AUC | Machine learning verification of DEG-derived and biologically ranked gene features in GSE152075 | 484 (COVID:control = 430:54) | not stated |
-
Differential expression was performed using limma on Log2-CPM normalized counts via NetworkAnalyst↳ Could also: DESeq2 or edgeR applied directly to raw integer counts using negative binomial models; or limma-voom, which applies precision weights to count data before the linear model step — DESeq2 and edgeR model count overdispersion explicitly and are frequently recommended as primary methods for bulk RNA-seq; limma-voom is a widely used hybrid; benchmark studies often compare all three, and reporting which was primary and which served as a sensitivity check is a common practice
-
Multiple parallel pathway enrichment tools (IPA, DAVID, ClueGO, GSEA) were applied without a stated primary test or cross-tool multiplicity correction↳ Could also: Designate one tool as the primary enrichment method (e.g., fgsea with BH-FDR) and use the remaining tools explicitly as secondary or confirmatory analyses — Formalizing a primary test clarifies the inferential framework and limits the risk of selectively highlighting findings that appear in at least one tool; treating additional tools as sensitivity or replication analyses is a widely used practice for transparency
-
Gene overlap across datasets and organs was identified and visualized using Venn diagrams↳ Could also: Hypergeometric test or Fisher's exact test to quantify whether the observed overlaps exceed what is expected by chance given the background gene universe — A statistical test on the overlap provides a probability-based measure for prioritizing shared genes rather than relying on visual inspection alone, which is especially useful when background gene-list sizes differ across datasets
-
GSEA used phenotype permutation with 1000 permutations↳ Could also: Gene-set permutation (permuting gene labels rather than sample labels) — Phenotype permutation requires adequate sample size for stable null distributions; for small datasets such as GSE169241 (n = 8), gene-set permutation is generally preferred as it is less sensitive to sample size, though it is more conservative
-
Class imbalance in machine learning (e.g., COVID:control = 430:54 in GSE152075) was addressed by considering SMOTE via imblearn↳ Could also: Class-weighted loss (XGBoost scale_pos_weight parameter) or stratified k-fold cross-validation without oversampling — Class weighting adjusts the learning objective directly without generating synthetic minority-class samples, avoiding potential overfitting to interpolated points; stratified k-fold ensures each fold preserves the original class ratio, which is a common complementary practice
-
Machine learning performance was evaluated with F1-score, MCC, and ROC-AUC↳ Could also: Precision-recall AUC (AUPRC) and calibration curves (reliability diagrams) as additional or primary metrics — For heavily imbalanced class distributions, AUPRC is more sensitive to minority-class performance than ROC-AUC, which can appear high even with poor minority-class recall; calibration curves assess whether model-predicted probabilities reflect observed event rates, which matters if predicted scores are used clinically
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38295140
Paper: Chang YY, Wei AC. Transcriptome and machine learning analysis of the impact of COVID-19 on mitochondria and multiorgan damage. PLoS One 2024. PMID 38295140 · PMCID PMC10830027 · DOI 10.1371/journal.pone.0297664.
Code: https://github.com/ntumitolab/ML-RNA-Seq (authors' own; MIT; commit
c6dbf5dcba2c4d694d91d6ddd0ad3e7547a00671, pushed 2023-12-19).
The repo ships the input data inside it (Data/GSE152075_raw_counts...csv.zip
etc.) → ideal: authors' own code + authors' own data.
Data: GEO GSE152075 — nasopharyngeal swab RNA-seq, 484 samples
(COVID:control = 430:54). Shipped as a CSV in the repo (genes × samples + a
target row encoding 1=COVID / 0=control).
Pipeline-derived results (IN SCOPE — attempt)
The core computational result is supervised ML classification of COVID vs control
from gene-expression features, evaluated by 10-fold stratified CV with SMOTE
oversampling of the training folds only. Pipeline = scikit-learn / imbalanced-learn
/ XGBoost. Fully deterministic in the code (random_state=42 everywhere; the
random-gene baseline uses random_state=28).
| ID | Reported in | What | Pinning | Confidence |
|---|---|---|---|---|
| T2-40 | Table 2, "40 genes" col | Acc/Sens/Spec/Prec/F1/MCC/AUC for XGBoost, RF, LogReg, SVM on the top-40 sig genes | 40-gene list hardcoded in Code/rna_seq_of_gse152075_for_40_sig_genes_and_validation...py |
HIGH — exact list + seeds + hyperparams |
| T2-rand40 | Table 2, "Randomly selected 40 genes" col | same metrics, 4 models | df.sample(n=40, random_state=28, axis=1) in ..._for_random_select_40_genes.py |
HIGH — fixed seed |
| T2-30, T2-20 | Table 2, "30"/"20 genes" cols | same metrics | first 30 / 20 of the ranked top-40 list (ordering assumption) | MED — list ordering assumption |
| T4 | Table 4 | same metrics on each IPA toxicity list (mito 32, heart 38, renal 55, liver 42) | gene lists from Table 3 (text) / 126-gene union hardcoded in tox script | MED — gene sets reconstructed from Table 3 |
| T5 | Table 5 | same metrics on mito+heart(66), mito+renal(83), mito+liver(73) | unions of Table 3 lists | MED |
| SHAP | Fig 6A | top SHAP features = CXCL10, ATP5F1E, ACE2 | 126-gene XGBoost in tox script | MED — qualitative top-feature ranking |
Primary FLOOR target: T2-40 (4 models × 7 metrics = 28 cells), the most clearly specified, fully-pinned result. Then push outward to T2-rand40/30/20, T4, T5, SHAP.
OUT OF SCOPE (not attempted — proprietary/manual/web tools)
- Derivation of the top-40 sig-gene list via NetworkAnalyst (limma, Log2-CPM, adj-p<0.05, |log2FC|>1.5). NetworkAnalyst is a web GUI; the result (the 40-gene list) is hardcoded in the repo, so we consume it as given rather than re-deriving. (Optional secondary: re-derive DEGs with limma on the raw counts to check the list.)
- IPA toxicity-list derivation (Fig 4, Table 3 list membership) — Ingenuity Pathway Analysis is proprietary (licensed). The resulting gene lists are in the text/code, so the ML on them is in scope; their derivation is not.
- DAVID / ClueGO / GSEA pathway figures (Fig 2B, 6B) — web/GUI, qualitative.
- The 100/80/60/50-gene columns of Table 2 — require the full ranked top-100 list, which is not shipped in the repo (only the top-40 is). Recorded as a gap, not attempted as 1:1.
Compute plan
All on «our HPC» (SLURM, account=kubisch_std, partition=std). conda env built inside
the job (python 3.10.11, scikit-learn 1.2.2, imbalanced-learn 0.10.1, xgboost
1.7.5, shap 0.41.0 — matching the paper's stated versions). Repo cloned + data
unzipped on «infra». A single driver replicates the authors' exact
evaluate_model (SMOTE-in-CV, confusion-matrix metrics, mean per-fold AUC).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a strong reproduction run on the authors' own code and data (GSE152075, 484=430:54 exact): 30/44 table cells exact and every order-invariant Logistic-Regression and SVM cell matches to 3 decimals, confirming data, labels, CV split, SMOTE and metric definitions reproduce 1:1. The non-exact cells are all XGBoost/RF and stem from gene-column-order and build/threading nondeterminism the paper doesn't pin — technical, expected, sub-1% deviations on our/implementation side, not authors' defects. One genuine authors-side under-specification: the reported AUC column matches no standard CV-ROC aggregation (deviations 0.01–0.03, model-dependent sign). The central conclusion holds fully; overall yellow only because of the explainable XGBoost deviations and the unpinned AUC recipe.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.