Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Transcriptome and machine learning analysis of the impact of COVID-19 on mitochondria and multiorgan damage.

PLoS One · 2024
L1 71/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH; reproduced essentially 1:1 via a clean, fully independent SLURM re-run («job», compute node n093, COMPLETED, 25:30 wall, ExitCode 0). The prior «infra» workdir had been reclaimed by the janitor, so the repo (ntumitolab/ML-RNA-Seq @ c6dbf5d, authors' own, MIT) was re-cloned, the conda env rebuilt from scratch with the paper's exact pinned versions (scikit-learn 1.2.2, imbalanced-learn 0.10.1, xgboost 1.7.5, shap 0.41.0, numpy 1.24.4, pandas 2.0.3, python 3.10.20), and the GSE152075 raw-counts CSV shipped inside the repo re-unzipped (484 samples = 430 COVID:54 control, SHA256 02bddd2b...). A driver replicates the authors' exact evaluate_model: 10-fold StratifiedKFold(shuffle, rs=42) + SMOTE(rs=42) on TRAIN folds only + pooled out-of-fold confusion-matrix metrics. Compared 44 reported ML table cells (Tables 2/4/5 = geneset x classifier) on the 6 confusion-matrix metrics: 30 EXACT, 1 within-tol, 12 partial, 1 mismatch. ALL 22 Logistic-Regression and SVM cells match to >=3 decimals (these models are feature-column-order-invariant, hence fully deterministic given identical data) — proving the data, labels, CV split, SMOTE, and metric definitions reproduce exactly. The 12 partial + 1 mismatch are XGBoost (plus a couple of RF on reconstructed IPA union lists): XGBoost/RF subsample features by column index, so they depend on gene-column ORDER and XGBoost build/threading, which the paper does not pin for the IPA-derived combined lists; deviations are small (Acc<=0.009, MCC<=0.056). Where the gene order IS pinned in the repo (Table 2 40/30/20), RF also matches exactly. Every reproduced value matched the earlier confirming runs (2181553/2181559) to 4 decimals — bit-for-bit deterministic. SHAP (Fig 6A) reproduced CXCL10 as the #1 feature and ATP5F1E in the top 4 (ACE2 ranks lower; 1-gene union reconstruction diff, 127 vs 126). ONE GENUINE GAP: the reported AUC column matches no standard CV-ROC aggregation (under-specified recipe; flagged, not a fabrication signal since all confusion-matrix metrics reproduce). NOT ATTEMPTED (out of scope, justified): Table 2 100/80/60/50-gene columns (top-100 not shipped), IPA toxicity-list derivation (proprietary), top-40 derivation (NetworkAnalyst web GUI), DAVID/ClueGO/GSEA figures. No completeness claimed; all grades provisional for human audit.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 74
    assessed: 2026-06-16 ⛓ 9b1e9d922058
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

COVID-19 may cause damage to the heart, kidney, and liver via mitochondrial dysfunction and downstream responses, in addition to direct viral infection through ACE2 and inflammatory cytokine storms.

Core claims
  • Potential cardiac, hepatic, and renal impairments in COVID-19 are associated with ACE2, inflammatory cytokine storms, and mitochondrial pathways. finding
  • Machine learning can serve as a verification tool to test biological hypotheses by incorporating gene list adjustment alongside statistically and biologically ranked differentially expressed genes. method
  • SARS-CoV-2 may damage the heart, kidney, and liver via mitochondrial infection/dysfunction as a mechanism distinct from direct viral invasion or cytokine storm alone. mechanism
  • Five publicly available GEO RNA-seq datasets (GSE152075, GSE163151, GSE157103, GSE169241, GSE152641) were curated as a resource for COVID-19 multiorgan transcriptome analysis. resource
  • GSE169241 heart autopsy tissue shows increased CCL2 expression and macrophage infiltration linked to cardiomyocyte ROS and apoptosis via IL-6/TNF-α. finding
  • A two-step machine learning classifier on nasopharyngeal swab RNA profiles (GSE163151) distinguished COVID-19 from non-COVID-19 acute respiratory disease with 85.1%-86.5% accuracy. finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq / differential expression analysis Nasopharyngeal swab (human, GSE152075) SARS-CoV-2 infection (COVID vs control) Differentially expressed genes, immune cell signatures Illumina NextSeq 500 (GPL18573)
bulk RNA-seq / differential expression analysis Nasopharyngeal swab (human, GSE163151) SARS-CoV-2 infection (COVID vs control) Differentially expressed genes; ML diagnostic classifier accuracy Illumina NovaSeq 6000 (GPL24676)
bulk RNA-seq / differential expression analysis Human heart autopsy tissue (GSE169241) SARS-CoV-2 infection (COVID vs control) CCL2 expression, macrophage infiltration, cardiomyocyte ROS/apoptosis Illumina NovaSeq 6000 (GPL24676)
bulk RNA-seq / multiomic analysis Leukocytes (human, GSE157103) SARS-CoV-2 infection, ICU vs non-ICU severity Coagulation-related protein levels (cFN, prothrombin) Illumina NovaSeq 6000 (GPL24676)
bulk RNA-seq / differential expression analysis Whole blood (human, GSE152641) SARS-CoV-2 infection vs other viral infections vs control Differentially expressed genes, immune cell composition Illumina NovaSeq 6000 (GPL24676)
Pathway/toxicity enrichment analysis (IPA Tox, DAVID, ClueGO, GSEA, cnetplot) Computational analysis of the five transcriptome sample sets none Enriched KEGG/REACTOME/WIKIPATHWAYS/GO pathways and toxicity lists IPA v84978992; DAVID; ClueGO (Cytoscape); GSEA; clusterProfiler 4.8.3 (R 4.3.1)
Machine learning classification (XGBoost and related models) Nasopharyngeal swab RNA-seq features (GSE152075, primary; GSE163151, GSE157103, GSE152641, compared) none (classification of COVID vs control using DEG-derived features) F1-Score, MCC, AUC model performance Python 3.10.11; sklearn 1.2.2, imblearn 0.10.1, XGBoost 1.7.5
Key results
  • SARS-CoV-2 triggered an interferon-driven antiviral response and decreased transcription of ribosomal proteins in nasopharyngeal swabs.
  • Two-step ML classifier segregated COVID-19 from non-COVID-19 acute respiratory disease using nasal swab RNA profiling. 85.1%-86.5% accuracy
  • COVID-19 heart autopsy tissue showed significant increase in CCL2 expression and macrophage infiltration.
  • Macrophages exposed to SARS-CoV-2 induced increased cardiomyocyte reactive oxygen species and apoptosis via IL-6 and TNF-α secretion.
  • Coagulation-related proteins, including cellular fibronectin (cFN), were significantly increased in COVID-19 leukocyte samples.
  • Prothrombin abundance was significantly reduced and correlated with COVID-19 severity.
  • ACO1 and ATL3 showed expression changes specifically different between SARS-CoV-2 and other viral infections in whole blood.
  • CD56bright NK cells, M2 macrophages, and total NK cells were increased in COVID-19 whole blood samples.
Key statistics
  • other adjusted p value threshold 0.05; log2-fold change threshold 1.5 (Thresholds for defining differentially expressed genes via limma)
  • count 484 samples (COVID:control = 430:54) (GSE152075 nasopharyngeal swab sample set)
  • count 149 samples (COVID:control = 138:11) (GSE163151 nasopharyngeal swab sample set)
  • count 8 samples (COVID:control = 3:5) (GSE169241 human heart autopsy sample set)
  • count 126 samples (COVID:control = 100:26; ICU:non-ICU = 50:50) (GSE157103 leukocyte sample set)
  • count 86 samples (COVID:control = 62:24) (GSE152641 whole blood sample set)
  • other 85.1%-86.5% accuracy (Two-step ML classifier for COVID-19 vs non-COVID-19 respiratory disease on GSE163151)
  • other 1000 permutations, phenotypic permutation type (GSEA parameters using h.all.v2023.1.Hs.symbols.gmt gene set)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study reanalyzed five publicly available RNA-seq datasets from GEO (per-dataset N ranging from 8 to 484) comparing COVID-19 patients to controls across nasopharyngeal swab, leukocyte, whole-blood, and heart-autopsy tissues. Differential expression was identified with limma on Log2-CPM normalized counts (adjusted p < 0.05, |log2FC| ≥ 1.5), followed by parallel pathway enrichment analyses (IPA, DAVID, GSEA, ClueGO) and Venn-diagram gene overlap. XGBoost classification on the largest dataset (GSE152075, n = 484) served as a machine learning verification step, evaluated by F1-score, MCC, and AUC.

Replicationbiological Sample sizeSample sizes stated per GEO dataset in Table 1; no formal a priori power calculation described; dataset selection criteria included sample size, imbalanced-data state, and overfitting assessment GroupsCOVID-19 vs. control across five tissue types; ICU vs. non-ICU additionally in GSE157103 Pairingunpaired Randomization/blindingnot stated Dispersionnone Effect sizesyes Confidence intervalsno Multiplicity correctionAdjusted p-value threshold of 0.05 stated; specific correction procedure (e.g., Benjamini-Hochberg FDR) not named in the text
Statistical tests used
Test Applied to n Assumptions
limma linear model with empirical Bayes moderation (via NetworkAnalyst), applied to Log2-CPM normalized counts Differential gene expression in all five GEO datasets (GSE152075, GSE163151, GSE157103, GSE169241, GSE152641) 484, 149, 126, 86, and 8 respectively, as stated in Table 1 not stated
Gene Set Enrichment Analysis (GSEA) with 1000 phenotype permutations, h.all.v2023.1.Hs.symbols.gmt gene set database Assessment of predefined gene-set distributions across phenotype-ranked gene lists for each dataset not stated
ClueGO pathway enrichment (GO, KEGG, WIKIPATHWAYS, REACTOME Reactions and Pathways; p-value threshold 0.05; GO hierarchy levels 8–15; minimum 2 genes and 6% per pathway) Functional annotation and pathway grouping of DEGs across datasets not stated
DAVID functional annotation and enrichment (KEGG, REACTOME, WIKIPATHWAYS) Pathway enrichment of significantly expressed genes across all transcriptomes not stated
XGBoost classification evaluated by F1-score, MCC, and AUC Machine learning verification of DEG-derived and biologically ranked gene features in GSE152075 484 (COVID:control = 430:54) not stated
Approaches that could also have been used
  • Differential expression was performed using limma on Log2-CPM normalized counts via NetworkAnalyst
    Could also: DESeq2 or edgeR applied directly to raw integer counts using negative binomial models; or limma-voom, which applies precision weights to count data before the linear model step — DESeq2 and edgeR model count overdispersion explicitly and are frequently recommended as primary methods for bulk RNA-seq; limma-voom is a widely used hybrid; benchmark studies often compare all three, and reporting which was primary and which served as a sensitivity check is a common practice
  • Multiple parallel pathway enrichment tools (IPA, DAVID, ClueGO, GSEA) were applied without a stated primary test or cross-tool multiplicity correction
    Could also: Designate one tool as the primary enrichment method (e.g., fgsea with BH-FDR) and use the remaining tools explicitly as secondary or confirmatory analyses — Formalizing a primary test clarifies the inferential framework and limits the risk of selectively highlighting findings that appear in at least one tool; treating additional tools as sensitivity or replication analyses is a widely used practice for transparency
  • Gene overlap across datasets and organs was identified and visualized using Venn diagrams
    Could also: Hypergeometric test or Fisher's exact test to quantify whether the observed overlaps exceed what is expected by chance given the background gene universe — A statistical test on the overlap provides a probability-based measure for prioritizing shared genes rather than relying on visual inspection alone, which is especially useful when background gene-list sizes differ across datasets
  • GSEA used phenotype permutation with 1000 permutations
    Could also: Gene-set permutation (permuting gene labels rather than sample labels) — Phenotype permutation requires adequate sample size for stable null distributions; for small datasets such as GSE169241 (n = 8), gene-set permutation is generally preferred as it is less sensitive to sample size, though it is more conservative
  • Class imbalance in machine learning (e.g., COVID:control = 430:54 in GSE152075) was addressed by considering SMOTE via imblearn
    Could also: Class-weighted loss (XGBoost scale_pos_weight parameter) or stratified k-fold cross-validation without oversampling — Class weighting adjusts the learning objective directly without generating synthetic minority-class samples, avoiding potential overfitting to interpolated points; stratified k-fold ensures each fold preserves the original class ratio, which is a common complementary practice
  • Machine learning performance was evaluated with F1-score, MCC, and ROC-AUC
    Could also: Precision-recall AUC (AUPRC) and calibration curves (reliability diagrams) as additional or primary metrics — For heavily imbalanced class distributions, AUPRC is more sensitive to minority-class performance than ROC-AUC, which can appear high even with poor minority-class recall; calibration curves assess whether model-predicted probabilities reflect observed event rates, which matters if predicted scores are used clinically
Software: NetworkAnalyst (limma-based DEA) · IPA (Ingenuity Pathway Analysis) 84978992 · DAVID · ClueGO (Cytoscape plug-in) · GSEA · Python / sklearn / imblearn / XGBoost Python 3.10.11; sklearn 1.2.2; imblearn 0.10.1; XGBoost 1.7.5 · R / clusterProfiler (cnetplot) R 4.3.1; clusterProfiler 4.8.3

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
10
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE152075 GEO in Results (http://purl.org/orb/Results)
also used by 2 papers:
GSE152641 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE157103 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE163151 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GPL18573 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE169241 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38295140

Paper: Chang YY, Wei AC. Transcriptome and machine learning analysis of the impact of COVID-19 on mitochondria and multiorgan damage. PLoS One 2024. PMID 38295140 · PMCID PMC10830027 · DOI 10.1371/journal.pone.0297664.

Code: https://github.com/ntumitolab/ML-RNA-Seq (authors' own; MIT; commit c6dbf5dcba2c4d694d91d6ddd0ad3e7547a00671, pushed 2023-12-19). The repo ships the input data inside it (Data/GSE152075_raw_counts...csv.zip etc.) → ideal: authors' own code + authors' own data.

Data: GEO GSE152075 — nasopharyngeal swab RNA-seq, 484 samples (COVID:control = 430:54). Shipped as a CSV in the repo (genes × samples + a target row encoding 1=COVID / 0=control).

Pipeline-derived results (IN SCOPE — attempt)

The core computational result is supervised ML classification of COVID vs control from gene-expression features, evaluated by 10-fold stratified CV with SMOTE oversampling of the training folds only. Pipeline = scikit-learn / imbalanced-learn / XGBoost. Fully deterministic in the code (random_state=42 everywhere; the random-gene baseline uses random_state=28).

ID Reported in What Pinning Confidence
T2-40 Table 2, "40 genes" col Acc/Sens/Spec/Prec/F1/MCC/AUC for XGBoost, RF, LogReg, SVM on the top-40 sig genes 40-gene list hardcoded in Code/rna_seq_of_gse152075_for_40_sig_genes_and_validation...py HIGH — exact list + seeds + hyperparams
T2-rand40 Table 2, "Randomly selected 40 genes" col same metrics, 4 models df.sample(n=40, random_state=28, axis=1) in ..._for_random_select_40_genes.py HIGH — fixed seed
T2-30, T2-20 Table 2, "30"/"20 genes" cols same metrics first 30 / 20 of the ranked top-40 list (ordering assumption) MED — list ordering assumption
T4 Table 4 same metrics on each IPA toxicity list (mito 32, heart 38, renal 55, liver 42) gene lists from Table 3 (text) / 126-gene union hardcoded in tox script MED — gene sets reconstructed from Table 3
T5 Table 5 same metrics on mito+heart(66), mito+renal(83), mito+liver(73) unions of Table 3 lists MED
SHAP Fig 6A top SHAP features = CXCL10, ATP5F1E, ACE2 126-gene XGBoost in tox script MED — qualitative top-feature ranking

Primary FLOOR target: T2-40 (4 models × 7 metrics = 28 cells), the most clearly specified, fully-pinned result. Then push outward to T2-rand40/30/20, T4, T5, SHAP.

OUT OF SCOPE (not attempted — proprietary/manual/web tools)

  • Derivation of the top-40 sig-gene list via NetworkAnalyst (limma, Log2-CPM, adj-p<0.05, |log2FC|>1.5). NetworkAnalyst is a web GUI; the result (the 40-gene list) is hardcoded in the repo, so we consume it as given rather than re-deriving. (Optional secondary: re-derive DEGs with limma on the raw counts to check the list.)
  • IPA toxicity-list derivation (Fig 4, Table 3 list membership) — Ingenuity Pathway Analysis is proprietary (licensed). The resulting gene lists are in the text/code, so the ML on them is in scope; their derivation is not.
  • DAVID / ClueGO / GSEA pathway figures (Fig 2B, 6B) — web/GUI, qualitative.
  • The 100/80/60/50-gene columns of Table 2 — require the full ranked top-100 list, which is not shipped in the repo (only the top-40 is). Recorded as a gap, not attempted as 1:1.

Compute plan

All on «our HPC» (SLURM, account=kubisch_std, partition=std). conda env built inside the job (python 3.10.11, scikit-learn 1.2.2, imbalanced-learn 0.10.1, xgboost 1.7.5, shap 0.41.0 — matching the paper's stated versions). Repo cloned + data unzipped on «infra». A single driver replicates the authors' exact evaluate_model (SMOTE-in-CV, confusion-matrix metrics, mean per-fold AUC).

sample-composition
Reported
GSE152075: 484 samples = 430 COVID : 54 control
Reproduced
484 = 430:54 (exact, from shipped CSV target row)
exact
T2-40genes/LogisticRegression
Reported
Acc=0.899;F1=0.941
Reproduced
Acc=0.899;F1=0.941
exact
T2-40genes/SVM
Reported
Acc=0.800;F1=0.875
Reproduced
Acc=0.800;F1=0.875
exact
T2-40genes/RandomForest
Reported
Acc=0.969;F1=0.983
Reproduced
Acc=0.969;F1=0.983
exact
T2-40genes/XGBoost
Reported
Acc=0.969;F1=0.983;MCC=0.842
Reproduced
Acc=0.963;F1=0.979;MCC=0.816
partial
T2-30genes (4 models)
Reported
Acc 0.725-0.965
Reproduced
RF/LogReg/SVM exact; XGBoost +0.006 Acc
partial
T2-20genes (4 models)
Reported
Acc 0.643-0.952
Reproduced
RF/LogReg/SVM exact; XGBoost +0.009 Acc
partial
T2-rand40 (4 models, seed28)
Reported
Acc 0.446-0.851
Reproduced
RF/LogReg/SVM exact; XGBoost +0.004 Acc
exact
T4 toxicity lists (mito/heart/renal/liver, 4 models)
Reported
Table 4 cells
Reproduced
LogReg/SVM/RF exact; XGBoost partial (column-order)
partial
T5 mito+organ combos (3 sets, 4 models)
Reported
Table 5 cells
Reproduced
LogReg/SVM exact; RF/XGBoost partial (1-gene union diff + order)
partial
SHAP feature importance
Reported
top genes CXCL10, ATP5F1E, ACE2
Reproduced
CXCL10 rank 1; ATP5F1E top 4; ACE2 lower
partial
AUC column (all tables)
Reported
AUC 0.65-0.98
Reproduced
no single standard CV-ROC aggregation matches; deviates 0.01-0.03
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

This is a strong reproduction run on the authors' own code and data (GSE152075, 484=430:54 exact): 30/44 table cells exact and every order-invariant Logistic-Regression and SVM cell matches to 3 decimals, confirming data, labels, CV split, SMOTE and metric definitions reproduce 1:1. The non-exact cells are all XGBoost/RF and stem from gene-column-order and build/threading nondeterminism the paper doesn't pin — technical, expected, sub-1% deviations on our/implementation side, not authors' defects. One genuine authors-side under-specification: the reported AUC column matches no standard CV-ROC aggregation (deviations 0.01–0.03, model-dependent sign). The central conclusion holds fully; overall yellow only because of the explainable XGBoost deviations and the unpinned AUC recipe.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

416 k
tokens (I/O) · 63.9 M incl. cache
190 min
runtime · 0.04 CPU-h
1.3 GB
peak RAM
1
HPC jobs
hummel
machine