Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Interpretable and integrative analysis of single-cell multiomics with scMKL.

Commun Biol · 2025
L1 65/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
65/100
Reproducibility score
0.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 27% of all assessed papers rank 843 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce: YES for the authors' own tool (scMKL, P16-valid), which ships a deterministic tutorial (random_state=5) running the full MKL pipeline on the paper's own MCF-7 data (Ors et al. 2022 = GSE154873, 1000-cell subset) with committed expected tables + result pickles. Installed scmkl from pinned commit 8118bd7 (v0.4.3) on «our HPC» and ran both example pipelines verbatim. RESULT: the ATAC tutorial reproduces ESSENTIALLY EXACTLY (max |dAUROC|=0.0003 over the full 10-point alpha grid; identical selected-group counts and top groups; alpha_star=0.77 matches). The RNA tutorial's committed numbers do NOT match a fresh run of the current code (reproduced AUROCs systematically higher, peak 0.992 vs 0.987, top group ESTROGEN_RESPONSE_EARLY vs LATE) -- but the committed notebook table and shipped MCF7_RNA_results.pkl agree with EACH OTHER while both disagree with the re-run, which indicates the RNA tutorial outputs are STALE relative to current code (RNA-default version drift; paper=0.1.6, pinned main=0.4.3; ATAC notebook pins scale_data/n_features and is unaffected). The PAPER's headline MCF-7 finding nonetheless reproduces: ~0.99 AUROC and estrogen-response as the top discriminating pathway. No fabrication signal -- all values are derivable from shipped data+code; the RNA discrepancy is intra-repo tutorial version drift. NOT ATTEMPTED (80/20): full-dataset figures (74k-cell LUAD GSE136246, LUSC, SLL, PCa, full MCF-7/T-47D multiome -- multi-GB downloads + long nested CV), the EasyMKL speed/memory benchmark (hardware-dependent), and transfer-learning/ablation/silhouette figures (full-data only).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 65
    assessed: 2026-06-14 ⛓ 9028b9343c09
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can an inherently interpretable multiple kernel learning framework (scMKL) integrate single-cell RNA and ATAC data using prior biological knowledge to accurately classify healthy versus cancerous cell states while transparently identifying the regulatory pathways and transcription factors driving those distinctions?

Core claims
  • scMKL combines multiple kernel learning with random Fourier features and group Lasso to jointly model transcriptomic and epigenomic single-cell data interpretably method
  • scMKL outperforms state-of-the-art supervised and unsupervised classifiers (MLP, XGBoost, SVM, EasyMKL) in accuracy across seven datasets and four cancer types finding
  • scMKL identifies interpretable transcriptomic, epigenetic, and multimodal pathway features driving cell-state classification without post-hoc explanation finding
  • scMKL uses random Fourier features to reduce kernel complexity from O(N^2) to O(N), enabling scalability to single-cell resolution method
  • TFBS-informed peak groupings (JASPAR) improve ATAC classification performance over Hallmark gene-set groupings finding
  • scMKL enables transfer learning, leveraging insights from one dataset to inform pathway analysis distinguishing treatment responses, tumor grades, and subtypes in new datasets finding
  • scMKL accurately identifies estrogen-response pathways and transcription factors in MCF-7 breast cancer cells finding
  • scMKL is robust to preprocessing strategy, maintaining high AUROC across diverse normalization workflows finding
Experimental setups
Assay System Perturbation Readout Platform
10x Multiome (scRNA-seq + scATAC-seq) MCF-7 breast cancer cell line estrogen treatment vs control binary classification of control vs estrogen-treated cells (AUROC) and pathway/TF weights
10x Multiome (scRNA-seq + scATAC-seq) T-47D breast cancer cell line estrogen treatment vs control binary classification of control vs estrogen-treated cells
10x Multiome (scRNA-seq + scATAC-seq) small lymphatic lymphoma (SLL) patient samples none (healthy vs tumor) classification of healthy vs tumor cells
scRNA-seq prostate cancer (PCa) patient cells none classification of non-malignant vs malignant cells
sciATAC-seq (scATAC-seq) prostate cancer (PCa) patient tumors none classification of Gleason 3 vs Gleason 4 grade tumors
scRNA-seq lung adenocarcinoma (LUAD) none classification of healthy vs cancer cells
scRNA-seq lung squamous cell carcinoma (LUSC) none classification of healthy vs cancer cells
Key results
  • scMKL achieved significantly better classification accuracy than MLP, XGBoost, and SVM despite using fewer (Hallmark) genes p<0.001
  • On the smallest dataset scMKL achieved superior accuracy while training faster and using less memory than EasyMKL 7x faster, 12x less memory
  • JASPAR TF groupings produced higher AUROC than Hallmark across all datasets ~2% (MCF-7, T-47D), ~1% (SLL), ~4% (PCa) median AUROC improvement
  • Cistrome TF groupings decreased ATAC performance for MCF-7 and T-47D 0.4% (MCF-7), 8% (T-47D)
  • TF-IDF normalization improved AUROC for MCF-7 and T-47D with Hallmark and JASPAR peak groups but decreased accuracy for other datasets 1-2%
  • scMKL remained consistently robust across preprocessing workflows (PCA, log-norm, raw, TF-IDF, LSI, binary) AUROC 0.9731-0.9999
  • scMKL performance declined on class-imbalanced LUAD and LUSC datasets class ratios 12:1 and 7:1
  • For ATAC data, SVM with LSI improved only marginally but required far more memory and compute ~18x more memory, 2x more compute time
Key statistics
  • other AUROC 0.9731–0.9999 (scMKL accuracy across datasets and preprocessing workflows)
  • fold_change 7x faster training, 12x less memory (scMKL vs EasyMKL on smallest dataset)
  • pvalue p<0.001 (scMKL significantly outperforming other algorithms)
  • other ~2%, ~1%, ~4% median AUROC improvement (JASPAR vs Hallmark groupings (MCF-7/T-47D, SLL, PCa))
  • count 100 independent models / 80-20 split repeated 100 times (nested cross-validation evaluation)
  • count 6438 cells, 36601 RNA features, 206167 ATAC features (MCF-7 multiome dataset)
  • count 74084 cells, 37698 RNA features (LUAD scRNA dataset, 68229/5855 class sizes)
  • fold_change 18x more memory, 2x more compute time (SVM with LSI on ATAC data)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

scMKL, a multiple-kernel-learning framework for single-cell multiomics, was benchmarked against MLP, XGBoost, SVM, and EasyMKL across seven cancer datasets using AUROC as the primary performance metric. Classification performance distributions were generated over 100 independent 80/20 train-test splits (different random seeds), with 4-fold cross-validation within each training set used to tune the regularization parameter λ. Pairwise performance differences between methods were assessed with Wilcoxon tests on the resulting distributions of 100 AUROC values per model per dataset. Significance was indicated by threshold-based asterisks only; supplementary metrics (F1, precision, recall) were reported in supplementary figures.

Replicationtechnical Sample size100 independent 80/20 train-test splits with different random seeds per model; 4-fold CV within each training set for λ tuning; per-class cell counts listed in Table 1 for each of seven datasets (ranging from 2,317 to 68,229 cells per class) GroupsscMKL vs. MLP, XGBoost, SVM, EasyMKL across seven datasets spanning four cancer types (breast, prostate, lymphoma, lung) and three sequencing platforms (scRNA-seq, scATAC-seq, multiome) Pairingunclear Randomization/blindingna DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Wilcoxon rank-sum test (exact variant and directionality not stated) Pairwise comparison of AUROC distributions between scMKL and benchmark classifiers (MLP, XGBoost, SVM) across seven datasets (Fig. 2a, b) 100 independent 80/20 train-test splits per model not stated
Wilcoxon rank-sum test Comparison of AUROC across ATAC peak groupings and data transformations (Hallmark vs. JASPAR vs. Cistrome; binary vs. TF-IDF) (Fig. 2e) 100 independent 80/20 train-test splits per model not stated
Approaches that could also have been used
  • Multiple Wilcoxon tests were performed across seven datasets and numerous method-pair comparisons without a stated multiple-testing correction
    Could also: Apply a false discovery rate procedure (e.g., Benjamini-Hochberg) or family-wise error rate correction (e.g., Bonferroni) across the full family of simultaneous comparisons — With many simultaneous hypothesis tests, the probability of at least one false positive grows; an explicit correction procedure controls this rate and is standard practice in multi-dataset benchmarking studies
  • AUROC was the sole primary metric for all seven datasets, including two with severe class imbalance (LUAD 12:1, LUSC 7:1)
    Could also: Also report area under the precision-recall curve (PR-AUC) or balanced accuracy for the class-imbalanced datasets — AUROC averages performance over all decision thresholds and can appear optimistic under severe imbalance; PR-AUC emphasizes the minority class and is widely recommended as a complementary metric in such settings
  • AUROC values from 100 repeated train-test splits on the same fixed dataset were compared with a Wilcoxon test
    Could also: Use a corrected repeated-holdout test (e.g., the Nadeau-Bengio variance correction) or a paired permutation test that explicitly accounts for the dependence among splits — AUROC values from repeated splits share training data and are therefore not independent; standard Wilcoxon tests assume independence, and corrected tests are available that adjust the variance estimate accordingly to avoid inflated type-I error
  • Exact p-values were not reported; significance was communicated only through threshold-based asterisk categories (*, **, ***)
    Could also: Report exact p-values alongside asterisk annotations — Exact p-values carry more information than categorical thresholds, allow readers to gauge effect magnitude in context, and align with increasingly common journal reporting guidelines advocating against p-value dichotomization
  • No rank-based or standardized effect size was reported alongside the Wilcoxon tests
    Could also: Report a rank-biserial correlation or common-language effect size for each Wilcoxon comparison — Statistical significance with n = 100 per model can be achieved for trivially small AUROC differences; an effect size quantifies practical magnitude independently of sample size and helps readers distinguish meaningful from negligible performance gaps
  • Dispersion around median AUROC was summarized as standard deviation (SD) in the perturbation analysis
    Could also: Report 95% bootstrap percentile confidence intervals (e.g., from the 100-split distribution) around median AUROC — CIs provide a direct interval estimate of estimation uncertainty that is more interpretable for quantifying performance precision and facilitate informal visual comparison between methods without requiring a separate significance test
Software: not stated in provided text

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 75/100
stands on reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40770488 (scMKL)

Paper: Kupp et al., "Interpretable and integrative analysis of single-cell multiomics with scMKL." Commun Biol 2025. PMID 40770488 / PMC12328712. Code: https://github.com/ohsu-cedar-comp-hub/scMKL (pinned commit 8118bd7ef50c609186e93ce85915cd935297790a, the authors' own tool). Paper-repro repo: scMKL_paper + Zenodo 10.5281/zenodo.15397923 (full-data).

What scMKL does (the pipeline)

Binary classification of single cells via Multiple Kernel Learning: features are grouped by prior knowledge (Hallmark gene sets for RNA; TFBS/region groupings for ATAC), each group becomes an RBF kernel approximated by Random Fourier Features, and Group Lasso (celer) selects informative groups while classifying. Sweeps a regularization grid (alpha), reports test AUROC, #selected groups, and top group per alpha. Deterministic given random_state.

In scope (attempted)

The repo ships a deterministic tutorial (example/RNA_analysis.ipynb, example/ATAC_analysis.ipynb) that runs the full scMKL pipeline on a 1,000-cell MCF-7 subset of the paper's own data (Ors et al. 2022 = GSE154873, one of the paper's seven datasets), with committed expected summary_df tables (random_state=5) and shipped result pickles (MCF7_{RNA,ATAC}_results.pkl).

  • C-RNA: MCF-7 RNA (Hallmark) AUROC / #groups / top-group across the alpha grid.
  • C-ATAC: MCF-7 ATAC (Hallmark, 5000 feats/group) AUROC / #groups / top-group.

These are pipeline-derived computational results on the paper's data, with an identifiable expected value → the clean 80/20 target. The RNA peak AUROC (~0.987) and top group (ESTROGEN_RESPONSE) are a faithful microcosm of the paper's MCF-7 finding (~0.98 AUROC; ER pathways top-ranked, Figs 2-3).

Out of scope (not attempted) — and why

  • Full-dataset figures (74k-cell LUAD GSE136246, LUSC, SLL, PCa scATAC/scRNA, full MCF-7/T-47D multiome): require multi-GB GEO/10x downloads + long nested-CV (100×80/20 splits) per dataset. This is the hard last ~20%; the tutorial subset already exercises the identical code path. Skipped per 80/20.
  • EasyMKL speed/memory benchmark (Fig 2 "7× faster, 12× less memory"): hardware- and implementation-dependent, not a portable reproducible number.
  • Wet-lab / data generation: none — paper reanalyzes public data only.
  • Transfer learning, ablation p-values, silhouette comparisons (Figs 3-6): derived from full-data runs; out of 80/20 scope.

Note on version

Paper used scmkl 0.1.6; committed tutorial outputs (our comparison target) are from current main (0.4.3). We install scmkl from the pinned source commit so our code == the code that produced the committed table → tests deterministic reproducibility of the tool on the paper's data. Flagged in AUDIT.md.

Figures / tables: Fig 2
atac_grid
Reported
MCF-7 ATAC peak AUROC 0.8837 @alpha=0.77; #groups 14..49; top group HALLMARK_TGF_BETA_SIGNALING (committed tutorial)
Reproduced
peak 0.8836 @alpha=0.77; identical #groups and top groups at every alpha; max |dAUROC|=0.0003
exact
rna_grid
Reported
MCF-7 RNA peak AUROC 0.9868 @alpha=0.53; top group ESTROGEN_RESPONSE_LATE; #groups 7..42 (committed tutorial)
Reproduced
peak 0.9921 @alpha=0.29; top group ESTROGEN_RESPONSE_EARLY; #groups 3..49; max |dAUROC|=0.064
did not match
paper_mcf7_science
Reported
Paper MCF-7 classification AUROC ~0.98-0.99 with estrogen-response pathways top-ranked (PMC12328712 Figs 2-3)
Reproduced
RNA peak AUROC 0.992 with HALLMARK_ESTROGEN_RESPONSE_EARLY top; ATAC peak 0.884
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 65/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

This is the authors' own tool (scMKL, P16-valid) run on its shipped deterministic MCF-7 tutorial. The ATAC pipeline reproduces essentially exactly (max ΔAUROC=0.0003 across the whole alpha grid, identical group counts and top groups), and the paper's headline finding (~0.99 AUROC, estrogen-response as the dominant pathway) reproduces. The only real deviation is that the repo's committed RNA tutorial outputs are stale (0.9868@0.53/LATE/30-groups vs fresh 0.9921@0.29/EARLY/19-groups, max ΔAUROC=0.064) due to intra-repo version drift, with both committed artifacts agreeing with each other — a documentation/version-hygiene issue on the authors' side, not fabrication, since every value is derivable from the shipped data+code. Overall yellow: a solid reproduction whose core claim is fully confirmed but with one explainable, non-trivial tutorial mismatch.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

97 k
tokens (I/O) · 7.9 M incl. cache
14 min
runtime · 0.13 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine