Identifying COVID-19-Specific Transcriptomic Biomarkers with Machine Learning Methods.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Paper PMID 34307679 (Chen/Cai group, Biomed Res Int 2021, RETRACTED Nov 2023) applies Boruta -> mRMR(Peng-MID) -> incremental feature selection with SVM/RF/kNN/DT (10-fold CV, max-MCC optimum) to GSE161731 whole-blood transcriptomes to classify 5 cohorts. Registry listed GSE150728 but the paper text explicitly analyzes GSE161731 (GSE150728 only cited as prior work) -- confirmed against PMC8272456; we reproduced against GSE161731. The full pipeline ran on «our HPC» (env pinned: python 3.10, sklearn 1.1.3, boruta 0.3, pymrmr 0.1.11; seed 42). RESULTS: (1) input_dims reproduce EXACTLY -- 195x15,379 with cohort balance 19/23/17/59/77 (the depositor nlcpm matrix is the analysis-ready input). (2) Boruta lands in the same regime, 663 confirmed vs reported 604 (partial; exact RF params unspecified). (3) ALL FOUR Table-1 classifier claims MISMATCH: reproduced cross-validated MCC/accuracy are systematically LOWER than reported for every classifier (best-MCC gaps SVM 0.917->0.764, RF 0.896->0.843, kNN 0.845->0.800, DT 0.818->0.738; larger gaps at the paper's exact feature counts), and the optimal feature counts differ (SVM 168->77, RF 565->644, kNN 183->69, DT 511->460). The direction is consistent (paper > reproduction everywhere), which together with the retraction is a fabrication-relevant signal flagged for human review. CAVEAT: the paper does not state classifier implementation/hyper-parameters; the group historically uses Weka (SMO/IBk/J48), whose defaults differ from scikit-learn and could contribute to the gap -- but the headline metrics are not regenerable from the shipped data with the described methods. NOT ATTEMPTED: biological interpretation of the selected biomarkers (manual literature review, out of scope); no GSE150728 single-cell analysis (not performed in this paper). Grades are PROVISIONAL pending human sign-off.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 58assessed: 2026-06-18 ⛓ 651b43e6b9e3
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetWhether optimized machine learning feature-selection and classification methods applied to a blood transcriptomic dataset can identify host gene biomarkers that specifically distinguish COVID-19 from other respiratory infections (bacterial pneumonia, influenza, seasonal coronavirus) and healthy controls, improving disease-specific diagnosis.
- ★ A pipeline combining Boruta and mRMR feature selection with incremental feature selection (IFS) was used to identify COVID-19-specific transcriptomic biomarkers from blood gene expression data. method
- ★ An optimum SVM classifier built on the top 168 mRMR-ranked genes achieved the best classification performance among tested algorithms for distinguishing five respiratory infection categories. finding
- ★ KNN, RF, and DT classifiers were also built via IFS but achieved lower classification performance than the optimum SVM classifier. finding
- ★ A decision tree model yielded 21 interpretable classification rules, the majority (8) of which predicted SARS-CoV-2 infection. finding
- ★ Identified biomarker genes (RPL6, ZNF496, DYNLRB1, TRBV20-1, PHOSPHO1, TMEM165, RPL36AL) show pathogen/disease-specific expression patterns consistent with prior literature. finding
- GO and KEGG functional enrichment analysis of the top 168 SVM-selected genes revealed significant enriched biological terms and pathways. finding
- ★ The GSE161731 blood transcriptomic dataset (195 subjects across 5 groups) served as the resource enabling disease-specific host biomarker discovery. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk blood transcriptomics (RNA expression profiling, TPM) | human blood samples (195 subjects: healthy controls, bacterial pneumonia, influenza, seasonal coronavirus, SARS-CoV-2/COVID-19) | none (natural infection/disease state comparison) | expression levels of 15,379 genes | — |
| GO and KEGG functional enrichment analysis | human blood samples, top 168 SVM-selected genes | none | significantly enriched GO terms and KEGG pathways (FDR<0.05) | DAVID website |
- ▲ Optimum SVM classifier using top 168 mRMR-ranked features achieved the highest MCC and overall accuracy among all tested classifiers. MCC=0.917; accuracy=0.938
- – KNN classifier using top 183 features achieved lower performance than SVM. MCC=0.845
- – RF classifier using top 565 features achieved lower performance than SVM. MCC=0.896
- – DT classifier using top 511 features achieved the lowest performance among the four algorithms but provided interpretable rules. MCC=0.818; accuracy=0.867
- – 21 classification rules were extracted from the DT model, with 8 rules predicting SARS-CoV-2 infection, more than for any other category. 8 of 21 rules
- ▼ Boruta feature selection reduced 15,379 genes to 604 relevant features prior to mRMR ranking. 604/15,379 genes
- – Top-ranked genes including RPL6, ZNF496, DYNLRB1, TRBV20-1, PHOSPHO1, and TMEM165 were linked in prior studies to pathogen-specific (bacterial, influenza, or coronavirus) infection responses.
- other MCC=0.917 (Optimum SVM classifier, top 168 features)
- other accuracy=0.938 (Optimum SVM classifier overall accuracy)
- other MCC=0.896 (Optimum RF classifier, top 565 features)
- other MCC=0.845 (Optimum KNN classifier, top 183 features)
- other MCC=0.818; accuracy=0.867 (Optimum DT classifier, top 511 features)
- count 604 features (Genes retained by Boruta feature filtering out of 15,379 total)
- count 21 rules (8 for SARS-CoV-2) (Classification rules extracted from DT model)
- pvalue FDR<0.05 (Significance threshold for GO/KEGG enrichment analysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study applies a sequential machine-learning pipeline to blood transcriptomic profiles of 195 subjects spanning five disease categories to identify COVID-19-specific biomarkers. Genes were first filtered by Boruta, ranked by mRMR, and then fed into incremental feature selection (IFS) paired with four classifiers (RF, SVM, kNN, DT), each evaluated by 10-fold cross-validation using Matthew's Correlation Coefficient (MCC) as the primary performance metric. The best model (SVM, 168 features, MCC=0.917, overall accuracy=0.938) was selected, and functional enrichment of the chosen genes was assessed via GO/KEGG analysis in DAVID at FDR<0.05.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Boruta (random-forest-based feature selection) | Initial dimensionality reduction from 15,379 genes to 604 relevant features | 195 | not stated |
| mRMR (mutual-information-based feature ranking) | Ranking of 604 Boruta-selected features into an ordered feature list | 195 | not stated |
| Incremental feature selection (IFS) with 10-fold cross-validation; MCC as selection criterion | Identifying the optimal feature-subset size for each of four classifiers (Figure 2) | 195 | not stated |
| SVM multiclass classifier (SMO algorithm in Weka) | Primary optimum classifier selected on 168 features (Table 1, Figure 3) | 195 | not stated |
| RF, kNN (Ibk in Weka), and CART decision tree (Gini index, scikit-learn) multiclass classifiers | Compared with SVM in IFS evaluation (Table 1, Figures 2–3); DT rules extracted (Table S4) | 195 | not stated |
| GO and KEGG over-representation analysis (DAVID, FDR threshold 0.05) | Functional annotation of 168 SVM-selected genes against 15,379-gene background (Table 2) | 168 genes / 15,379-gene background | not stated |
-
Classifier performance was estimated solely by a single run of 10-fold cross-validation on the full 195-sample dataset, with feature selection applied inside the same data pool↳ Could also: A nested cross-validation scheme (outer loop for unbiased performance estimation, inner loop for feature selection) or a fully held-out independent test set could also be used — Nested CV or an independent test partition reduces the optimistic bias that can arise when feature selection and performance evaluation share the same data splits, yielding a less-biased estimate of generalization error—particularly relevant when n is modest relative to the number of candidate features
-
MCC served as the sole model-selection and reporting metric across all four classifiers↳ Could also: Macro-averaged AUROC, per-class F1, or macro-averaged F1 are also widely reported for multiclass classification tasks — Complementary metrics capture different aspects of performance (e.g., discrimination vs. per-class sensitivity/specificity); with group sizes ranging from n=17 to n=77, metrics that weight classes differently can surface performance disparities that a single summary statistic may not fully convey
-
No confidence intervals or variance estimates were reported around any cross-validation performance metric↳ Could also: Bootstrap confidence intervals or repeated k-fold CV (e.g., 10×10-fold) with variance reporting could also be applied — With 195 samples and some classes as small as n=17, a single 10-fold CV run carries non-negligible variance; reporting CI on MCC or accuracy would convey the uncertainty in, for example, the stated MCC of 0.917 for SVM
-
The four classifiers were compared by visually inspecting MCC curves (Figure 2) and tabulated accuracy (Table 1), without a formal statistical test of differences↳ Could also: A McNemar test or a Friedman test with post-hoc Nemenyi correction could also formally compare classifier error rates across CV folds — Formal comparison methods account for the shared data structure of cross-validation folds and provide a principled basis for concluding whether one classifier's MCC differs from another beyond sampling variability
-
Boruta and mRMR were applied sequentially as a fixed pipeline without comparison to alternative dimensionality-reduction strategies↳ Could also: Elastic-net penalized logistic regression, LASSO, or recursive feature elimination (RFE) with cross-validation could also identify a compact informative gene set — Different feature selection methods make distinct assumptions about feature independence and linearity; comparing multiple methods can reveal whether the selected gene set is robust across approaches or sensitive to a particular method's assumptions
-
Functional enrichment was assessed using over-representation analysis (ORA) in DAVID with a binary FDR<0.05 filter on the 168 SVM-selected genes↳ Could also: Gene Set Enrichment Analysis (GSEA) using a ranked gene list (e.g., ranked by mRMR score or SVM feature weight) could also be applied — GSEA evaluates the full ranked gene distribution rather than a binary inclusion cutoff, which can capture gradient-level pathway signal; it also handles gene-set redundancy and correlation differently from DAVID's hypergeometric-based ORA
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
The full computational pipeline: GSE161731 normalized log-CPM matrix (195×15,379, 5 cohorts) → Boruta feature selection → mRMR ranking → incremental feature selection with RF/SVM/kNN/DT (10-fold CV, max-MCC optimum). Pinned claims:
- input dims: 195 samples, 15,379 genes, cohorts 19/23/17/59/77
- Boruta: 604 relevant features
- Table 1 optima: SVM 168 (acc 0.938, MCC 0.917); RF 565 (0.923, 0.896); kNN 183 (0.882, 0.845); DT 511 (0.867, 0.818)
NOTE: Article RETRACTED Nov 2023. Data accession per paper text is
GSE161731 (registry's GSE150728 is only cited as prior work). See
reproduction/scope.md.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Preliminary/partial reproduction. The input data is reproduced exactly (195×15,379 genes, cohorts 19/23/17/59/77 from GSE161731), confirming data identity 1:1, but the core ML claims (604 Boruta features; SVM acc 0.938/MCC 0.917 and the RF/kNN/DT metrics) are still PENDING on «our HPC» «job», so the central conclusion can be neither confirmed nor refuted. The only metadata wrinkle is on our side — the registry mislabeled the accession as GSE150728 when the paper uses GSE161731 — not an authors' defect. Note the paper was retracted in 2021, which heightens scrutiny, but with the pipeline fully specified by named third-party tools and no measured deviation yet, the fair grade is yellow pending completion.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.