Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identifying COVID-19-Specific Transcriptomic Biomarkers with Machine Learning Methods.

Biomed Res Int · 2021
L1 32/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
32/100
Reproducibility score
2.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 2% of all assessed papers rank 1144 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Paper PMID 34307679 (Chen/Cai group, Biomed Res Int 2021, RETRACTED Nov 2023) applies Boruta -> mRMR(Peng-MID) -> incremental feature selection with SVM/RF/kNN/DT (10-fold CV, max-MCC optimum) to GSE161731 whole-blood transcriptomes to classify 5 cohorts. Registry listed GSE150728 but the paper text explicitly analyzes GSE161731 (GSE150728 only cited as prior work) -- confirmed against PMC8272456; we reproduced against GSE161731. The full pipeline ran on «our HPC» (env pinned: python 3.10, sklearn 1.1.3, boruta 0.3, pymrmr 0.1.11; seed 42). RESULTS: (1) input_dims reproduce EXACTLY -- 195x15,379 with cohort balance 19/23/17/59/77 (the depositor nlcpm matrix is the analysis-ready input). (2) Boruta lands in the same regime, 663 confirmed vs reported 604 (partial; exact RF params unspecified). (3) ALL FOUR Table-1 classifier claims MISMATCH: reproduced cross-validated MCC/accuracy are systematically LOWER than reported for every classifier (best-MCC gaps SVM 0.917->0.764, RF 0.896->0.843, kNN 0.845->0.800, DT 0.818->0.738; larger gaps at the paper's exact feature counts), and the optimal feature counts differ (SVM 168->77, RF 565->644, kNN 183->69, DT 511->460). The direction is consistent (paper > reproduction everywhere), which together with the retraction is a fabrication-relevant signal flagged for human review. CAVEAT: the paper does not state classifier implementation/hyper-parameters; the group historically uses Weka (SMO/IBk/J48), whose defaults differ from scikit-learn and could contribute to the gap -- but the headline metrics are not regenerable from the shipped data with the described methods. NOT ATTEMPTED: biological interpretation of the selected biomarkers (manual literature review, out of scope); no GSE150728 single-cell analysis (not performed in this paper). Grades are PROVISIONAL pending human sign-off.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 58
    assessed: 2026-06-18 ⛓ 651b43e6b9e3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Whether optimized machine learning feature-selection and classification methods applied to a blood transcriptomic dataset can identify host gene biomarkers that specifically distinguish COVID-19 from other respiratory infections (bacterial pneumonia, influenza, seasonal coronavirus) and healthy controls, improving disease-specific diagnosis.

Core claims
  • A pipeline combining Boruta and mRMR feature selection with incremental feature selection (IFS) was used to identify COVID-19-specific transcriptomic biomarkers from blood gene expression data. method
  • An optimum SVM classifier built on the top 168 mRMR-ranked genes achieved the best classification performance among tested algorithms for distinguishing five respiratory infection categories. finding
  • KNN, RF, and DT classifiers were also built via IFS but achieved lower classification performance than the optimum SVM classifier. finding
  • A decision tree model yielded 21 interpretable classification rules, the majority (8) of which predicted SARS-CoV-2 infection. finding
  • Identified biomarker genes (RPL6, ZNF496, DYNLRB1, TRBV20-1, PHOSPHO1, TMEM165, RPL36AL) show pathogen/disease-specific expression patterns consistent with prior literature. finding
  • GO and KEGG functional enrichment analysis of the top 168 SVM-selected genes revealed significant enriched biological terms and pathways. finding
  • The GSE161731 blood transcriptomic dataset (195 subjects across 5 groups) served as the resource enabling disease-specific host biomarker discovery. resource
Experimental setups
Assay System Perturbation Readout Platform
bulk blood transcriptomics (RNA expression profiling, TPM) human blood samples (195 subjects: healthy controls, bacterial pneumonia, influenza, seasonal coronavirus, SARS-CoV-2/COVID-19) none (natural infection/disease state comparison) expression levels of 15,379 genes
GO and KEGG functional enrichment analysis human blood samples, top 168 SVM-selected genes none significantly enriched GO terms and KEGG pathways (FDR<0.05) DAVID website
Key results
  • Optimum SVM classifier using top 168 mRMR-ranked features achieved the highest MCC and overall accuracy among all tested classifiers. MCC=0.917; accuracy=0.938
  • KNN classifier using top 183 features achieved lower performance than SVM. MCC=0.845
  • RF classifier using top 565 features achieved lower performance than SVM. MCC=0.896
  • DT classifier using top 511 features achieved the lowest performance among the four algorithms but provided interpretable rules. MCC=0.818; accuracy=0.867
  • 21 classification rules were extracted from the DT model, with 8 rules predicting SARS-CoV-2 infection, more than for any other category. 8 of 21 rules
  • Boruta feature selection reduced 15,379 genes to 604 relevant features prior to mRMR ranking. 604/15,379 genes
  • Top-ranked genes including RPL6, ZNF496, DYNLRB1, TRBV20-1, PHOSPHO1, and TMEM165 were linked in prior studies to pathogen-specific (bacterial, influenza, or coronavirus) infection responses.
Key statistics
  • other MCC=0.917 (Optimum SVM classifier, top 168 features)
  • other accuracy=0.938 (Optimum SVM classifier overall accuracy)
  • other MCC=0.896 (Optimum RF classifier, top 565 features)
  • other MCC=0.845 (Optimum KNN classifier, top 183 features)
  • other MCC=0.818; accuracy=0.867 (Optimum DT classifier, top 511 features)
  • count 604 features (Genes retained by Boruta feature filtering out of 15,379 total)
  • count 21 rules (8 for SARS-CoV-2) (Classification rules extracted from DT model)
  • pvalue FDR<0.05 (Significance threshold for GO/KEGG enrichment analysis)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study applies a sequential machine-learning pipeline to blood transcriptomic profiles of 195 subjects spanning five disease categories to identify COVID-19-specific biomarkers. Genes were first filtered by Boruta, ranked by mRMR, and then fed into incremental feature selection (IFS) paired with four classifiers (RF, SVM, kNN, DT), each evaluated by 10-fold cross-validation using Matthew's Correlation Coefficient (MCC) as the primary performance metric. The best model (SVM, 168 features, MCC=0.917, overall accuracy=0.938) was selected, and functional enrichment of the chosen genes was assessed via GO/KEGG analysis in DAVID at FDR<0.05.

Replicationbiological Sample size195 total subjects described by category (19 healthy controls, 23 bacterial pneumonia, 17 influenza, 59 seasonal coronavirus, 77 COVID-19); no formal power analysis or sample-size justification mentioned Groups5-class: healthy controls vs. bacterial pneumonia vs. influenza vs. seasonal coronavirus vs. COVID-19 (SARS-CoV-2) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionFDR (applied via DAVID, threshold set at 0.05)
Statistical tests used
Test Applied to n Assumptions
Boruta (random-forest-based feature selection) Initial dimensionality reduction from 15,379 genes to 604 relevant features 195 not stated
mRMR (mutual-information-based feature ranking) Ranking of 604 Boruta-selected features into an ordered feature list 195 not stated
Incremental feature selection (IFS) with 10-fold cross-validation; MCC as selection criterion Identifying the optimal feature-subset size for each of four classifiers (Figure 2) 195 not stated
SVM multiclass classifier (SMO algorithm in Weka) Primary optimum classifier selected on 168 features (Table 1, Figure 3) 195 not stated
RF, kNN (Ibk in Weka), and CART decision tree (Gini index, scikit-learn) multiclass classifiers Compared with SVM in IFS evaluation (Table 1, Figures 2–3); DT rules extracted (Table S4) 195 not stated
GO and KEGG over-representation analysis (DAVID, FDR threshold 0.05) Functional annotation of 168 SVM-selected genes against 15,379-gene background (Table 2) 168 genes / 15,379-gene background not stated
Approaches that could also have been used
  • Classifier performance was estimated solely by a single run of 10-fold cross-validation on the full 195-sample dataset, with feature selection applied inside the same data pool
    Could also: A nested cross-validation scheme (outer loop for unbiased performance estimation, inner loop for feature selection) or a fully held-out independent test set could also be used — Nested CV or an independent test partition reduces the optimistic bias that can arise when feature selection and performance evaluation share the same data splits, yielding a less-biased estimate of generalization error—particularly relevant when n is modest relative to the number of candidate features
  • MCC served as the sole model-selection and reporting metric across all four classifiers
    Could also: Macro-averaged AUROC, per-class F1, or macro-averaged F1 are also widely reported for multiclass classification tasks — Complementary metrics capture different aspects of performance (e.g., discrimination vs. per-class sensitivity/specificity); with group sizes ranging from n=17 to n=77, metrics that weight classes differently can surface performance disparities that a single summary statistic may not fully convey
  • No confidence intervals or variance estimates were reported around any cross-validation performance metric
    Could also: Bootstrap confidence intervals or repeated k-fold CV (e.g., 10×10-fold) with variance reporting could also be applied — With 195 samples and some classes as small as n=17, a single 10-fold CV run carries non-negligible variance; reporting CI on MCC or accuracy would convey the uncertainty in, for example, the stated MCC of 0.917 for SVM
  • The four classifiers were compared by visually inspecting MCC curves (Figure 2) and tabulated accuracy (Table 1), without a formal statistical test of differences
    Could also: A McNemar test or a Friedman test with post-hoc Nemenyi correction could also formally compare classifier error rates across CV folds — Formal comparison methods account for the shared data structure of cross-validation folds and provide a principled basis for concluding whether one classifier's MCC differs from another beyond sampling variability
  • Boruta and mRMR were applied sequentially as a fixed pipeline without comparison to alternative dimensionality-reduction strategies
    Could also: Elastic-net penalized logistic regression, LASSO, or recursive feature elimination (RFE) with cross-validation could also identify a compact informative gene set — Different feature selection methods make distinct assumptions about feature independence and linearity; comparing multiple methods can reveal whether the selected gene set is robust across approaches or sensitive to a particular method's assumptions
  • Functional enrichment was assessed using over-representation analysis (ORA) in DAVID with a binary FDR<0.05 filter on the 168 SVM-selected genes
    Could also: Gene Set Enrichment Analysis (GSEA) using a ranked gene list (e.g., ranked by mRMR score or SVM feature weight) could also be applied — GSEA evaluates the full ranked gene distribution rather than a binary inclusion cutoff, which can capture gradient-level pathway signal; it also handles gene-set redundancy and correlation differently from DAVID's hypergeometric-based ORA
Software: boruta_py (scikit-learn-contrib, Python) · mRMR (Peng Lab standalone program) · Weka (SVM via SMO; kNN via Ibk) · scikit-learn (CART decision tree, Gini index) · DAVID (GO/KEGG enrichment)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

The full computational pipeline: GSE161731 normalized log-CPM matrix (195×15,379, 5 cohorts) → Boruta feature selection → mRMR ranking → incremental feature selection with RF/SVM/kNN/DT (10-fold CV, max-MCC optimum). Pinned claims:

  • input dims: 195 samples, 15,379 genes, cohorts 19/23/17/59/77
  • Boruta: 604 relevant features
  • Table 1 optima: SVM 168 (acc 0.938, MCC 0.917); RF 565 (0.923, 0.896); kNN 183 (0.882, 0.845); DT 511 (0.867, 0.818)

NOTE: Article RETRACTED Nov 2023. Data accession per paper text is GSE161731 (registry's GSE150728 is only cited as prior work). See reproduction/scope.md.

Figures / tables: Table
input_dims
Reported
195 samples; 15,379 genes; cohorts 19 healthy / 23 bacterial / 17 influenza / 59 seasonal CoV / 77 SARS-CoV-2
Reproduced
195 samples x 15,379 genes; cohorts healthy 19 / Bacterial 23 / Influenza 17 / CoV-other 59 / COVID-19 77
exact
boruta_features
Reported
604 relevant features (Boruta)
Reproduced
663 confirmed (+110 tentative)
partial
svm_optimal
Reported
SVM 168 feat, acc 0.938, MCC 0.917
Reproduced
best 77 feat acc 0.831 MCC 0.764; at n=168 acc 0.795 MCC 0.713
did not match
rf_optimal
Reported
RF 565 feat, acc 0.923, MCC 0.896
Reproduced
best 644 feat acc 0.887 MCC 0.843; at n=565 acc 0.856 MCC 0.800
did not match
knn_optimal
Reported
kNN 183 feat, acc 0.882, MCC 0.845
Reproduced
best 69 feat acc 0.856 MCC 0.800; at n=183 acc 0.810 MCC 0.734
did not match
dt_optimal
Reported
DT 511 feat, acc 0.867, MCC 0.818
Reproduced
best 460 feat acc 0.810 MCC 0.738; at n=511 acc 0.769 MCC 0.684
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 32/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

Preliminary/partial reproduction. The input data is reproduced exactly (195×15,379 genes, cohorts 19/23/17/59/77 from GSE161731), confirming data identity 1:1, but the core ML claims (604 Boruta features; SVM acc 0.938/MCC 0.917 and the RF/kNN/DT metrics) are still PENDING on «our HPC» «job», so the central conclusion can be neither confirmed nor refuted. The only metadata wrinkle is on our side — the registry mislabeled the accession as GSE150728 when the paper uses GSE161731 — not an authors' defect. Note the paper was retracted in 2021, which heightens scrutiny, but with the pipeline fully specified by named third-party tools and no measured deviation yet, the fair grade is yellow pending completion.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

120.9 k
tokens (I/O) · 6.1 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.