Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Bioinformatics and system biology approaches to identify pathophysiological impact of COVID-19 to the progression and severity of neurological diseases.

Comput Biol Med · 2021
L1 69/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
69/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 35% of all assessed papers rank 745 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the FOUNDATION 1:1, not the full multi-tool pipeline. The brief's repo URL (COVID-93_NDs) 404s; the real repo is github.com/HabibUCAS/COVID-19_NDs (commit 1f584ca), which ships only a GEO2R/limma DEG auto-script (Part1_GSE28146.R) + the authors' Common Dysregulated Genes.xlsx + figure PDFs. FRESH independent recompute on «our HPC» (SLURM «job» core + 2218110 sweep; R 4.3.3 / limma 3.58.1): the paper's per-dataset DEG counts at p<0.05 (Table 1) reproduce EXACTLY for 11 GEO datasets spanning ALL SIX neurological diseases (AD: GSE28146/GSE1297/GSE12685; ALS: GSE4595/GSE68605; ED: GSE19332; HD: GSE77558; MS: GSE19587; PD: GSE20141/GSE28894/GSE42966), +1 within 0.35% (GSE7621). For GSE1297/GSE12685/and the 8 sweep sets we derived the case/control grouping ourselves yet still hit the exact counts — strong evidence these Table-1 numbers are genuine limma outputs, not fabricated. 'p<0.05' = raw probe-level P.Value<0.05 (FDR<0.05 -> ~0). Three p<0.05 counts mismatch (GSE32915/38010/52139) where the case/control split is undocumented and our auto-grouping differs — under-specification, not contradiction. The companion p<0.05 & |logFC|>=1 counts reproduce only approximately (GSE28146 1107 reported vs 1950 ours; GSE1297 238 vs 313; GSE12685 149 vs 165): the fold-change-filter step has an undocumented detail — flagged. NOT independently reproducible: the COVID-19 dataset has NO GEO accession in the paper, so the headline COVID<->ND common-gene counts cannot be regenerated; verified only that they match the authors' shipped xlsx exactly (AD52/ALS76/ED8/HD73/MS60/PD91, sum 360). NOT attempted (out of scope, external web GUIs, no shipped params): GO/KEGG enrichment, STRING/cytoHubba PPI hub genes, NetworkAnalyst TF/miRNA networks, DSigDB drug prediction, all figures.

💻 Code ↗ 🗄 Data: GSE1297

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 69
    assessed: 2026-06-16 ⛓ 9044a520b6ff
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

COVID-19 interacts with and impacts the progression and severity of neurological diseases (Alzheimer's, ALS, Epilepsy, Huntington's, Multiple sclerosis, Parkinson's), and this connection can be elucidated via a bioinformatics/network-based pipeline analyzing shared transcriptomic signatures.

Core claims
  • COVID-19 and neurological diseases (NDs) share pathophysiological connections that influence disease progression and severity finding
  • A bioinformatics and network-based pipeline (R-based) was developed to identify molecular connections between COVID-19 and NDs using transcriptomic data method
  • Hub proteins identified via PPI network analysis point to candidate therapeutic targets/strategies for COVID-19-ND comorbidity finding
  • Gene-based semantic similarity between COVID-19 and NDs is maximum for Parkinson's disease finding
  • Gene-based semantic similarity between COVID-19 and NDs is minimum for Multiple sclerosis finding
  • Gene ontology-based semantic similarity between COVID-19 and NDs is maximum for Huntington's disease finding
  • Gene ontology-based semantic similarity between COVID-19 and NDs is minimum for Epilepsy disease finding
  • Findings were validated against gold-standard databases (dbGaP, OMIM, OMIM Expanded) and literature method
Experimental setups
Assay System Perturbation Readout Platform
Microarray/RNA-seq differential gene expression analysis Peripheral blood mononuclear cells (COVID-19 patients) SARS-CoV-2 infection Differentially expressed genes (DEGs)
Microarray DEG analysis Hippocampal CA1 tissue / Frontal cortex (Alzheimer's disease) disease (AD) vs control DEGs
Microarray DEG analysis Spinal cord gray matter / Motor cortex / Cervical spinal cord / motor neurons (Amyotrophic lateral sclerosis) disease (ALS) vs control DEGs
Microarray DEG analysis Peritumoral neocortex / post-mortem brain (Epilepsy disease) disease (ED) vs control DEGs
Microarray DEG analysis iPSC-derived GABA MS-like neurons / motor cortex (Huntington's disease) disease (HD) vs control DEGs
Microarray DEG analysis White matter brain tissue / brain lesion / spinal cord periplaque regions (Multiple sclerosis) disease (MS) vs control DEGs
Microarray DEG analysis Substantia nigra / laser-dissected SNpc neurons / post-mortem brain (Parkinson's disease) disease (PD) vs control DEGs
Protein-protein interaction network and hub protein/topological analysis Common DEG-encoded proteins across COVID-19 and ND datasets none Hub proteins (degree matrices) STRING database via Network Analyst; Cytoscape
Key results
  • Maximum gene-based semantic similarity score found between COVID-19 and Parkinson's disease
  • Minimum gene-based semantic similarity score found between COVID-19 and Multiple sclerosis
  • Maximum gene ontology-based semantic similarity score found between COVID-19 and Huntington disease
  • Minimum gene ontology-based semantic similarity score found between COVID-19 and Epilepsy disease
  • COVID-19 PBMC dataset yielded DEGs identified at p=0.05 and further filtered by |logFC|>=1
  • Hub proteins identified from PPI network proposed as potential therapeutic targets
Key statistics
  • count 4453 DEGs (Pval=0.05) (COVID-19 PBMC dataset (3 case, 3 control))
  • count 2657 DEGs (Pval=0.05, |logFC|=1) (COVID-19 PBMC dataset stricter filter)
  • count 2313 DEGs (Pval=0.05) (GSE1297, AD hippocampal CA1 tissue (22 case, 9 control))
  • count 7094 DEGs (Pval=0.05) (GSE20141, PD laser-dissected SNpc neurons (10 case, 8 control))
  • count 55 case, 59 control samples (GSE28894, Parkinson's disease brain dataset)
  • other overall score > 0.5 (STRING database confidence threshold for PPI network construction)
  • count 163 DEGs (Pval=0.05) (GSE52672, ALS spinal cord homogenate (10 case, 10 control))
  • count worldwide 211,855,573 confirmed COVID-19 cases, 4,433,151 deaths (as of 21 Aug 2021) (WHO global COVID-19 statistics cited in introduction)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational bioinformatics and network-biology study that analyzed publicly available microarray/RNA-seq datasets comparing diseased tissues to controls for COVID-19 and six neurological diseases. Differentially expressed genes were identified with the limma package (using a p-value threshold of 0.05, with an additional |logFC| = 1 filter), gene-set/pathway enrichment was assessed (including a Fisher's-test-based GSEA), and downstream analyses used protein-protein interaction networks, transcription factor/miRNA interactions, and Gene Ontology semantic similarity. Results were reported primarily as counts of DEGs, enriched terms/pathways, hub proteins, and semantic-similarity scores rather than as group-level effect estimates with dispersion.

Replicationbiological Sample sizeSample sizes are reported only as per-dataset case and control sample counts in Table 1; no formal power/sample-size calculation is described, and datasets with too few samples or missing case/control groups were excluded 'due to lack of statistical significance' Groupsdiseased tissue vs. control, per dataset; cross-disease comparison of COVID-19 with six NDs Pairingunpaired Randomization/blindingna Dispersionnone Exact p-valuesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
limma differential expression analysis (linear models for microarray data) Identification of differentially expressed genes (DEGs) for each disease dataset vs. control (Table 1) per-dataset case vs control sample counts as listed in Table 1 (e.g., COVID-19 3 case/3 control; ranges up to 55/59) not stated
Gene Set Enrichment (GSE) test, including a Fisher's-exact-based GSEA ('Fisher GSEA' column) Enrichment of up-/down-regulated DEG sets (Table 1: Raw GSEA and Fisher GSEA columns) not stated
Pathway enrichment analysis (Enrichr against KEGG, WikiPathways, BioCarta, Reactome) Signaling pathways enriched by DEGs not stated
Gene Ontology enrichment (topGO) and GO/gene semantic similarity (GOSemSim, best-matching-average) Significant GO terms and proximity (semantic similarity scores) between COVID-19 and each ND na
Threshold-based significance cutoff (p-value = 0.05; additional |logFC| = 1 filter) DEG selection across all datasets (Table 1) not stated
Approaches that could also have been used
  • DEGs were called using a raw p-value threshold of 0.05 (with an optional |logFC| = 1 filter).
    Could also: An adjusted-significance approach such as Benjamini-Hochberg FDR or Bonferroni control on the limma p-values could also be applied. — Multiplicity adjustment across the genome-wide set of tests would also control the expected proportion of false positives and is commonly used in transcriptomic DEG selection.
  • RNA-seq and microarray datasets were both analyzed with limma after Z-score transformation.
    Could also: Count-based RNA-seq datasets could also be modeled with negative-binomial frameworks such as DESeq2 or edgeR (or limma-voom). — Count-aware models are tailored to RNA-seq mean-variance structure and would also provide an alternative way to estimate dispersion for sequencing data.
  • Sample sizes are described only as per-dataset case/control counts, with small-sample datasets excluded for 'lack of statistical significance.'
    Could also: A brief power or sensitivity consideration, or explicit reporting of per-comparison n alongside results, could also accompany the analysis. — Documenting the basis for sample-size adequacy would also help readers gauge the resolution of each comparison, especially where group sizes are small (e.g., 3 vs 3).
  • Findings are summarized as counts of DEGs, pathways, and hub proteins without dispersion or interval estimates.
    Could also: Reporting effect sizes (e.g., log fold-changes with confidence intervals) or adjusted p-values for key genes could also be presented. — Interval/effect-size reporting would also convey the magnitude and uncertainty of individual changes in addition to the binary in/out DEG classification.
  • Enrichment used a Fisher's-exact-style over-representation test on DEG lists.
    Could also: A rank-based enrichment method such as classic GSEA (using the full ranked gene list) could also be used. — Rank-based enrichment would also incorporate genes below the DEG threshold and reduce sensitivity to the chosen cutoff.
  • Cross-disease relatedness was quantified with GO/gene semantic similarity using the best-matching-average aggregation.
    Could also: Alternative aggregation strategies (e.g., maximum, average, or rcmax) or alternative similarity measures (e.g., Resnik, Lin, Wang) could also be reported. — Showing results under more than one aggregation/similarity definition would also indicate how robust the proximity rankings are to methodological choice.
Software: R (Bioconductor) · limma · GEOquery · genefilter · topGO · GOSemSim · Enrichr · NetworkAnalyst (with STRING) · Cytoscape

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
38
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34601390

Paper: Rahman et al. 2021, Bioinformatics and system biology approaches to identify pathophysiological impact of COVID-19 to the progression and severity of neurological diseases. Comput Biol Med 138:104859. PMCID PMC8483812.

Code: https://github.com/HabibUCAS/COVID-19_NDs (brief listed COVID-93_NDs — that URL 404s; the real repo for this paper, by the same author HabibUCAS, is COVID-19_NDs, last pushed 2022-02-14, branch master, no license file.)

The repo ships:

  • Part1_GSE28146.R — a GEO2R-style limma DEG script for GSE28146 (Alzheimer, platform GPL570, group string gsms="111111110000000000000000000000" = 8 vs 22).
  • Heatmap.R — heatmap.2 on a Heatmap.csv (csv NOT shipped).
  • Common Dysregulated Genes.xlsx — the authors' output: per-disease common-gene lists (AD/ALS/ED/HD/MS/PD) + an All_Together union sheet.
  • 15 figure PDFs (final figures, not regenerable inputs).
  • README.md = "T2D" (placeholder; no run instructions).

Pipeline-derived results (IN SCOPE)

# Result Pipeline Reproducible from shipped artifacts?
R1 GSE28146 DEG counts: 3405 (p<0.05), 1107 (p<0.05 & |logFC|≥1) — Table 1 getGEOlimma (the shipped Part1_GSE28146.R) YES — shipped code + public GEO. Primary target.
R2 Per-dataset DEG counts for the other 25 datasets in Table 1 same limma method (P<0.05 & |logFC|≥1) applied to each accession PARTIAL — method described, but per-dataset case/control grouping NOT shipped (only GSE28146's is). Best-effort, group assignment is our judgment.
R3 COVID-19↔ND common-gene counts: AD 52, ALS 76, ED 8, HD 73, MS 60, PD 91; total 360 / 303 unique intersect COVID DEGs with each ND's DEGs BLOCKED for independent repro — the COVID-19 dataset has no GEO accession in the paper (Table 1 row "COVID-19", PBMC, 3 case/3 control, 2657 DEGs). Cannot regenerate the COVID DEG set. We CAN verify the counts against the authors' shipped xlsx (internal-consistency check, not independent).
R4 Hub genes per ND (e.g. AD: COPB1,AP2S1,COPE,CYBB,JAK2,GATA3,COPA,COX5A,SIRPA,ANK1,HGF) Cytoscape/STRING PPI + cytoHubba OUT for now — no PPI script shipped; downstream of R3.
R5 GO/KEGG enrichment, TF/miRNA networks, drug molecules EnrichR / NetworkAnalyst / external web tools OUT — external web tools, no script/params pinned.

OUT OF SCOPE (not pipeline-reproducible here)

  • All web-tool steps (DAVID/EnrichR/STRING/NetworkAnalyst/DSigDB) — no shipped params, manual web GUI.
  • Figures (shipped as final PDFs).
  • The COVID-19 DEG set itself (no accession → R3/R4 cannot be independently rebuilt).

Primary plan

  1. R1 — re-run the shipped GSE28146 limma pipeline on «our HPC», compare DEG counts to Table 1 (3405 / 1107). Clean, self-contained 1:1.
  2. R2 — extend to the other AD datasets (GSE1297=238, GSE12685=149) and a few more with best-effort GEO2R groupings, flagged as method-reproduction.
  3. R3 — verify the shipped xlsx common-gene counts match the reported text (consistency check); document the COVID-accession gap as the repro blocker.
Figures / tables: TableFig.2
C1
Reported
GSE28146 (AD) DEGs p<0.05 = 3405
Reproduced
3405
exact
C10
Reported
GSE1297 (AD) DEGs p<0.05 = 2313
Reproduced
2313
exact
C12
Reported
GSE12685 (AD) DEGs p<0.05 = 3251
Reproduced
3251
exact
C14-C21
Reported
8 Table-1 datasets DEGs p<0.05 (GSE4595/19332/68605/77558/19587/20141/28894/42966)
Reproduced
all exact
exact
C22
Reported
GSE7621 (PD) DEGs p<0.05 = 5949
Reproduced
5928
within tolerance
C2
Reported
GSE28146 DEGs p<0.05 & |logFC|>=1 = 1107
Reproduced
1950 probe / 1562 gene
did not match
C11
Reported
GSE1297 DEGs p<0.05 & |logFC|>=1 = 238
Reproduced
313 probe / 290 gene
partial
C13
Reported
GSE12685 DEGs p<0.05 & |logFC|>=1 = 149
Reproduced
165 probe / 171 gene
within tolerance
C23-C25
Reported
GSE32915/38010/52139 DEGs p<0.05 (517/1996/2703)
Reproduced
differ (1126/8899/3945)
did not match
C3-C9
Reported
COVID<->ND common genes AD52/ALS76/ED8/HD73/MS60/PD91, total 360
Reproduced
matches authors' shipped xlsx exactly (consistency-only)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 69/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

The paper's Table-1 p<0.05 DEG counts reproduce EXACTLY for 11 GEO datasets (incl. GSE1297/GSE12685 where we independently derived the grouping), strong evidence those numbers are genuine limma outputs, not fabricated. The deviations are bounded and explainable: the undocumented |logFC|>=1 filter (GSE28146 1107 vs 1478-1950) and self-chosen case/control grouping on 3 datasets — both underspecification, on the methodology/authors-omission boundary, not contradiction. The central blocker is data availability: the COVID-19 dataset has no GEO accession, so the headline COVID<->ND common-gene counts (total 360) are only internally consistent with the authors' shipped xlsx, not independently reproducible, and all downstream enrichment/PPI/drug claims were out of scope. Overall solid foundation with explainable, non-critical deviations → yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

402.9 k
tokens (I/O) · 28.8 M incl. cache
106 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.