Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

TSUNAMI: Translational Bioinformatics Tool Suite for Network Analysis and Mining.

Genomics Proteomics Bioinformatics · 2021
L1 79/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
How its reproducibility compares
79/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 55% of all assessed papers rank 514 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce without author contact; essentially 1:1. TSUNAMI is a Shiny wrapper around the authors' lmQCM algorithm; we ran lmQCM 0.2.4 (CRAN, same algorithm) on the paper's own data GSE17537 with the verbatim preprocessing + parameters from the repo. C1 (dataset 54675x55, GPL570), C2 (15-gene module) and C3 (the 9 Table-1 interferon genes, 9/9) reproduced EXACTLY; C4 top GO term (Type I interferon signaling GO:0060337) and its 9 overlap genes reproduced exactly, with the exact p-value differing only because the paper used the now-unavailable Enrichr GO_BP_2018 library (within-tol). Only the Enrichr z-score (C4z, a known version-dependent soft value) does not match numerically. NOT attempted: survival KM curves (Fig 6, qualitative/no pinnable number) and the web UI/Circos rendering (interface, not numeric results). No fabrication concerns: all reported values are derivable from the shipped algorithm + public data.

💻 Code ↗ 🗄 Data: GSE17537

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 79
    assessed: 2026-06-14 ⛓ 30871629b604
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

There is no existing tool that provides a complete, easy-to-use pipeline for mining relatively small, tightly connected (and potentially overlapping) gene co-expression network modules from transcriptomic data and directly linking them to downstream gene set enrichment, visualization, and survival analysis; TSUNAMI was built to fill this gap.

Core claims
  • TSUNAMI is a freely accessible web-based tool suite that mines gene co-expression network (GCN) modules from public (GEO, TCGA) or user-uploaded numerical omics data and performs downstream gene set enrichment analysis. resource
  • The lmQCM algorithm mines smaller, more densely connected GCN modules than WGCNA and allows overlapping module membership, better reflecting biological pathway co-regulation. method
  • TSUNAMI provides direct search/retrieval interfaces to NCBI GEO and TCGA databases as well as user file upload (CSV, TSV, XLSX, TXT). method
  • TSUNAMI integrates Enrichr for downstream gene set enrichment analysis across 14 default categories. method
  • TSUNAMI generates Circos plots (via R package circlize) to visualize chromosomal locations and pairwise relationships of genes within a GCN module. method
  • TSUNAMI includes a survival analysis module that correlates GCN module eigengenes (dichotomized at the median) with patient overall/event-free survival using the log-rank test. method
  • In the GSE17537 example dataset, the 36th lmQCM-derived GCN module (15 genes) is significantly enriched for the type I interferon signaling pathway. finding
  • Smaller gene modules derived from lmQCM tend to generate more biologically meaningful gene set enrichment results than larger WGCNA-style modules. finding
Experimental setups
Assay System Perturbation Readout Platform
Gene co-expression network mining (lmQCM algorithm) GSE17537 microarray dataset, 55 colorectal cancer patients (Vanderbilt Medical Center) none Merged GCN modules (gene clusters) sorted by size; eigengene matrix Affymetrix HU133 2.0 Plus GeneChip (GPL570)
Gene co-expression network mining (WGCNA) GSE17537 microarray dataset none Hierarchical clustering of gene modules R package WGCNA (Bioconductor)
Gene set enrichment analysis 36th GCN module (15 genes) from GSE17537, mined by lmQCM none Enriched terms, P value, Z-score, overlapping genes Enrichr
Circos plot gene locus visualization 36th GCN module (15 genes), human genome hg38/hg19 none Chromosomal positions and pairwise gene links R package circlize; UCSC refGene database
Survival analysis (log-rank test, Kaplan-Meier) GSE17537, 55 samples with overall survival data none P value of log-rank test comparing high vs. low eigengene groups (median split) for OS/EFS R package survival (survdiff function)
GEO dataset retrieval/processing test First 1000 GSE datasets from NCBI GEO none Number of datasets successfully processed vs. failed
Key results
  • The 36th GCN module (15 genes) is highly enriched in the GO Biological Process term 'type I interferon signaling pathway (GO:0060337)' 9/148 genes overlap, P=2.51E-16
  • Only a small portion of tested GSE datasets failed to process with TSUNAMI, mostly legacy microarray data with excessive missing data or small sample size 12 out of first 1000 GSE datasets
  • Kaplan-Meier survival analysis was generated for the 36th GCN module eigengene dichotomized at the median into high/low groups
Key statistics
  • pvalue 2.51E-16 (Enrichr enrichment of 36th GCN module for type I interferon signaling pathway (GO:0060337), 9/148 gene overlap)
  • other Z-score = -3.2821 (Enrichment result for type I interferon signaling pathway in 36th GCN module)
  • pvalue 1.80E-09 (Enrichment for 'cellular response to type I interferon (GO:0071357)', 4/23 overlap)
  • count 12 out of 1000 (GSE datasets from GEO that could not be processed by TSUNAMI)
  • count 55 (Colorectal cancer patients in example dataset GSE17537 used for survival analysis)
  • count 54,675 probesets (Number of probesets on the Affymetrix HU133 2.0 Plus GeneChip platform used for GSE17537)
  • other γ=0.7, λ=1, t=1, β=0.4, minimum cluster size=10 (Default lmQCM parameters used for GSE17537 example analysis)
  • other power=10, reassign threshold=0/1, merge cut height=0.25, minimum module size=10 (WGCNA parameters used for GSE17537 example analysis)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper describes TSUNAMI, a bioinformatics web tool suite for gene co-expression network (GCN) mining and downstream analysis; statistical methods are demonstrated on a single illustrative dataset (GSE17537, n=55 colorectal cancer patients) rather than testing a primary biological hypothesis. GCN modules are constructed using Pearson or Spearman correlation-based algorithms (lmQCM or WGCNA). Downstream analyses include gene set enrichment via Enrichr (reporting P values and Z-scores across 14 databases) and Kaplan-Meier survival analysis comparing median-split eigengene groups via the log-rank test.

Replicationunclear Sample size55 colorectal cancer patients from GSE17537 (Vanderbilt Medical Center, Affymetrix HU133 2.0 Plus) used as a single illustrative demonstration; no formal power analysis or sample size justification stated GroupsHigh vs. low GCN module eigengene expression defined by median split, compared on overall survival Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Log-rank test (R survdiff) Overall survival comparison between high vs. low GCN module eigengene groups (Figure 6, 36th module from GSE17537) 55 not stated
Enrichr enrichment analysis (Fisher's exact test, internal to Enrichr) Gene set enrichment of GCN module gene lists across 14 databases (Table 1 and associated panels) Gene overlap counts stated per term (e.g., 9/148 for top GO term) not stated
Pearson correlation coefficient (PCC) Pairwise gene correlation for GCN module construction via lmQCM (default setting, GSE17537 example) 55 not stated
Spearman rank correlation coefficient (SCC) Alternative pairwise gene correlation for GCN module construction (recommended for RNA-seq data per tool documentation) not stated
Singular value decomposition (first principal component as eigengene) Summarization of gene expression within each GCN module into a single eigengene value (Figure 5B) 55 na
Approaches that could also have been used
  • Survival analysis groups patients into high and low eigengene expression by median split before applying the log-rank test
    Could also: Cox proportional hazards regression with eigengene as a continuous predictor could also be used — Retaining eigengene as a continuous variable avoids information loss from dichotomization and yields a hazard ratio with confidence interval, which quantifies effect magnitude; median split can produce different results depending on the specific cut point chosen
  • The log-rank test is applied to compare survival between the two eigengene groups without covariate adjustment
    Could also: Multivariable Cox regression adjusting for available clinical covariates (e.g., stage, age) could also be applied — Covariate-adjusted models can separate the independent prognostic contribution of a GCN module from confounders, which is relevant when demonstrating clinical utility of a biomarker module
  • Fourteen enrichment analyses are performed concurrently across databases without an explicitly stated multiple-testing correction
    Could also: Applying Benjamini-Hochberg FDR correction within each database (already implemented internally by Enrichr as adjusted P values) and noting this correction explicitly would also be standard practice — Reporting the adjusted P values available from Enrichr alongside raw P values would make the degree of correction transparent to readers and facilitate comparison across studies
  • Pearson correlation is used as the default measure for GCN construction on the microarray demonstration dataset
    Could also: Spearman rank correlation (also available in TSUNAMI) or mutual information-based measures could also be used — Spearman correlation is more robust to outliers and non-normality; mutual information captures non-linear dependencies; the paper itself recommends Spearman for RNA-seq, so the choice of measure can be matched to the distributional properties of the input data
  • Tool performance and enrichment results are demonstrated on a single example dataset (GSE17537, n=55)
    Could also: Benchmarking on multiple datasets with known biological ground truth (e.g., simulated data or datasets with validated pathway activity) could also be used — Multi-dataset demonstration would allow readers to assess the consistency and generalizability of the tool's outputs across different platforms, sample sizes, and disease contexts
  • GCN module eigengenes are derived as the first principal component of within-module expression via SVD without variance explained being reported
    Could also: Reporting the proportion of variance explained by the first PC for each module could also accompany the eigengene output — The variance-explained fraction conveys how well a single eigengene summarizes the module; a low value would indicate heterogeneous expression within the module, which is useful context for interpreting downstream survival or enrichment results
Software: R/Shiny · lmQCM (CRAN) · WGCNA (Bioconductor) · R/survival (survdiff) · Enrichr · R/circlize · BiocGenerics (R/Bioconductor)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
16
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (1)

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GO:0060339 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
also used by 1 paper:
GO:0034340 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0044828 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0045071 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0045869 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0060337 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0060338 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0060340 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0071357 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:1902046 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE17537 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33705981 (TSUNAMI)

Paper: Huang Z. et al. TSUNAMI: Translational Bioinformatics Tool Suite for Network Analysis and Mining. Genomics Proteomics Bioinformatics 2021. PMID 33705981 · PMCID PMC9403021 · DOI 10.1016/j.gpb.2019.05.006 Code: https://github.com/huangzhii/TSUNAMI (R Shiny app; default branch master) Data: GEO GSE17537 (colorectal cancer, Affymetrix HG-U133 Plus 2.0 / GPL570)

Nature of the paper

TSUNAMI is a web tool (R Shiny) wrapping the authors' own lmQCM co-expression module-mining algorithm (also on CRAN as package lmQCM, same authors), plus Enrichr enrichment, Circos visualisation and survival analysis. There is no wet-lab content. The only concrete, reported numeric results are in the case study on GSE17537 (Section "An example", Figures 5–6, Table 1). Per P16 of the brief, applying the authors' own published tool to the paper's own data is a fully valid reproduction.

In scope (pipeline-derived, attempted)

The case-study pipeline is fully specified in the paper + source:

  1. Data ingestgetGEO("GSE17537", GSEMatrix=TRUE, AnnotGPL=FALSE); expression matrix exprs(); probe→gene-symbol map from GPL570 platform table column "Gene Symbol" (server.R:240,448-472).
  2. Preprocessing (server.R:575-655, utils.R:varFilter2):
    • mean filter: remove genes below the 50th percentile of row means (J=50);
    • variance filter: remove genes below the 10th percentile of row variance (K=10);
    • drop empty symbols; collapse duplicate gene symbols keeping the max-mean probe.
  3. lmQCM module mining (server.R:710-...): Pearson correlation, γ=0.7, t=1, λ=1, β=0.4, minClusterSize=10, normalization=FALSE; modules sorted by size descending.
  4. GO enrichment of a module via Enrichr (GO Biological Process).

Claims to reproduce (see original/claims.tsv)

  • C1 GSE17537 = 55 samples × 54,675 probesets on GPL570 (paper text).
  • C2 A co-expression module of 15 genes (the "36th module", Fig 5C).
  • C3 That module contains the 9 genes listed in Table 1: SP100, RSAD2, STAT2, MX1, ISG15, SAMHD1, XAF1, IFIT1, IFIT3.
  • C4 GO enrichment of that module → top term Type I interferon signaling pathway (GO:0060337), overlap 9/148, p = 2.51E-16 (Table 1).

Out of scope / not attempted (with reason)

  • Survival analysis (Fig 6) — KM curves are qualitative (no reported number to pin); log-rank on a module eigengene is straightforward but yields no specific printed value to compare → no_expected_result for that panel. May add as a bonus if core claims reproduce.
  • The web UI / Circos rendering / GEO browser — interface features, not numeric results.
  • The Z-score (-3.2821) in Table 1 is Enrichr's combined-score component; it is sign/version-dependent, recorded but treated as soft.

Heavy-compute note

Compute is modest (55 samples; ~13k-gene Pearson matrix) but per the hard rule it runs on «our HPC» inside a conda env built on the compute node (internet there), not on «host».

Figures / tables: Fig 5CTable
C1
Reported
55 samples; 54675 probesets; GPL570
Reproduced
54675 probesets x 55 samples; GPL570
exact
C2
Reported
15-gene co-expression module (the 36th module)
Reproduced
15-gene lmQCM module (our #34)
exact
C3
Reported
module genes SP100,RSAD2,STAT2,MX1,ISG15,SAMHD1,XAF1,IFIT1,IFIT3
Reproduced
all 9 present (9/9); module = those 9 + CMPK2,IFI44,IFI44L,SP110,TNFSF13B,USP18
exact
C4
Reported
top GO term Type I interferon signaling (GO:0060337); overlap 9/148; p=2.51E-16
Reproduced
top GO term Type I interferon signaling (GO:0060337); overlap n=9 (same genes); p=1.12E-19 (Enrichr GO_BP_2021)
within tolerance
C4z
Reported
Enrichr z-score -3.2821
Reproduced
533.81 (Enrichr GO_BP_2021)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 79/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

This is an essentially 1:1 reproduction. The dataset dimensions (C1), 15-gene module (C2), the 9 Table-1 interferon genes (C3, 9/9), and the top GO term Type I interferon signaling GO:0060337 with the same 9 overlap genes (C4) all reproduced exactly from public GSE17537 and the authors' own lmQCM algorithm. The only non-matching values are the exact Enrichr p-value (2.51E-16 vs 1.12E-19) and z-score (-3.2821 vs 533.81), both soft and explained by Enrichr's GO_BP_2018→2021 library/formula change — a technical version drift on the annotation-tool side, not an authors' defect or fabrication concern. Central conclusion fully confirmed; severity negligible.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

159.1 k
tokens (I/O) · 11.4 M incl. cache
23 min
runtime · 0.06 CPU-h
7.6 GB
peak RAM
1
HPC jobs
hummel
machine