TSUNAMI: Translational Bioinformatics Tool Suite for Network Analysis and Mining.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡A deviation arose in the data or preprocessing
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce without author contact; essentially 1:1. TSUNAMI is a Shiny wrapper around the authors' lmQCM algorithm; we ran lmQCM 0.2.4 (CRAN, same algorithm) on the paper's own data GSE17537 with the verbatim preprocessing + parameters from the repo. C1 (dataset 54675x55, GPL570), C2 (15-gene module) and C3 (the 9 Table-1 interferon genes, 9/9) reproduced EXACTLY; C4 top GO term (Type I interferon signaling GO:0060337) and its 9 overlap genes reproduced exactly, with the exact p-value differing only because the paper used the now-unavailable Enrichr GO_BP_2018 library (within-tol). Only the Enrichr z-score (C4z, a known version-dependent soft value) does not match numerically. NOT attempted: survival KM curves (Fig 6, qualitative/no pinnable number) and the web UI/Circos rendering (interface, not numeric results). No fabrication concerns: all reported values are derivable from the shipped algorithm + public data.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 79assessed: 2026-06-14 ⛓ 30871629b604
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThere is no existing tool that provides a complete, easy-to-use pipeline for mining relatively small, tightly connected (and potentially overlapping) gene co-expression network modules from transcriptomic data and directly linking them to downstream gene set enrichment, visualization, and survival analysis; TSUNAMI was built to fill this gap.
- ★ TSUNAMI is a freely accessible web-based tool suite that mines gene co-expression network (GCN) modules from public (GEO, TCGA) or user-uploaded numerical omics data and performs downstream gene set enrichment analysis. resource
- ★ The lmQCM algorithm mines smaller, more densely connected GCN modules than WGCNA and allows overlapping module membership, better reflecting biological pathway co-regulation. method
- ★ TSUNAMI provides direct search/retrieval interfaces to NCBI GEO and TCGA databases as well as user file upload (CSV, TSV, XLSX, TXT). method
- ★ TSUNAMI integrates Enrichr for downstream gene set enrichment analysis across 14 default categories. method
- TSUNAMI generates Circos plots (via R package circlize) to visualize chromosomal locations and pairwise relationships of genes within a GCN module. method
- ★ TSUNAMI includes a survival analysis module that correlates GCN module eigengenes (dichotomized at the median) with patient overall/event-free survival using the log-rank test. method
- In the GSE17537 example dataset, the 36th lmQCM-derived GCN module (15 genes) is significantly enriched for the type I interferon signaling pathway. finding
- ★ Smaller gene modules derived from lmQCM tend to generate more biologically meaningful gene set enrichment results than larger WGCNA-style modules. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Gene co-expression network mining (lmQCM algorithm) | GSE17537 microarray dataset, 55 colorectal cancer patients (Vanderbilt Medical Center) | none | Merged GCN modules (gene clusters) sorted by size; eigengene matrix | Affymetrix HU133 2.0 Plus GeneChip (GPL570) |
| Gene co-expression network mining (WGCNA) | GSE17537 microarray dataset | none | Hierarchical clustering of gene modules | R package WGCNA (Bioconductor) |
| Gene set enrichment analysis | 36th GCN module (15 genes) from GSE17537, mined by lmQCM | none | Enriched terms, P value, Z-score, overlapping genes | Enrichr |
| Circos plot gene locus visualization | 36th GCN module (15 genes), human genome hg38/hg19 | none | Chromosomal positions and pairwise gene links | R package circlize; UCSC refGene database |
| Survival analysis (log-rank test, Kaplan-Meier) | GSE17537, 55 samples with overall survival data | none | P value of log-rank test comparing high vs. low eigengene groups (median split) for OS/EFS | R package survival (survdiff function) |
| GEO dataset retrieval/processing test | First 1000 GSE datasets from NCBI GEO | none | Number of datasets successfully processed vs. failed | — |
- ▲ The 36th GCN module (15 genes) is highly enriched in the GO Biological Process term 'type I interferon signaling pathway (GO:0060337)' 9/148 genes overlap, P=2.51E-16
- – Only a small portion of tested GSE datasets failed to process with TSUNAMI, mostly legacy microarray data with excessive missing data or small sample size 12 out of first 1000 GSE datasets
- – Kaplan-Meier survival analysis was generated for the 36th GCN module eigengene dichotomized at the median into high/low groups
- pvalue 2.51E-16 (Enrichr enrichment of 36th GCN module for type I interferon signaling pathway (GO:0060337), 9/148 gene overlap)
- other Z-score = -3.2821 (Enrichment result for type I interferon signaling pathway in 36th GCN module)
- pvalue 1.80E-09 (Enrichment for 'cellular response to type I interferon (GO:0071357)', 4/23 overlap)
- count 12 out of 1000 (GSE datasets from GEO that could not be processed by TSUNAMI)
- count 55 (Colorectal cancer patients in example dataset GSE17537 used for survival analysis)
- count 54,675 probesets (Number of probesets on the Affymetrix HU133 2.0 Plus GeneChip platform used for GSE17537)
- other γ=0.7, λ=1, t=1, β=0.4, minimum cluster size=10 (Default lmQCM parameters used for GSE17537 example analysis)
- other power=10, reassign threshold=0/1, merge cut height=0.25, minimum module size=10 (WGCNA parameters used for GSE17537 example analysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper describes TSUNAMI, a bioinformatics web tool suite for gene co-expression network (GCN) mining and downstream analysis; statistical methods are demonstrated on a single illustrative dataset (GSE17537, n=55 colorectal cancer patients) rather than testing a primary biological hypothesis. GCN modules are constructed using Pearson or Spearman correlation-based algorithms (lmQCM or WGCNA). Downstream analyses include gene set enrichment via Enrichr (reporting P values and Z-scores across 14 databases) and Kaplan-Meier survival analysis comparing median-split eigengene groups via the log-rank test.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Log-rank test (R survdiff) | Overall survival comparison between high vs. low GCN module eigengene groups (Figure 6, 36th module from GSE17537) | 55 | not stated |
| Enrichr enrichment analysis (Fisher's exact test, internal to Enrichr) | Gene set enrichment of GCN module gene lists across 14 databases (Table 1 and associated panels) | Gene overlap counts stated per term (e.g., 9/148 for top GO term) | not stated |
| Pearson correlation coefficient (PCC) | Pairwise gene correlation for GCN module construction via lmQCM (default setting, GSE17537 example) | 55 | not stated |
| Spearman rank correlation coefficient (SCC) | Alternative pairwise gene correlation for GCN module construction (recommended for RNA-seq data per tool documentation) | — | not stated |
| Singular value decomposition (first principal component as eigengene) | Summarization of gene expression within each GCN module into a single eigengene value (Figure 5B) | 55 | na |
-
Survival analysis groups patients into high and low eigengene expression by median split before applying the log-rank test↳ Could also: Cox proportional hazards regression with eigengene as a continuous predictor could also be used — Retaining eigengene as a continuous variable avoids information loss from dichotomization and yields a hazard ratio with confidence interval, which quantifies effect magnitude; median split can produce different results depending on the specific cut point chosen
-
The log-rank test is applied to compare survival between the two eigengene groups without covariate adjustment↳ Could also: Multivariable Cox regression adjusting for available clinical covariates (e.g., stage, age) could also be applied — Covariate-adjusted models can separate the independent prognostic contribution of a GCN module from confounders, which is relevant when demonstrating clinical utility of a biomarker module
-
Fourteen enrichment analyses are performed concurrently across databases without an explicitly stated multiple-testing correction↳ Could also: Applying Benjamini-Hochberg FDR correction within each database (already implemented internally by Enrichr as adjusted P values) and noting this correction explicitly would also be standard practice — Reporting the adjusted P values available from Enrichr alongside raw P values would make the degree of correction transparent to readers and facilitate comparison across studies
-
Pearson correlation is used as the default measure for GCN construction on the microarray demonstration dataset↳ Could also: Spearman rank correlation (also available in TSUNAMI) or mutual information-based measures could also be used — Spearman correlation is more robust to outliers and non-normality; mutual information captures non-linear dependencies; the paper itself recommends Spearman for RNA-seq, so the choice of measure can be matched to the distributional properties of the input data
-
Tool performance and enrichment results are demonstrated on a single example dataset (GSE17537, n=55)↳ Could also: Benchmarking on multiple datasets with known biological ground truth (e.g., simulated data or datasets with validated pathway activity) could also be used — Multi-dataset demonstration would allow readers to assess the consistency and generalizability of the tool's outputs across different platforms, sample sizes, and disease contexts
-
GCN module eigengenes are derived as the first principal component of within-module expression via SVD without variance explained being reported↳ Could also: Reporting the proportion of variance explained by the first PC for each module could also accompany the eigengene output — The variance-explained fraction conveys how well a single eigengene summarizes the module; a low value would indicate heterogeneous expression within the module, which is useful context for interpreting downstream survival or enrichment results
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
cellular response to type I interferon (GO:0071357) is enriched in co-expression module 36 from colorectal cancer microarray data (P = 1.80e-09)other human colorectal-cancer 2021×1papers★ This paper is the founder (earliest)
-
type I interferon signaling pathway (GO:0060337) is significantly enriched in a co-expression network module derived from colorectal cancer microarray data (P = 2.51e-16)other human colorectal-cancer 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33705981 (TSUNAMI)
Paper: Huang Z. et al. TSUNAMI: Translational Bioinformatics Tool Suite for
Network Analysis and Mining. Genomics Proteomics Bioinformatics 2021.
PMID 33705981 · PMCID PMC9403021 · DOI 10.1016/j.gpb.2019.05.006
Code: https://github.com/huangzhii/TSUNAMI (R Shiny app; default branch master)
Data: GEO GSE17537 (colorectal cancer, Affymetrix HG-U133 Plus 2.0 / GPL570)
Nature of the paper
TSUNAMI is a web tool (R Shiny) wrapping the authors' own lmQCM
co-expression module-mining algorithm (also on CRAN as package lmQCM, same
authors), plus Enrichr enrichment, Circos visualisation and survival analysis.
There is no wet-lab content. The only concrete, reported numeric results are
in the case study on GSE17537 (Section "An example", Figures 5–6, Table 1).
Per P16 of the brief, applying the authors' own published tool to the paper's own
data is a fully valid reproduction.
In scope (pipeline-derived, attempted)
The case-study pipeline is fully specified in the paper + source:
- Data ingest —
getGEO("GSE17537", GSEMatrix=TRUE, AnnotGPL=FALSE); expression matrixexprs(); probe→gene-symbol map from GPL570 platform table column "Gene Symbol" (server.R:240,448-472). - Preprocessing (
server.R:575-655,utils.R:varFilter2):- mean filter: remove genes below the 50th percentile of row means (J=50);
- variance filter: remove genes below the 10th percentile of row variance (K=10);
- drop empty symbols; collapse duplicate gene symbols keeping the max-mean probe.
- lmQCM module mining (
server.R:710-...): Pearson correlation, γ=0.7, t=1, λ=1, β=0.4, minClusterSize=10, normalization=FALSE; modules sorted by size descending. - GO enrichment of a module via Enrichr (GO Biological Process).
Claims to reproduce (see original/claims.tsv)
- C1 GSE17537 = 55 samples × 54,675 probesets on GPL570 (paper text).
- C2 A co-expression module of 15 genes (the "36th module", Fig 5C).
- C3 That module contains the 9 genes listed in Table 1: SP100, RSAD2, STAT2, MX1, ISG15, SAMHD1, XAF1, IFIT1, IFIT3.
- C4 GO enrichment of that module → top term Type I interferon signaling pathway (GO:0060337), overlap 9/148, p = 2.51E-16 (Table 1).
Out of scope / not attempted (with reason)
- Survival analysis (Fig 6) — KM curves are qualitative (no reported number to
pin); log-rank on a module eigengene is straightforward but yields no specific
printed value to compare →
no_expected_resultfor that panel. May add as a bonus if core claims reproduce. - The web UI / Circos rendering / GEO browser — interface features, not numeric results.
- The Z-score (-3.2821) in Table 1 is Enrichr's combined-score component; it is sign/version-dependent, recorded but treated as soft.
Heavy-compute note
Compute is modest (55 samples; ~13k-gene Pearson matrix) but per the hard rule it runs on «our HPC» inside a conda env built on the compute node (internet there), not on «host».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is an essentially 1:1 reproduction. The dataset dimensions (C1), 15-gene module (C2), the 9 Table-1 interferon genes (C3, 9/9), and the top GO term Type I interferon signaling GO:0060337 with the same 9 overlap genes (C4) all reproduced exactly from public GSE17537 and the authors' own lmQCM algorithm. The only non-matching values are the exact Enrichr p-value (2.51E-16 vs 1.12E-19) and z-score (-3.2821 vs 533.81), both soft and explained by Enrichr's GO_BP_2018→2021 library/formula change — a technical version drift on the annotation-tool side, not an authors' defect or fabrication concern. Central conclusion fully confirmed; severity negligible.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.