DESE: estimating driver tissues by selective expression of genes associated with complex diseases or traits.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL: 1:1 on the headline claim, honest mismatch on one sub-claim, rest out of reach of the shipped data. Reproduced DESE headless via KGGSEE --gene-assoc-condi on the shipped SCZ chr1 GWAS + GTEx v8 selective-expression on a «our HPC» compute node («job»). C1 (headline) REPRODUCES: all top-10 driver tissues for schizophrenia are brain regions at BOTH gene and transcript level, with frontal cortex BA9 in the top tier (gene rank #3, adj p=0.0134) and non-brain tissues far below (adj p>0.59); exact rank-1 shifts within the brain set (BA24/Hippocampus vs BA9) as expected from a chr1-only GWAS + GTEx v8 vs the paper's genome-wide GWAS + GTEx v7. C8 (iterative convergence) REPRODUCES exactly. C2 (transcript more powerful than gene for BA9) does NOT reproduce on the chr1 subset - gene-level BA9 is actually stronger here; an expected, honestly-recorded limitation of one-chromosome data, not a fabrication flag. C3-C7 NOT attempted: those trait GWAS are not shipped. Method described well enough to reproduce; code is a third-party-maintained CLI of the authors' own tool (KGGSEE), valid per P16. Exact genome-wide p-values were never expected to match; the qualitative driver-tissue claim - the paper's actual science - holds.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-20 ⛓ f98fd62ac2fc
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetTissue-selective expression of a disease's susceptibility genes (identified via GWAS) can be used to computationally identify the causal/driver tissues in which complex diseases or traits primarily arise.
- ★ DESE is a unified iterative framework that estimates driver tissues of complex diseases/traits from tissue-selective expression of GWAS-associated genes, and outputs prioritized susceptibility genes as a byproduct method
- ★ The robust-regression z-score, derived from Huber robust linear regression on ranked expression values, is a new, more powerful measure of tissue-selective expression than the conventional z-score, especially with multiple selectively expressed tissues method
- ★ Transcript-level selective expression detects more selectively expressed genes and yields higher statistical significance for driver-tissue estimation than gene-level selective expression finding
- ★ The lung is estimated as a driver tissue of rheumatoid arthritis, consistent with known involvement of lung autoimmune response in RA pathogenesis finding
- ★ Frontal cortex and anterior cingulate cortex are the top estimated driver brain regions for both schizophrenia and bipolar disorder finding
- ★ DESE-estimated driver tissues show high concordance with independently derived tissues from two existing methods (Ongen et al. eQTL-based method and LDSC-SEG) finding
- Liver is identified as the major driver tissue for total cholesterol, consistent with its role in endogenous cholesterol synthesis finding
- DESE is implemented in the KGG platform, and a webserver (REZ) is provided for online query of robust selective expression across tissues/cell types resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA-seq (bulk, gene- and transcript-level) | 50 human tissues, GTEx Project (V7) | none | robust-regression z-score of tissue-selective expression | GTEx RNA-Seq |
| GWAS summary statistics / conditional gene-based association test | Human, schizophrenia meta-GWAS cohort | none | driver tissue p-values and prioritized susceptibility genes | — |
| GWAS summary statistics analysis | Human, bipolar disorder cohort (20,129 cases, 54,065 controls) | none | driver tissue p-values | — |
| GWAS summary statistics analysis | Human, coronary artery disease cohort | none | driver tissue p-values | — |
| GWAS summary statistics analysis | Human, rheumatoid arthritis cohort | none | driver tissue p-values | — |
| GWAS summary statistics analysis | Human, total cholesterol trait cohort | none | driver tissue p-values | — |
| GWAS summary statistics analysis | Human, height (anthropometric trait) cohort | none | driver tissue p-values | — |
| Gene expression profiling (independent replication dataset) | Human tissues, GEO gene-level expression data | none | gene-level selective expression for driver tissue validation | GEO |
- ▲ Top driver tissue for schizophrenia had far higher significance using transcript-level vs gene-level selective expression 5.3E-13 (transcript) vs 2.0E-5 (gene)
- ▲ Transcript-level expression detected substantially more selectively expressed genes than gene-level expression across tissues on average 54% extra genes (up to 5.5-fold more unique genes)
- ▲ Lung ranked as second most significant driver tissue for rheumatoid arthritis p=4.2E-9
- ▲ Spleen and lymphocytes (immune tissues) among top driver tissues for rheumatoid arthritis p=7E-8 and p=1.3E-6
- ▲ Cerebellar hemisphere was top driver tissue for bipolar disorder p=1.3E-09 (transcript), 9.0E-06 (gene)
- ▲ Coronary artery was top driver tissue for coronary artery disease, with higher significance at transcript level 4.3E-6 (transcript) vs 9E-4 (gene)
- ▲ Liver was by far the most significant driver tissue for total cholesterol, with a large drop-off to the second-ranked tissue second tissue p=6.9E-8 vs 3.3E-5 relative to liver
- – Top 10 driver tissues for schizophrenia by DESE and by LDSC-SEG were both brain regions, with brain frontal cortex (BA9) ranked top by both tools
- pvalue 5.3E-13 (top schizophrenia driver tissue, transcript-level robust-regression z-score)
- pvalue 2.0E-5 (top schizophrenia driver tissue, gene-level robust-regression z-score)
- fold_change on average 54% extra selectively expressed genes (5.5-fold more unique genes) (transcript-level vs gene-level detection across 50 tissues)
- count 20,129 cases; 54,065 controls (bipolar disorder GWAS sample size)
- pvalue 4.2E-9 (lung as driver tissue for rheumatoid arthritis, transcript-level)
- correlation r ∈ [0.3, 0.6] (Spearman) (tissues with moderate correlation between original and selective expression)
- count 27 significant tissues (p < 10^-3) (driver tissues detected for height at transcript-level selective expression)
- pvalue 6.9E-8 vs 3.3E-5 (liver vs second-ranked tissue significance for total cholesterol)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational/statistical genomics methods paper introducing DESE, a framework that estimates disease-driver tissues from GWAS summary statistics and tissue gene-expression profiles (GTEx). The core statistic is a novel robust-regression z-score (built on Huber robust regression) for quantifying tissue-selective expression, combined with a previously published conditional gene-based association test applied to GWAS p-values; results across six diseases/traits are reported primarily as tissue-level significance p-values, benchmarked against two other published methods (Ongen et al. and LDSC-SEG) and against simulation and correlation analyses (Pearson/Spearman) of expression profiles.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Robust-regression z-score (Huber robust linear regression extension) | quantifying tissue-selective expression of genes across 50 GTEx tissues (gene- and transcript-level) | — | not stated |
| Conditional gene-based association test (previously published method, ref. 22) | detecting susceptibility genes from GWAS summary p-values for six diseases/traits | — | not stated |
| Alternative selective-expression measures: conventional z-score, MAD robust z-score, ratio of vector-scalar projection | comparison of driver-tissue significance across schizophrenia, bipolar disorder, CAD, RA, total cholesterol, height | — | not stated |
| Combined ranking by averaging -log10(p) across the four selective-expression measures | final driver-tissue prioritization shown in Fig. 3 for all six phenotypes | — | na |
| Bonferroni correction | correcting for multiple transcripts tested per gene when detecting selectively expressed genes | — | stated |
| Pearson correlation | comparing tissue-pair similarity based on robust-regression z-scores vs. original TPM expression values | — | not stated |
| Spearman correlation | comparing original expression values vs. selective expression values within the same tissue (Additional file 1: Figures S6–S7) | — | not stated |
-
Multiple transcripts per gene were corrected for using the Bonferroni method.↳ Could also: A false discovery rate (FDR/Benjamini-Hochberg) correction — FDR-based correction is often preferred when testing many transcripts/genes because it can offer greater power to detect true selectively expressed transcripts while still controlling the expected proportion of false positives, which can be useful when Bonferroni's conservatism is a concern for exploratory discovery.
-
Four independent selective-expression measures (robust-regression z-score, conventional z-score, MAD robust z-score, vector-scalar projection ratio) were combined by averaging their -log10(p) values into one ranking.↳ Could also: A formal p-value combination method such as Fisher's combined probability test or Stouffer's Z-score method — These established meta-analytic approaches combine p-values from multiple tests with defined statistical properties and significance thresholds, which can complement a simple averaging of -log10(p) for summarizing agreement across methods.
-
Agreement between DESE and the two comparator methods (Ongen et al., LDSC-SEG) was assessed by visually comparing overlap in top-ranked tissues.↳ Could also: A formal concordance statistic, such as Spearman/Kendall rank correlation or a hypergeometric enrichment test for overlap in top-N tissue lists — A quantitative concordance test would provide a p-value or effect size summarizing how much agreement between methods exceeds what would be expected by chance, complementing the qualitative overlap description.
-
Correlation coefficients (Pearson, Spearman) between tissues and between expression/selective-expression values were reported as point estimates.↳ Could also: Bootstrap or asymptotic confidence intervals around the correlation coefficients — Reporting an interval alongside each correlation coefficient would convey the precision of the estimate, which can be informative when coefficients are compared across many tissue pairs.
-
Statistical significance of driver tissues was derived analytically from the robust-regression z-score's theoretical null distribution (checked via QQ plots).↳ Could also: A permutation-based null distribution (e.g., shuffling gene-phenotype associations or expression labels) — Permutation approaches can serve as an additional or alternative way to empirically validate the null distribution and calibrate p-values, particularly useful as a complement when assumptions underlying a parametric approach are of interest to examine further.
-
The number of significant driver tissues per phenotype (e.g., 27 tissues for height) was reported using a fixed p-value threshold (p < 10^-3).↳ Could also: An explicit multiple-testing-adjusted threshold (e.g., FDR q-value) applied across all 50 tissues tested per phenotype — Since many tissues are tested simultaneously for each phenotype, applying and reporting a family-wise or FDR-adjusted threshold across tissues would provide an additional lens on how many tissues remain significant after accounting for the number of tissues examined.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-31694669 (DESE; Jiang et al., Genome Biol 2019)
Paper: DESE: estimating driver tissues by selective expression of genes associated with complex diseases or traits. Jiang L, Xue C, Dai S, Chen S, Chen P, Sham PC, Wang H, Li M. Genome Biol 2019;20:233. PMID 31694669 · PMCID PMC6836538 · DOI 10.1186/s13059-019-1801-5.
Method (DESE). A unified iterative framework that detects causal/"driver" tissues of a complex trait from GWAS summary statistics + multi-tissue expression profiles. Three components run to convergence:
- Conditional gene-based association (ECS — effective chi-square test) to call trait-associated genes from GWAS p-values, conditioning on LD.
- Driver-tissue estimation — Wilcoxon/Mann–Whitney U test of whether the associated genes show elevated selective expression in each tissue (a robust-regression z-score per gene×tissue).
- Gene re-ranking by selective expression in the prioritized tissues, fed back into (1) until tissue p-values stabilize.
Implemented originally in KGG v4.1 (Java GUI; zenodo 10.5281/zenodo.3367790
= source only). The same lab's command-line reimplementation is KGGSEE
(github.com/pmglab/KGGSEE), which exposes DESE via --gene-assoc-condi. Per
brief rule P16, applying the same-method maintained tool to the paper's data is
a valid reproduction path; KGGSEE is the only practical headless route (KGG 4.1
is a GUI NetBeans app).
IN SCOPE (pipeline-derived → attempted)
| # | Reported result | Pipeline |
|---|---|---|
| C1 | Schizophrenia driver tissues are brain regions (all top-10 brain); top = frontal cortex BA9 | DESE (ECS + Wilcoxon selective-expression) |
| C2 | SCZ transcript-level selective expression more powerful than gene-level (e.g. BA9 5.3E-13 vs 2.0E-5) | DESE gene vs transcript |
| C3 | Total cholesterol top driver tissue = liver | DESE |
| C4 | Coronary artery disease top tissue = coronary artery / artery / adipose | DESE |
| C5 | Rheumatoid arthritis top tissues = immune + lung + GI (spleen, lung, ileum, colon) | DESE |
| C6 | Bipolar disorder top tissue = brain (cerebellar hemisphere / cortex) | DESE |
| C7 | Height: many (~27) significant tissues incl. cardiovascular/fibroblast | DESE |
Primary attempt: C1/C2 (schizophrenia) using the KGGSEE-shipped SCZ GWAS + GTEx v8 gene- and transcript-level selective-expression resources — this is the documented DESE tutorial and directly tests the paper's headline claim. Additional traits attempted as data permits.
OUT OF SCOPE / not attempted (and why)
- Exact p-value 1:1 match. Original used KGG v4.1 + GTEx v7 + PGC SCZ2 (2014); the reproducible toolchain is KGGSEE + GTEx v8 + a newer SCZ GWAS (the shipped sumstats are the PGC3-era EUR set, Nca=53386/Nco=77258). Different expression build, gene models, and GWAS → exact p-values are NOT expected to match. The reproducible target is the qualitative driver-tissue ranking (which tissues top the list), which is the paper's actual scientific claim.
- The "robust-regression z-score" selective-expression values themselves
(provided as a precomputed resource at the now-defunct HKU
rezserver) — not recomputed; we consume the maintained GTEx-v8 selective-expression resource. - Wet-lab / literature-support gene counts (e.g. "40 vs 17 RA genes with literature support") — manual literature curation, not pipeline output.
- BrainSpan and GEO expression analyses — secondary validations; primary resource (GTEx) suffices to test the core claim.
Data
- Software: KGGSEE jar (pmglab.top), KGG v4.1 source (zenodo 3367790).
- Expression resource: GTEx v8 TMM selective-expression mean/SE, gene (54 tissues) + transcript level (KGGSEE resources.zip).
- GWAS: schizophrenia EUR summary statistics (KGGSEE tutorials.zip; chr1 subset shipped — see limitation note in AUDIT.md).
- LD reference: 1000 Genomes Phase3 EUR (503 individuals; chr1 subse
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is an incomplete, in-progress reproduction: the DESE/KGGSEE run was still executing, no claim was graded (claims_graded=false, agreement.json status not-run-yet), and C1/C2/C8 remain PENDING. The limitations are on our/data-availability side, not the authors': only the SCZ GWAS ships with the tool (C3–C7 not attempted for lack of deposited GWAS), and GTEx v8 + a newer GWAS mean even the SCZ check is only an indirect magnitude/direction comparison. There is no fabrication signal and no observed discrepancy — the headline brain-driver-tissue claim is simply unconfirmed rather than contradicted, so everything is graded yellow pending completion.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.