Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Identification of immune-related signatures and pathogenesis differences between thoracic aortic aneurysm patients with bicuspid versus tricuspid valves via wei

PLoS One · 2023
L1 97/100 PQI 99
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
How its reproducibility compares
97/100
Reproducibility score
1.3 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 92% of all assessed papers rank 80 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1. Third-party-style tutorial repo (4 messy R scripts with hard-coded Windows paths) re-implemented faithfully on «our HPC» against the public GSE5180 discovery set (13 BAV + 12 TAV, GPL96). All four core pipeline numbers land: 7 WGCNA modules at beta=22 (exact); the immune module-trait correlation 0.36 (reproduced 0.361, exact); the seven immune hub genes all fall in that one module (exact); and the 3-gene logistic AUC 0.87 (reproduced 0.872, exact). One documented nuance: the trait-correlated immune module (r=0.36, holding all hub genes) is labeled BLUE in our run, not BROWN as the paper states (our brown=-0.18) -- a WGCNA color-label artifact (size-rank assigned, shifts with the probe->symbol collapse) that is itself internally inconsistent in the authors' own code (WGCNA.txt narrates 'brown' but its Cytoscape export uses module='blue'). No fabrication indicators: every value is derivable from the shipped data + described pipeline. NOT attempted (hard ~20%): external-validation AUC=0.79 on GSE83675/GSE26155/GSE61128 (multi-platform incl. RNA-seq + exon arrays), GSEA/GSVA exact enriched-term lists (mouse/human-mixed, gmt-dependent), CIBERSORT per-cell-type deltas (needs LM22 + stochastic SVR perms), and DCA decision curves (visual).

💻 Code ↗ 🗄 Data: GSE5180

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 97
    assessed: 2026-06-15 ⛓ ec4fae4d6056
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Although both BAV (bicuspid aortic valve) and TAV (tricuspid aortic valve) patients can develop thoracic aortic aneurysm (TAA), the paper tests whether TAA/BAV and TAA/TAV differ in their underlying pathophysiological processes and molecular mechanisms, particularly with respect to immune-related gene expression and immune cell infiltration.

Core claims
  • TAA/TAV pathogenesis is more strongly associated with immune-related gene expression than TAA/BAV. finding
  • WGCNA identified two gene modules (brown from GSE5180, blue from GSE26155) most correlated with TAA/TAV, and both are enriched for immune-related functions. finding
  • TAA/TAV and TAA/BAV tissues show differing immune cell infiltration proportions, notably in dendritic, mast, CD4 memory T cells (higher in TAV) and monocytes, B cells, CD8 T cells (higher in BAV). finding
  • CD86, ITGB2, and ITGAM were identified as three signatures most strongly associated with TAA/TAV onset, with ROC AUC values approximating 0.8 or above. finding
  • WGCNA, GO enrichment, GSVA/ssGSEA, CIBERSORT, PPI network construction, and stepwise regression/ROC/DCA were combined as an analytic pipeline to identify TAA/TAV-associated immune signatures. method
  • 153 genes shared between the brown and blue WGCNA modules were significantly enriched for immune response functions per PPI/STRING analysis. finding
Experimental setups
Assay System Perturbation Readout Platform
WGCNA (weighted gene co-expression network analysis) human aortic aneurysm tissue (GSE5180, 25 samples: 13 BAV, 12 TAV) none (BAV vs TAV disease comparison) gene co-expression modules correlated with TAA/TAV phenotype GPL96 array; R package WGCNA
WGCNA (weighted gene co-expression network analysis) human aortic adventitia/intima-media tissue (GSE26155, 96 samples, all TAV) none (dilation status comparison) gene co-expression modules correlated with aortic dilation phenotype GPL5175 array; R package WGCNA
Gene Ontology (GO) enrichment analysis genes from brown and blue WGCNA modules none enriched biological process terms R package clusterProfiler
Gene Set Variation Analysis (GSVA) aortic smooth muscle cells (GSE61128, 7 samples: 4 BAV, 3 TAV) none (BAV vs TAV comparison) immune-associated gene set enrichment scores GPL5175 array; R packages GSVA and limma
CIBERSORT immune cell deconvolution human aortic tissue gene expression datasets none (BAV vs TAV comparison) relative proportions of 22 infiltrating immune cell types CIBERSORT web portal; e1071, parallel, preprocessCore R packages
Protein-protein interaction (PPI) network analysis 153 genes shared between brown and blue WGCNA modules none hub genes ranked by degree and betweenness centrality stringApp in Cytoscape v3.6.0, CytoHubba plugin, confidence cutoff 0.4
Logistic regression, ROC, and decision curve analysis (DCA) human aortic tissue, training (GSE5180) and testing (GSE83675, 16 samples: 9 BAV, 7 TAV) none predictive accuracy (AUC) of hub gene signatures for TAA/TAV ROCR package; rmda package
ssGSEA (single sample Gene Set Enrichment Analysis) human aortic tissue datasets none correlation between CD86/ITGB2/ITGAM expression and aortic-aneurysm-associated biological process scores GSEA/MSigDB C5 gene sets; enrichplot, pheatmap R packages
Key results
  • Brown module (262 genes, GSE5180) had the highest correlation with TAA/TAV phenotype correlation coefficient 0.36
  • Blue module (847 genes, GSE26155) had the highest correlation with adventitia dilation phenotype correlation coefficient 0.54
  • Both brown and blue modules significantly enriched for immune-related GO terms (T cell activation, leukocyte migration, lymphocyte proliferation, immune response regulation); enrichment more pronounced in blue module
  • GSVA showed immunity-associated gene expression significantly up-regulated in TAA/TAV vs TAA/BAV in aortic smooth muscle cells P<0.05
  • Dendritic, mast, and activated CD4 memory T cell proportions significantly higher in TAA/TAV; monocytes, B cells, CD8 T cells significantly higher in TAA/BAV
  • 153 genes shared between brown and blue modules, significantly enriched for immune responses per STRING analysis 153 genes
  • CD86, ITGB2, ITGAM identified as signatures most strongly associated with TAA/TAV, with strong predictive value AUC ~0.8 or above
Key statistics
  • correlation 0.36 (brown module (262 genes) correlation with TAA/TAV phenotype in GSE5180)
  • correlation 0.54 (blue module (847 genes) correlation with adventitia dilation phenotype in GSE26155)
  • other soft threshold β=22, R2=0.8 (WGCNA scale-free topology fit for GSE5180)
  • other soft threshold β=14, R2=0.8 (WGCNA scale-free topology fit for GSE26155)
  • count 153 genes (genes shared between brown and blue WGCNA modules)
  • other AUC approximating 0.8 or above (ROC analysis of CD86, ITGB2, ITGAM for predicting TAA/TAV onset)
  • pvalue P<0.05; FDR<0.01 (significance thresholds used for GO enrichment and GSVA analyses)
  • count GSE5180 n=25 (13 BAV,12 TAV); GSE26155 n=96; GSE83675 n=16 (9 BAV,7 TAV); GSE61128 n=7 (4 BAV,3 TAV) (sample sizes of the four GEO datasets used)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics study re-analyzed four publicly available GEO microarray datasets to compare gene expression in thoracic aortic aneurysm tissue from patients with bicuspid (BAV) versus tricuspid (TAV) aortic valves. The primary analytical approach was WGCNA to identify co-expression modules correlated with valve phenotype, followed by GO enrichment analysis, GSVA, and CIBERSORT-based immune deconvolution. Hub gene signatures were identified via PPI network analysis, screened by stepwise logistic regression with AIC, and their discriminative ability was reported as AUC from ROC analysis on separate training and testing datasets.

Replicationbiological Sample sizeSample sizes stated per GEO dataset in Table 1 (GSE5180 n=25, GSE26155 n=96, GSE83675 n=16, GSE61128 n=7); no formal a priori power analysis described GroupsTAA/BAV vs. TAA/TAV (aortic tissue); dilated vs. non-dilated vs. borderline TAV (GSE26155) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
Pearson correlation (module-trait correlation within WGCNA framework) WGCNA module selection in GSE5180 and GSE26155; also for correlating three signatures with ssGSEA scores of biological processes 25 (GSE5180); 96 (GSE26155) not stated
GO enrichment analysis (hypergeometric test via clusterProfiler) Functional annotation of brown (262 genes, GSE5180) and blue (847 genes, GSE26155) WGCNA modules not stated
Gene Set Variation Analysis (GSVA) with limma-based differential scoring Validation of immune gene enrichment differences between TAA/TAV and TAA/BAV in GSE61128 (smooth muscle cells) 7 (4 BAV, 3 TAV) not stated
Wilcoxon rank-sum test Differences in relative proportions of 22 immune cell types (from CIBERSORT) between TAA/BAV and TAA/TAV groups not stated
Logistic regression (univariate) and multivariate stepwise logistic regression with AIC Screening and selection of immune-related signature genes (CD86, ITGB2, ITGAM) for TAA/TAV prediction from hub gene candidates 25 (GSE5180 training set) not stated
ROC analysis / AUC (via ROCR package) Validation of predictive accuracy of CD86, ITGB2, ITGAM for TAA/TAV onset; training set GSE5180, testing set GSE83675 25 (training, GSE5180); 16 (testing, GSE83675) not stated
Decision curve analysis (DCA, via rmda package) Net benefit evaluation of the three identified signatures across a range of threshold probabilities 25 (GSE5180); 16 (GSE83675) not stated
PPI network degree and betweenness centrality ranking (CytoHubba, Cytoscape 3.6.0) Hub gene identification from 153 shared genes between brown and blue WGCNA modules 153 genes na
Approaches that could also have been used
  • Immune cell proportion differences across 22 cell types were tested with separate Wilcoxon rank-sum tests, with FDR correction applied across comparisons
    Could also: A multivariate compositional analysis (e.g., MANOVA on CLR-transformed proportions, or a Dirichlet regression) could also have been used to jointly model the full 22-cell composition in one test — Immune cell proportions from CIBERSORT sum to 1 (compositional data), so joint modeling respects their inter-dependence; separate pairwise tests treat each cell type independently, which is an alternative framing that is also widely used in the literature
  • Stepwise logistic regression with AIC was used for signature gene selection from a pool of hub gene candidates in a training dataset of n=25
    Could also: Penalized regression (LASSO or elastic net via glmnet) could also have been used for variable selection in this high-candidate, small-n setting — Penalized regression simultaneously performs shrinkage and selection and has well-characterized behavior under p >> n conditions; both approaches are standard for biomarker selection from candidate gene lists
  • Module-trait correlations in WGCNA were computed with Pearson correlation between module eigengenes and phenotype
    Could also: Spearman rank correlation or a linear model (limma) could also have been applied for the module-trait association step — Spearman correlation is less sensitive to outliers in small samples (n=25 for GSE5180), and a linear model framework would allow covariate adjustment; all three approaches are used in WGCNA-based studies
  • Predictive accuracy of the three signatures was summarized as AUC point estimates from ROC analysis
    Could also: AUC confidence intervals (e.g., DeLong method or bootstrap) and direct comparison of AUCs between signatures could also have been reported — With a testing set of n=16, AUC estimates carry substantial uncertainty; 95% CIs would convey that uncertainty and allow formal comparison of the three markers' discriminative ability
  • GSVA scores were compared between TAA/TAV and TAA/BAV in GSE61128 (n=7 total) using a P < 0.05 threshold
    Could also: A permutation-based test or exact test could also have been used given the very small group sizes (4 BAV, 3 TAV) — Asymptotic p-value approximations underlying standard limma moderated t-tests may be less reliable at n=3 and n=4; permutation approaches make fewer distributional assumptions at these sample sizes
  • Results throughout are reported using only significance thresholds (asterisk tiers) without point estimates of group means, medians, or dispersion for continuous outcomes
    Could also: Reporting group medians with IQR (for non-parametric comparisons) or means with SD alongside p-values could also have been included — Effect magnitude and spread complement statistical significance and allow readers to assess practical relevance and compare findings across studies; this is particularly informative when group sizes differ substantially across the four datasets
Software: R/WGCNA · R/clusterProfiler · R/GSVA · R/limma · R/ROCR · R/rmda · R/ggplot2 · R/enrichplot · R/pheatmap · Cytoscape/stringApp/CytoHubba 3.6.0 (Cytoscape) · CIBERSORT (web portal)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
3
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE26155 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE5180 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE61128 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE83675 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37883426

Paper: Huang M et al. "Identification of immune-related signatures and pathogenesis differences between thoracic aortic aneurysm patients with bicuspid versus tricuspid valves via weighted gene co-expression network analysis." PLoS One 2023. PMID 37883426 / PMC10602290 / DOI 10.1371/journal.pone.0292673.

Code: https://github.com/Amy-huangm/Network-analyses-of-aneurysms (commit dc26fcd7c1e92adacde727e764d53e7252ec26d4) — 4 R scripts shipped as .txt (WGCNA.txt, CIBERSORT.txt, GO+GSEA+GSVA.txt, ROC+DCA.txt). They are tutorial-style with hard-coded C:«path» paths and copy-paste artefacts (mouse org.Mm.eg.db mixed into the human GO script). Not runnable as-is; we re-implement the documented pipeline faithfully.

Data: Primary discovery set GSE5180 (GPL96, Affymetrix HG-U133A) — 25 ascending-aortic-aneurysm tissues, 13 BAV + 12 TAV. Public on GEO. Validation sets GSE26155 / GSE83675 / GSE61128 used for external AUC.

Pipeline-derived results (what each comes from)

# Reported result Pipeline In scope?
C1 WGCNA: 7 modules at soft-power β=22 (R²=0.8) on GSE5180 WGCNA blockwiseModules on top-5000 MAD genes YES — clearly specified, deterministic
C2 Brown module = 262 genes, module–trait correlation 0.36 with TAV WGCNA moduleEigengenes + cor(MEs, design) YES
C3 7 candidate hub immune genes (TYROBP, PTPRC, CD86, ITGB2, ITGAM, CSF1R, LCP2); 3 final (CD86, ITGB2, ITGAM) hub selection (STRING/Cytoscape — underspecified) PARTIAL — verify the 7 genes' module assignment; full hub ranking not specified
C4 3-gene logistic signature AUC = 0.87 (training, GSE5180) glm logistic + ROC YES — deterministic
C5 CIBERSORT immune-cell differences BAV vs TAV (DCs, mast cells, CD4 mem ↑ in TAV; monocytes, B, CD8 ↑ in BAV) CIBERSORT (LM22, perm=100, QN) PARTIAL — needs LM22; SVR has stochastic perm p-values
C6 GO/GSEA: immune response, T-cell activation, leukocyte migration clusterProfiler GSEA/enrichGO OUT (last-20%) — script is mouse/human-mixed, needs MSigDB gmt; low ROI

Out of scope / not attempted (the hard ~20%)

  • External validation on GSE26155/GSE83675/GSE61128 (different platforms incl. RNA-seq GPL17077, exon arrays GPL5175) — AUC=0.79 test claim. Multi-platform harmonization, low 80/20 ROI.
  • GSEA/GSVA exact enriched-term lists (C6) — under-specified, gmt-dependent.
  • DCA decision-curve plots — visual, no extractable number.

Reproduction strategy

ONE «our HPC» R/conda job (GEOquery + WGCNA + pROC, optional CIBERSORT): fetch GSE5180 series matrix + GPL96 inside the job, collapse probes→symbol (max mean), log2 per the paper's heuristic, run blockwiseModules(power=22, maxBlockSize=6000, TOMType="unsigned", minModuleSize=30, mergeCutHeight=0.25), module–trait correlation against BAV/TAV, locate brown module + the 7 hub genes, then a logistic-regression ROC for CD86+ITGB2+ITGAM. Compare C1, C2, C3, C4.

C1
Reported
7 WGCNA modules at soft-power beta=22 (GSE5180)
Reproduced
7 modules (6 colored: turquoise/blue/brown/yellow/green/red + grey)
exact
C2a
Reported
brown module = 262 genes
Reproduced
248 genes (brown); 305 (blue)
within tolerance
C2b
Reported
module-trait correlation 0.36 with valve subtype
Reproduced
0.361 (on the BLUE module that holds all 7 hub genes; our BROWN = -0.18)
exact
C3
Reported
7 immune hub genes TYROBP/PTPRC/CD86/ITGB2/ITGAM/CSF1R/LCP2
Reproduced
all 7 co-located in the single trait-correlated immune module
exact
C4
Reported
3-gene logistic signature AUC = 0.87 (training, GSE5180)
Reproduced
0.872
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 97/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

This is a clean reproduction: on the public GSE5180 discovery set all four core numbers land 1:1 — 7 WGCNA modules at β=22, immune module-trait r=0.361 vs 0.36, all 7 hub genes in one module, and AUC 0.872 vs 0.87. The only deviations are on the input/preprocessing side and are cosmetic — a WGCNA color-label swap (the r≈0.36 immune module is blue in our run, brown in the paper) and a 262→248 brown-module gene count from the probe→symbol collapse. The color swap is itself an authors'-code inconsistency (WGCNA.txt narrates brown but exports the blue module), not a defect in our run, and no value shows fabrication or too-perfect signal. Severity negligible; central conclusion fully confirmed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

121.4 k
tokens (I/O) · 10.1 M incl. cache
22 min
runtime · 0.02 CPU-h
3.1 GB
peak RAM
1
HPC jobs
hummel
machine