Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identification of the shared hub gene signatures and molecular mechanisms between HIV-1 and pulmonary arterial hypertension.

Sci Rep · 2024
L1 71/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the core, low-hanging pipeline outputs 1:1. The registry code (junjunlab/GseaVis) is a third-party GSEA plotting package, NOT the authors' analysis pipeline, so per P16 we ran the described limma DEG pipeline on the paper's own public GEO data (GSE77939 HIV, GSE703 PAH). Results: IFI27 diagnostic ROC-AUC in GSE703 reproduced EXACTLY (0.940 vs 0.940); the two reported DEG counts reproduced within 7-11% with matching up/down direction (GSE77939 68 genes vs 61; GSE703 557 genes vs 519) using standard microarray limma (paper said 'voom', which is for RNA-seq counts, not these microarrays); ISG15 & IFI27 confirmed as up-regulated DEGs in both datasets. Overall: partial reproduction, NO fabrication concern — every value is derivable from the shipped public data and the independent exact AUC match strongly corroborates the paper. NOT attempted (the hard ~20%): WGCNA modules (11/21) & the 109 shared genes (many unstated params), ClueGO GO/KEGG (interactive Cytoscape GUI), GSEA top-5 IFN pathways (needs exact ranked list + MSigDB v7.4), single-cell 50,524-cell clustering, CIBERSORT, GeneMANIA network.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 71
    assessed: 2026-06-15 ⛓ 0e8abf5ebfc0
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What shared molecular mechanisms and hub genes link HIV-1 infection to pulmonary arterial hypertension (PAH)? The authors hypothesize that a heightened type I interferon response in HIV-1 is a crucial susceptibility factor for PAH.

Core claims
  • HIV-1 and PAH share 109 co-expressed genes primarily enriched in type I interferon (IFN) pathways. finding
  • ISG15 and IFI27 are pivotal shared hub genes between HIV-1 and PAH. finding
  • A heightened type I IFN response in HIV-1 may be a crucial susceptibility factor for PAH. mechanism
  • Monocytes are pivotal cells involved in the shared type I IFN response pathway linking HIV-1 and PAH. mechanism
  • WGCNA combined with DEG intersection and external validation can identify shared disease hub genes. method
  • Shared genes are enriched in NOD-like and RIG-I-like receptor signaling and defense response to virus. finding
  • ISG15 and IFI27 show diagnostic potential (ROC/AUC) for HIV-1 and PAH across datasets. resource
Experimental setups
Assay System Perturbation Readout Platform
Gene expression microarray (WGCNA co-expression analysis) Human PBMC (GSE140713 HIV; GSE33463 PAH) disease vs control (HIV-1 / PAH) module eigengene–trait correlation, shared module genes GPL6480 (Agilent); GPL6947 (Illumina)
Microarray differential expression analysis (limma/voom) Human PBMC (GSE77939 HIV; GSE703 PAH) disease vs control differentially expressed genes (DEGs) GPL15207; GPL80
Microarray hub gene expression / ROC diagnostic validation Human PBMC (GSE140713, GSE77939, GSE2171, GSE33463, GSE703, GSE19617) disease vs control (incl. ART and IPAH/SPAH subgroups) hub gene expression levels (t-test) and AUC GPL6480, GPL15207, GPL201, GPL6947, GPL80
GO/KEGG and PPI enrichment analysis 109 shared genes (in silico) none enriched GO-BP terms and KEGG pathways; PPI network ClueGO/Cytoscape; GeneMania; Metascape
Gene set enrichment analysis (GSEA) Human PBMC (GSE140713 HIV; GSE33463 PAH) ISG15/IFI27 high vs low expression groups hallmark pathway normalized enrichment score (NES) clusterProfiler / MSigDB v7.4
Immune cell infiltration analysis (CIBERSORT) Human PBMC (GSE140713 HIV; GSE33463 PAH) disease vs control relative abundance of 22 immune cell types; hub gene–cell correlation CIBERSORT R package
Single-cell RNA-seq (scRNA-seq) Human PBMC (10X 8381 PBMCs; GSE157829; GSM6647828) healthy vs HIV-1 infected donors cell cluster identification and hub gene/pathway signature expression 10X Genomics Chromium 3'; Seurat/AUCell
Key results
  • 109 shared genes overlap between HIV-1 and PAH positively correlated modules 109 genes
  • ISG15 and IFI27 identified as hub genes by intersecting upregulated DEGs and WGCNA shared genes
  • Shared genes enriched in type I IFN / defense response to virus / toll-like receptor signaling
  • GSE140713 green module strongly positively correlated with HIV-1 r=0.65, P<0.001
  • GSE33463 turquoise module strongly positively correlated with PAH r=0.74, P<0.001
  • GSE77939 yielded 61 DEGs (40 up, 21 down) in HIV-1 61 DEGs
  • GSE703 yielded 519 DEGs (389 up, 130 down) in PAH 519 DEGs
  • Monocytes implicated as pivotal cells in type I IFN response by CIBERSORT and scRNA-seq
Key statistics
  • correlation r=0.74, P<0.001 (GSE33463 turquoise module vs PAH)
  • correlation r=0.65, P<0.001 (GSE140713 green module vs HIV-1)
  • correlation r=0.59, P<0.001 (GSE33463 blue module vs PAH)
  • correlation r=0.56, P<0.001 (GSE140713 blue module vs HIV-1)
  • correlation r=0.50, P<0.001 (GSE33463 greenyellow module vs PAH)
  • count 109 shared genes (overlap of HIV-1 and PAH modules)
  • count 519 DEGs (389 up, 130 down) (GSE703 PAH=14, control=6)
  • fold_change |log2FC|>0.585, P<0.05 (DEG screening criteria for GSE77939 and GSE703)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This computational bioinformatics study used six public PBMC microarray datasets to identify shared transcriptomic signatures between HIV-1 infection and pulmonary arterial hypertension (PAH). WGCNA was applied to the two largest datasets for co-expression module discovery, followed by limma/voom differential expression analysis on validation datasets to narrow hub gene candidates; results were further evaluated with ROC/AUC, GSEA (Benjamini-Hochberg corrected), CIBERSORT immune deconvolution, and scRNA-seq clustering. Group-level comparisons used Student's t-test and Wilcoxon rank sum test, and inter-variable associations used Spearman correlation.

Replicationbiological Sample sizeSample sizes reported per dataset in Table 1; no formal a priori power calculation described GroupsHIV-1 patients vs. healthy controls; PAH patients vs. healthy controls; subgroup analyses by ART status (ART-IF, ART-N, ART-R) and PAH subtype (IPAH, SPAH) Pairingunpaired Randomization/blindingnot stated Dispersionnone Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR correction
Statistical tests used
Test Applied to n Assumptions
WGCNA with Pearson correlation (network construction) and Spearman correlation (module-trait association) Co-expression network construction and module-phenotype correlation in GSE140713 (HIV) and GSE33463 (PAH) GSE140713: 50 HIV + 7 controls; GSE33463: 72 PAH + 41 controls not stated
limma/voom linear model with empirical Bayes moderation Differential expression analysis in validation datasets GSE77939 (HIV) and GSE703 (PAH); threshold |log2FC| > 0.585 and P < 0.05 GSE77939: 17 HIV + 4 controls; GSE703: 14 PAH + 6 controls not stated
Student's t-test (unpaired, two-group) Hub gene (ISG15, IFI27) expression comparison between control and HIV-1/PAH groups across all six datasets and ART/PAH-subtype subgroup analyses varies by dataset; see Table 1 (control n ranges from 4 to 41) not stated
Wilcoxon rank sum test Immune cell infiltration abundance (CIBERSORT output) comparison between control and HIV-1/PAH groups in GSE140713 and GSE33463 GSE140713: 50 HIV + 7 controls; GSE33463: 72 PAH + 41 controls not stated
Spearman correlation Correlation between hub gene expression and relative immune cell type abundance (CIBERSORT); significance threshold |cor| > 0.3 and P < 0.05 null not stated
ROC curve / AUC (via pROC package) Biomarker diagnostic performance of ISG15 and IFI27 evaluated across all six datasets varies by dataset; see Table 1 na
GSEA (clusterProfiler, MSigDB v7.4 hallmark gene sets) with Benjamini-Hochberg FDR adjustment Pathway enrichment in high- vs. low-ISG15/IFI27 expression subgroups within GSE140713 and GSE33463; significance threshold adjusted P < 0.05 null not stated
Approaches that could also have been used
  • Student's t-test was used to compare hub gene expression between cases and controls in datasets where control group sizes are small (n = 4 or n = 7)
    Could also: Wilcoxon rank sum test (Mann-Whitney U) or a permutation-based test — Non-parametric alternatives make fewer distributional assumptions and are often preferred when sample sizes are too small to verify normality reliably; the study already applied the Wilcoxon test for CIBERSORT comparisons, so the approach is established within the same workflow
  • Multiple hub-gene expression comparisons were performed across six datasets and multiple subgroups using uncorrected P < 0.05
    Could also: Bonferroni, Holm, or Benjamini-Hochberg correction applied across the family of hub-gene × dataset comparisons — An explicit correction over this comparison family would bound the type I error rate; the study already applies BH correction within GSEA, so extension to other comparison families would be consistent with that choice
  • WGCNA was applied to only two of the six available datasets (one per disease, selected for largest n) for module discovery
    Could also: Consensus WGCNA across all same-disease datasets after ComBat or limma-based batch correction — Consensus WGCNA is specifically designed to find modules reproducible across multiple datasets and can reduce dataset-specific noise; using all available datasets could increase the robustness of identified modules
  • Immune cell deconvolution was performed with CIBERSORT alone
    Could also: xCell, TIMER2.0, or MCP-counter applied in parallel or as a sensitivity check — Different deconvolution algorithms use distinct reference signatures and estimation methods; concordance of cell-type proportion estimates across multiple tools is commonly used to increase confidence in findings
  • Group comparison figures used threshold-based significance symbols (* P < 0.05, ** P < 0.01) rather than exact p-values for hub gene and immune cell comparisons
    Could also: Report exact p-values alongside effect size estimates (e.g., log2FC, rank-biserial correlation) and 95% confidence intervals — Exact p-values and effect sizes with confidence intervals allow readers to assess the precision and practical magnitude of differences and are recommended by many statistical reporting guidelines (e.g., APA, ICMJE)
  • Measures of dispersion were not reported for hub gene expression group comparisons
    Could also: Report SD, SEM, or 95% CI alongside group means or medians, particularly in box or violin plots — Displaying spread is especially informative when control group sizes are small (n = 4–12 across datasets), helping readers gauge the reliability of central tendency estimates and the degree of overlap between groups
Software: R/RStudio 4.2.1 · WGCNA (R package) · limma (R package) · affy (R package) · pROC (R package) · clusterProfiler (R package) · GseaVis (R package) · Seurat (R package) 4.3.0 · DoubletFinder (R package) 2.0.3 · AUCell (R/Bioconductor package) 1.12.0 · CIBERSORT (R package) · ClueGO (Cytoscape plugin) 2.5.9 · Cytoscape 3.9.1 · Metascape (online platform)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
4
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GPL201 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
GPL80 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
GPL15207 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE140713 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE157829 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE19617 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE2171 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE33463 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE703 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE77939 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM6647828 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38528047

Paper: Mai et al. 2024, Sci Rep 14:7048. "Identification of the shared hub gene signatures and molecular mechanisms between HIV-1 and pulmonary arterial hypertension." DOI 10.1038/s41598-024-55645-x · PMCID PMC10963360.

Registry code: https://github.com/junjunlab/GseaVis — this is a third-party GSEA visualization R package, NOT the authors' analysis pipeline. The authors ship no analysis code. Per BRIEF rule 2 (P16), reproducing by running the standard pipeline they describe on their own public GEO data is equally valid. We do that.

Pipeline-derived results in the paper (candidate claims)

# Reported result Pipeline In scope?
C1 GSE77939 (HIV): 61 DEGs = 40 up + 21 down at |log2FC|>0.585 & P<0.05, limma limma DEG YES — exact method + threshold + dataset stated
C2 GSE703 (PAH): 519 DEGs = 389 up + 130 down, same threshold/method limma DEG YES — exact method + threshold + dataset stated
C3 IFI27 diagnostic AUC = 0.940 in GSE703 pROC ROC-AUC YES (stretch) — single clean number on a public dataset
C4 ISG15 & IFI27 are the 2 hub genes (up-regulated, IFN-related) sanity check partial — verify they appear up-regulated in the DEG lists
109 shared WGCNA genes; module counts (11 / 21); β=14/6 WGCNA OUT — many unstated params (cut height, merge, module-trait), high variance, the hard 20%
GO/KEGG terms (ClueGO Cytoscape plugin) manual GUI OUT — interactive plugin, not scriptable 1:1
GSEA top-5 IFN pathways (GseaVis plots) clusterProfiler OUT — depends on upstream ranked list + MSigDB v7.4 exactly
scRNA: 50,524 cells → 15 clusters → 6 populations; CIBERSORT scanpy/Seurat OUT — separate single-cell dataset, heavy, many unstated params
GeneMANIA associated proteins web tool OUT — external web service, not a reproducible compute step

What we attempt (the 80, not the 20)

Re-run the explicitly-specified limma DEG analysis on the two datasets where the paper gives both the method and the exact threshold and an exact DEG count (C1, C2), plus the one clean diagnostic number (C3) and a hub-gene sanity check (C4). These are the low-hanging, clearly-specified, 1:1-checkable outputs.

Method faithfulness notes / caveats (recorded up front, not hidden)

  • Paper says "voom method via limma". voom is for RNA-seq counts; GSE77939 (GPL15207) and GSE703 (GPL80) are microarrays with log-intensity values, so voom is not strictly applicable. We run the standard microarray limma pipeline (lmFit + eBayes on log2 intensities) — the correct method for these data — and flag the voom mention as a possible methods inconsistency.
  • "P < 0.05" is taken as the raw p-value (as printed). We also report the adjusted-p count for context.
  • DEG counts can be reported at probe vs gene level; we report both (probe count and unique-gene-symbol count) so the human auditor can see which the paper's number matches.

All compute on «our HPC» (SLURM). Data fetched via GEOquery inside the compute job; large intermediates stay on «infra»; only small result JSONs come back.

Figures / tables: Fig.2Fig.6
C1
Reported
GSE77939 (HIV) DEGs: 61 total (40 up, 21 down)
Reproduced
68 genes / 75 probes (42 up, 33 down)
partial
C2
Reported
GSE703 (PAH) DEGs: 519 total (389 up, 130 down)
Reproduced
557 genes / 624 probes (450 up, 174 down)
partial
C3
Reported
IFI27 ROC-AUC in GSE703 = 0.940
Reproduced
0.940
exact
C4
Reported
ISG15 & IFI27 are the 2 hub genes, up-regulated
Reproduced
both up-regulated DEGs in both GSE77939 and GSE703
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

This is a solid partial reproduction with no fabrication concern: all reported values are derivable from the paper's own public GEO data, and the IFI27 GSE703 ROC-AUC reproduced exactly (0.940) while both DEG counts came within 7–11% with matching up/down direction. The deviations sit on the input/preprocessing side and stem from our self-chosen steps under an underspecified method (paper's 'voom' is wrong for microarrays; normalization and probe-collapse unspecified) — not from the authors' side and not from data unavailability. The authors' actual analysis code was never shipped (the linked repo is a generic plotting package), and the hard ~20% (WGCNA, GSEA, single-cell, CIBERSORT) was not attempted, so overall quality is yellow — confirmed core, explainable gaps, incomplete scope.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

116.1 k
tokens (I/O) · 9.7 M incl. cache
15 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine