Comprehensive Analysis of Cell Population Dynamics and Related Core Genes During Vitiligo Development.
The main results reproduced, with only marginal, non-material deviations.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Partial reproduction (third-party tool valid per P16). PRIMARY DE result: 3 of 4 limma DEG counts reproduce within tolerance with near-exact up/down splits - GSE53146 771 vs 761 and GSE90880 94 vs 91 (raw P<0.05 & |FC|>=1.5, gene-level), GSE75819 2669 vs 2783 (matches under BH-adjusted P). GSE80009 does NOT reproduce (2073 vs 335; right down-dominant direction but ~6x too many; unusual OneArray series-matrix scale). DECONVOLUTION (EPIC, == immunedeconv method=epic): macrophage increase in vitiligo reproduces across ALL 4 datasets (skin+blood); the secondary claims (B/NK in skin, CD8+T in blood) are directionally inconsistent between the two datasets of each tissue and non-significant. DATA INTEGRITY notes for reviewer: (1) GSE90880 paper text swaps sample counts (says 6 vitiligo/8 healthy; GEO data is 8 vitiligo/6 control) yet DEG count matches -> real data used, reporting error not fabrication; (2) GSE80009 series is a superset (12 samples; paper used 8 non-segmental). All reported DE numbers are derivable from shipped GEO data (3/4 closely) -> no fabrication indicated. NOT attempted: GO/KEGG (discontinued KOBAS web tool), PCA, co-expression core genes. Grades PROVISIONAL; human reviewer decides.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-18 ⛓ 1c22abfd5a4f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe pathogenesis of vitiligo involves a relationship between immune cell infiltration and differential gene expression, which can be characterized by combining bioinformatics analysis of published expression datasets with immune cell deconvolution and coexpression analysis.
- ★ Immune cell infiltration and abnormal gene expression are closely related to vitiligo pathogenesis mechanism
- ★ Macrophages, B cells and NK cells are increased in lesional skin of vitiligo patients compared to healthy controls finding
- ★ CD8+ T cells and macrophages are significantly increased in peripheral blood of vitiligo patients compared to controls finding
- ★ Differentially expressed immune and inflammatory response genes show a strong positive correlation with macrophage abundance finding
- ★ TLR4 receptor pathway, interferon gamma-mediated signaling pathway and lipopolysaccharide-related pathway are positively correlated with CD4+ T cell abundance finding
- ★ Specific immune/inflammatory response genes (e.g., IFITM2, TNFSF10, GZMA, CEBPB, ADAM8, CXCR3, CXCL10) are associated with macrophage, CD4+ T cell, or CD8+ T cell abundance finding
- ★ Upregulated genes in vitiligo skin lesions and peripheral blood leukocytes are highly enriched in immune response and inflammatory response signaling pathways finding
- EPIC deconvolution via immunedeconv software can estimate immune cell population proportions from bulk gene expression profiles of vitiligo samples method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| gene expression microarray (Illumina HumanHT-12 WG-DASL V4.0 R2 beadchip, GPL14951) | human skin (epidermis and superficial dermis), GSE53146 | vitiligo lesion vs healthy control | differentially expressed genes | Illumina HumanHT-12 WG-DASL V4.0 R2 expression beadchip |
| gene expression microarray (Illumina HumanWG-6 v3.0 beadchip, GPL6884) | human skin epidermis, GSE75819 | lesional vs non-lesional epidermis in vitiligo patients | differentially expressed genes | Illumina HumanWG-6 v3.0 expression beadchip |
| gene expression microarray (Affymetrix Human Genome U95 v2 array, GPL8300) | human peripheral blood mononuclear cells (PBMCs), GSE90880 | vitiligo vs healthy control | differentially expressed genes | Affymetrix Human Genome U95 version 2 array |
| gene expression microarray (Phalanx Human OneArray ver.6, GPL16951) | human peripheral blood leukocytes (PBLs), GSE80009 | vitiligo vs healthy control | differentially expressed genes | Phalanx Human OneArray ver. 6 release 1 |
| immune cell deconvolution (EPIC method via immunedeconv) | vitiligo and healthy skin and peripheral blood samples (all four datasets) | vitiligo vs healthy control (and lesional vs non-lesional) | estimated proportions of immune cell populations (macrophages, B cells, NK cells, CD4+/CD8+ T cells) | immunedeconv software, EPIC method |
| coexpression analysis (Pearson correlation) | vitiligo and healthy skin and peripheral blood samples | none (correlative analysis) | correlation between immune cell population fractions and DEG/pathway expression | — |
| GO and KEGG functional enrichment analysis | DEG gene lists from skin and peripheral blood datasets | none | enriched biological processes and pathways | KOBAS 2.0 server |
- – 761 DEGs identified in GSE53146 skin dataset (433 upregulated, 328 downregulated)
- – 2,783 DEGs identified in GSE75819 skin dataset (1,594 upregulated, 1,189 downregulated)
- – 335 DEGs identified in GSE80009 peripheral blood dataset (95 upregulated, 240 downregulated)
- – 91 DEGs identified in GSE90880 peripheral blood dataset (1 upregulated, 90 downregulated)
- ▲ Macrophage, B cell and NK cell populations increased in vitiligo skin vs healthy controls; no difference between lesional and non-lesional skin
- ▲ CD8+ T cell and macrophage populations increased in peripheral blood of vitiligo patients vs controls, higher overall in PBLs than PBMCs
- ▲ Upregulated skin DEGs enriched in inflammatory response, T cell costimulation, T cell activation, T cell receptor signaling and related immune pathways (GSE53146)
- ▼ Downregulated PBL genes highly enriched in innate immune response and inflammatory response pathways (GSE90880), similar to skin findings
- pvalue P<0.05 (threshold for DEG identification (LIMMA), combined with |log2FC|>=1.5 or <=2/3)
- pvalue p<0.005 (cutoff for GO/KEGG functional enrichment significance)
- count 761 DEGs (433 up, 328 down) (GSE53146 skin dataset (5 vitiligo, 5 healthy))
- count 2,783 DEGs (1,594 up, 1,189 down) (GSE75819 skin dataset (lesional vs non-lesional epidermis, 15 patients))
- count 335 DEGs (95 up, 240 down) (GSE80009 peripheral blood leukocyte dataset (4 vitiligo, 4 healthy))
- count 91 DEGs (1 up, 90 down) (GSE90880 PBMC dataset (6 vitiligo, 8 healthy))
- other |log2FC| >= 1.5 or <= 2/3 (fold-change threshold for DEG calling)
- count ~9% (approximate proportion of vitiligo patients with a family history of the condition (cited from prior literature))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This bioinformatics study re-analyzed four publicly available Illumina/Affymetrix microarray datasets (GSE53146, GSE75819, GSE80009, GSE90880) from GEO to identify differentially expressed genes (DEGs) in vitiligo skin and peripheral blood versus healthy controls. DEGs were called with the LIMMA moderated t-test (p<0.05, |log2FC|≥1.5 or ≤2/3); GO/KEGG enrichment was tested by hypergeometric test with Benjamini-Hochberg FDR in KOBAS 2.0; immune cell fractions were estimated by EPIC deconvolution and compared between groups with Student's t-tests; and Pearson correlation coefficients were used to link cell fractions to DEG expression across samples.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| LIMMA moderated t-test (empirical Bayes linear model) | Identification of DEGs in all four datasets (GSE53146, GSE75819, GSE80009, GSE90880) | GSE53146: n=10 (5 vitiligo, 5 controls); GSE75819: n=15 (lesional vs non-lesional); GSE80009: n=8 (4+4); GSE90880: n=14 (6 vitiligo, 8 controls) | not stated |
| Hypergeometric test with Benjamini-Hochberg FDR correction | GO term and KEGG pathway enrichment analysis of DEG lists via KOBAS 2.0 server | DEG set sizes per dataset (761 in GSE53146; 2,783 in GSE75819; 335 in GSE80009; 91 in GSE90880) | not stated |
| Student's t-test (two-sample) | Comparison of EPIC-estimated immune cell fractions between vitiligo and healthy control groups (Figure 3D scatter plot axis described as 'log p value using Students t-test') | As per respective datasets (minimum n=8 for GSE80009) | not stated |
| Pearson correlation coefficient | Coexpression analysis between EPIC-estimated immune cell fractions and immune/inflammatory response DEG expression in three datasets | Sample sizes of contributing datasets (n=8 to n=15) | not stated |
| Principal component analysis (PCA) | Visualization of overall gene expression patterns and cell-type composition in all four datasets; used for QC and sample grouping rather than inference | As per respective datasets | na |
-
DEGs were called using a raw p-value threshold (p<0.05) combined with a fold-change filter (|log2FC|≥1.5), without FDR correction applied to the DEG list itself↳ Could also: Apply Benjamini-Hochberg FDR correction to the per-probe p-values from LIMMA and use an FDR q-value cutoff (e.g., q<0.05 or q<0.10) for DEG calling — With hundreds of thousands of probes tested simultaneously, FDR-adjusted q-values directly bound the expected proportion of false positives among declared DEGs; this approach is the standard complement to the enrichment-level B-H correction already applied downstream
-
Student's t-test was used to compare EPIC-estimated immune cell fractions between groups with sample sizes as small as n=4 per group↳ Could also: Use the Mann-Whitney U (Wilcoxon rank-sum) test as a non-parametric alternative — With minimum group sizes of n=4, the normality assumption underlying the t-test cannot be robustly assessed; rank-based tests make no distributional assumption and are widely used for small-sample comparisons of deconvolved cell fractions, which are bounded and often skewed
-
Pearson correlation was used to quantify coexpression between immune cell fractions and gene expression across samples↳ Could also: Use Spearman rank correlation as an alternative — Spearman correlation does not assume bivariate normality and is less sensitive to outliers, which is relevant given small sample sizes (n=8–15) and the bounded, compositional nature of EPIC-estimated cell fractions
-
The four datasets were analyzed independently and results described narratively side-by-side↳ Could also: Apply a fixed- or random-effects meta-analysis to pool log fold-change estimates (and their standard errors from LIMMA) across datasets for each gene — A formal meta-analysis yields a single weighted effect-size estimate with a confidence interval and a heterogeneity statistic (I²), enabling a more rigorous distinction between findings that replicate consistently across datasets and those that are dataset-specific
-
Multiple immune cell types were compared between groups per dataset without correction for the family of cell-type tests↳ Could also: Apply a multiple-testing correction (e.g., Bonferroni or Benjamini-Hochberg) across the set of cell types evaluated within each dataset — Testing six or more cell types simultaneously increases the probability of at least one false positive; a correction across the cell-type family per dataset would formally account for this, consistent with the correction already applied in the enrichment analyses
-
Immune cell deconvolution relied solely on EPIC, selected after informal comparison with other methods↳ Could also: Report results from two or more deconvolution methods (e.g., CIBERSORT, MCP-counter) in parallel and quantify concordance — Cell-fraction estimates can vary between algorithms due to differences in reference signatures and mathematical assumptions; reporting inter-method concordance allows readers to assess how much the biological conclusions depend on the choice of deconvolution tool
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.