Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comprehensive Analysis of Cell Population Dynamics and Related Core Genes During Vitiligo Development.

Front Genet · 2021
L1 67/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
67/100
Reproducibility score
0.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 29% of all assessed papers rank 795 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Partial reproduction (third-party tool valid per P16). PRIMARY DE result: 3 of 4 limma DEG counts reproduce within tolerance with near-exact up/down splits - GSE53146 771 vs 761 and GSE90880 94 vs 91 (raw P<0.05 & |FC|>=1.5, gene-level), GSE75819 2669 vs 2783 (matches under BH-adjusted P). GSE80009 does NOT reproduce (2073 vs 335; right down-dominant direction but ~6x too many; unusual OneArray series-matrix scale). DECONVOLUTION (EPIC, == immunedeconv method=epic): macrophage increase in vitiligo reproduces across ALL 4 datasets (skin+blood); the secondary claims (B/NK in skin, CD8+T in blood) are directionally inconsistent between the two datasets of each tissue and non-significant. DATA INTEGRITY notes for reviewer: (1) GSE90880 paper text swaps sample counts (says 6 vitiligo/8 healthy; GEO data is 8 vitiligo/6 control) yet DEG count matches -> real data used, reporting error not fabrication; (2) GSE80009 series is a superset (12 samples; paper used 8 non-segmental). All reported DE numbers are derivable from shipped GEO data (3/4 closely) -> no fabrication indicated. NOT attempted: GO/KEGG (discontinued KOBAS web tool), PCA, co-expression core genes. Grades PROVISIONAL; human reviewer decides.

💻 Code ↗ 🗄 Data: GSE53146

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ 1c22abfd5a4f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The pathogenesis of vitiligo involves a relationship between immune cell infiltration and differential gene expression, which can be characterized by combining bioinformatics analysis of published expression datasets with immune cell deconvolution and coexpression analysis.

Core claims
  • Immune cell infiltration and abnormal gene expression are closely related to vitiligo pathogenesis mechanism
  • Macrophages, B cells and NK cells are increased in lesional skin of vitiligo patients compared to healthy controls finding
  • CD8+ T cells and macrophages are significantly increased in peripheral blood of vitiligo patients compared to controls finding
  • Differentially expressed immune and inflammatory response genes show a strong positive correlation with macrophage abundance finding
  • TLR4 receptor pathway, interferon gamma-mediated signaling pathway and lipopolysaccharide-related pathway are positively correlated with CD4+ T cell abundance finding
  • Specific immune/inflammatory response genes (e.g., IFITM2, TNFSF10, GZMA, CEBPB, ADAM8, CXCR3, CXCL10) are associated with macrophage, CD4+ T cell, or CD8+ T cell abundance finding
  • Upregulated genes in vitiligo skin lesions and peripheral blood leukocytes are highly enriched in immune response and inflammatory response signaling pathways finding
  • EPIC deconvolution via immunedeconv software can estimate immune cell population proportions from bulk gene expression profiles of vitiligo samples method
Experimental setups
Assay System Perturbation Readout Platform
gene expression microarray (Illumina HumanHT-12 WG-DASL V4.0 R2 beadchip, GPL14951) human skin (epidermis and superficial dermis), GSE53146 vitiligo lesion vs healthy control differentially expressed genes Illumina HumanHT-12 WG-DASL V4.0 R2 expression beadchip
gene expression microarray (Illumina HumanWG-6 v3.0 beadchip, GPL6884) human skin epidermis, GSE75819 lesional vs non-lesional epidermis in vitiligo patients differentially expressed genes Illumina HumanWG-6 v3.0 expression beadchip
gene expression microarray (Affymetrix Human Genome U95 v2 array, GPL8300) human peripheral blood mononuclear cells (PBMCs), GSE90880 vitiligo vs healthy control differentially expressed genes Affymetrix Human Genome U95 version 2 array
gene expression microarray (Phalanx Human OneArray ver.6, GPL16951) human peripheral blood leukocytes (PBLs), GSE80009 vitiligo vs healthy control differentially expressed genes Phalanx Human OneArray ver. 6 release 1
immune cell deconvolution (EPIC method via immunedeconv) vitiligo and healthy skin and peripheral blood samples (all four datasets) vitiligo vs healthy control (and lesional vs non-lesional) estimated proportions of immune cell populations (macrophages, B cells, NK cells, CD4+/CD8+ T cells) immunedeconv software, EPIC method
coexpression analysis (Pearson correlation) vitiligo and healthy skin and peripheral blood samples none (correlative analysis) correlation between immune cell population fractions and DEG/pathway expression
GO and KEGG functional enrichment analysis DEG gene lists from skin and peripheral blood datasets none enriched biological processes and pathways KOBAS 2.0 server
Key results
  • 761 DEGs identified in GSE53146 skin dataset (433 upregulated, 328 downregulated)
  • 2,783 DEGs identified in GSE75819 skin dataset (1,594 upregulated, 1,189 downregulated)
  • 335 DEGs identified in GSE80009 peripheral blood dataset (95 upregulated, 240 downregulated)
  • 91 DEGs identified in GSE90880 peripheral blood dataset (1 upregulated, 90 downregulated)
  • Macrophage, B cell and NK cell populations increased in vitiligo skin vs healthy controls; no difference between lesional and non-lesional skin
  • CD8+ T cell and macrophage populations increased in peripheral blood of vitiligo patients vs controls, higher overall in PBLs than PBMCs
  • Upregulated skin DEGs enriched in inflammatory response, T cell costimulation, T cell activation, T cell receptor signaling and related immune pathways (GSE53146)
  • Downregulated PBL genes highly enriched in innate immune response and inflammatory response pathways (GSE90880), similar to skin findings
Key statistics
  • pvalue P<0.05 (threshold for DEG identification (LIMMA), combined with |log2FC|>=1.5 or <=2/3)
  • pvalue p<0.005 (cutoff for GO/KEGG functional enrichment significance)
  • count 761 DEGs (433 up, 328 down) (GSE53146 skin dataset (5 vitiligo, 5 healthy))
  • count 2,783 DEGs (1,594 up, 1,189 down) (GSE75819 skin dataset (lesional vs non-lesional epidermis, 15 patients))
  • count 335 DEGs (95 up, 240 down) (GSE80009 peripheral blood leukocyte dataset (4 vitiligo, 4 healthy))
  • count 91 DEGs (1 up, 90 down) (GSE90880 PBMC dataset (6 vitiligo, 8 healthy))
  • other |log2FC| >= 1.5 or <= 2/3 (fold-change threshold for DEG calling)
  • count ~9% (approximate proportion of vitiligo patients with a family history of the condition (cited from prior literature))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics study re-analyzed four publicly available Illumina/Affymetrix microarray datasets (GSE53146, GSE75819, GSE80009, GSE90880) from GEO to identify differentially expressed genes (DEGs) in vitiligo skin and peripheral blood versus healthy controls. DEGs were called with the LIMMA moderated t-test (p<0.05, |log2FC|≥1.5 or ≤2/3); GO/KEGG enrichment was tested by hypergeometric test with Benjamini-Hochberg FDR in KOBAS 2.0; immune cell fractions were estimated by EPIC deconvolution and compared between groups with Student's t-tests; and Pearson correlation coefficients were used to link cell fractions to DEG expression across samples.

Replicationbiological Sample sizeSample sizes stated per dataset; no formal power calculation or sample size justification described GroupsVitiligo patients vs healthy controls (primary, three datasets); lesional vs non-lesional epidermis within vitiligo patients (secondary, GSE75819) Pairingmixed Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
LIMMA moderated t-test (empirical Bayes linear model) Identification of DEGs in all four datasets (GSE53146, GSE75819, GSE80009, GSE90880) GSE53146: n=10 (5 vitiligo, 5 controls); GSE75819: n=15 (lesional vs non-lesional); GSE80009: n=8 (4+4); GSE90880: n=14 (6 vitiligo, 8 controls) not stated
Hypergeometric test with Benjamini-Hochberg FDR correction GO term and KEGG pathway enrichment analysis of DEG lists via KOBAS 2.0 server DEG set sizes per dataset (761 in GSE53146; 2,783 in GSE75819; 335 in GSE80009; 91 in GSE90880) not stated
Student's t-test (two-sample) Comparison of EPIC-estimated immune cell fractions between vitiligo and healthy control groups (Figure 3D scatter plot axis described as 'log p value using Students t-test') As per respective datasets (minimum n=8 for GSE80009) not stated
Pearson correlation coefficient Coexpression analysis between EPIC-estimated immune cell fractions and immune/inflammatory response DEG expression in three datasets Sample sizes of contributing datasets (n=8 to n=15) not stated
Principal component analysis (PCA) Visualization of overall gene expression patterns and cell-type composition in all four datasets; used for QC and sample grouping rather than inference As per respective datasets na
Approaches that could also have been used
  • DEGs were called using a raw p-value threshold (p<0.05) combined with a fold-change filter (|log2FC|≥1.5), without FDR correction applied to the DEG list itself
    Could also: Apply Benjamini-Hochberg FDR correction to the per-probe p-values from LIMMA and use an FDR q-value cutoff (e.g., q<0.05 or q<0.10) for DEG calling — With hundreds of thousands of probes tested simultaneously, FDR-adjusted q-values directly bound the expected proportion of false positives among declared DEGs; this approach is the standard complement to the enrichment-level B-H correction already applied downstream
  • Student's t-test was used to compare EPIC-estimated immune cell fractions between groups with sample sizes as small as n=4 per group
    Could also: Use the Mann-Whitney U (Wilcoxon rank-sum) test as a non-parametric alternative — With minimum group sizes of n=4, the normality assumption underlying the t-test cannot be robustly assessed; rank-based tests make no distributional assumption and are widely used for small-sample comparisons of deconvolved cell fractions, which are bounded and often skewed
  • Pearson correlation was used to quantify coexpression between immune cell fractions and gene expression across samples
    Could also: Use Spearman rank correlation as an alternative — Spearman correlation does not assume bivariate normality and is less sensitive to outliers, which is relevant given small sample sizes (n=8–15) and the bounded, compositional nature of EPIC-estimated cell fractions
  • The four datasets were analyzed independently and results described narratively side-by-side
    Could also: Apply a fixed- or random-effects meta-analysis to pool log fold-change estimates (and their standard errors from LIMMA) across datasets for each gene — A formal meta-analysis yields a single weighted effect-size estimate with a confidence interval and a heterogeneity statistic (I²), enabling a more rigorous distinction between findings that replicate consistently across datasets and those that are dataset-specific
  • Multiple immune cell types were compared between groups per dataset without correction for the family of cell-type tests
    Could also: Apply a multiple-testing correction (e.g., Bonferroni or Benjamini-Hochberg) across the set of cell types evaluated within each dataset — Testing six or more cell types simultaneously increases the probability of at least one false positive; a correction across the cell-type family per dataset would formally account for this, consistent with the correction already applied in the enrichment analyses
  • Immune cell deconvolution relied solely on EPIC, selected after informal comparison with other methods
    Could also: Report results from two or more deconvolution methods (e.g., CIBERSORT, MCP-counter) in parallel and quantify concordance — Cell-fraction estimates can vary between algorithms due to differences in reference signatures and mathematical assumptions; reporting inter-method concordance allows readers to assess how much the biological conclusions depend on the choice of deconvolution tool
Software: R/affy · R/LIMMA · R/ggplot2 · R/pheatmap · R/factoextra · KOBAS 2.0 · immunedeconv/EPIC · Reactome

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

deg_GSE53146
Reported
761 DEGs (433 up / 328 down)
Reproduced
771 (423 up / 348 down)
within tolerance
deg_GSE75819
Reported
2783 DEGs (1594 up / 1189 down)
Reproduced
2669 (1603 up / 1066 down) under BH-adjusted P
within tolerance
deg_GSE80009
Reported
335 DEGs (95 up / 240 down)
Reproduced
2073 (564 up / 1509 down); no standard threshold lands near 335
did not match
deg_GSE90880
Reported
91 DEGs (1 up / 90 down)
Reproduced
94 (2 up / 92 down)
within tolerance
immune_skin_macrophage
Reported
macrophages increased in vitiligo skin
Reproduced
UP in both GSE53146(+0.0026) and GSE75819(+0.0001)
within tolerance
immune_skin_Bcell_NK
Reported
B cells and NK cells increased in vitiligo skin
Reproduced
inconsistent: B up GSE75819/down GSE53146; NK up GSE53146/flat GSE75819
partial
immune_blood_macrophage
Reported
macrophages increased in vitiligo blood
Reproduced
UP in both GSE80009(+0.0094) and GSE90880(+0.0024)
within tolerance
immune_blood_CD8T
Reported
CD8+ T cells increased in vitiligo blood
Reproduced
inconsistent: up GSE90880(+0.0244)/down GSE80009(-0.0163)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 67/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

267.2 k
tokens (I/O) · 13.1 M incl. cache
57 min
runtime · 0.07 CPU-h
1.9 GB
peak RAM
4
HPC jobs
hummel
machine