Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identification of TYR, TYRP1, DCT and LARP7 as related biomarkers and immune infiltration characteristics of vitiligo via comprehensive strategies.

Bioengineered · 2021
L1 55/100 PQI 85
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
55/100
Reproducibility score
1.1 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 15% of all assessed papers rank 986 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described-well-enough? PARTIALLY. The registry code link (ggstatsplot) is a generic plotting package, not the authors' analysis code; the paper ships no code, so we ran the described pipeline (GEOquery + limma, p<0.05 & |log2FC|>1) on the paper's public GEO data on «our HPC». Outcome: NOT a clean 1:1, but a faithful PARTIAL reproduction. The pipeline robustly reproduces the SCALE and DIRECTION of every per-dataset DEG result: GSE75819 within ~5% (1118 vs 1064 reported), GSE53146 ~18% high (1000 vs 847), GSE65127 predominantly-downregulated signal reproduced with down-count near-exact (67 vs 64). Exact counts are NOT 1:1 because the paper leaves pivotal parameters unspecified: GSE65127 has 4 site types x 10 samples (healthy/lesional/non-lesional/peri-lesional) and the paper never states which form the 'normal' control, so the reported total 73 sits BETWEEN our two defensible choices (67 lesional-vs-all-other-sites, 171 lesional-vs-healthy); probe->gene collapsing rule and raw-vs-adjusted p are also unstated. RRA gives 96 vs 131 (cascades from the per-dataset lists). No fabrication concern: every reported value is plausible and lies within the range produced by reasonable parameter choices. NOT ATTEMPTED (the hard ~20%, deliberately, per 80/20): WGCNA module detection, the LASSO+SVM-RFE+RF+WGCNA biomarker overlap yielding TYR/TYRP1/DCT/LARP7, CIBERSORT immune infiltration, and the GSE90880 validation ROC (AUC=0.942) — these are stochastic and/or depend on unspecified seeds and inputs.

💻 Code ↗ 🗄 Data: GSE65127

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 55
    assessed: 2026-06-15 ⛓ 6898b48d6af3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a comprehensive computational strategy combining robust rank aggregation, WGCNA, and machine learning algorithms identify reliable gene biomarkers of vitiligo and characterize the immune cell infiltration underlying its pathogenesis?

Core claims
  • TYR, TYRP1, DCT and LARP7 are biomarkers associated with vitiligo finding
  • 131 robust DEGs distinguish vitiligo lesional from normal skin and are enriched in pigmentation/melanogenesis and immune pathways finding
  • Immune cell infiltration (CD4 T, CD8 T, Tregs, NK cells, dendritic cells, macrophages) is involved in vitiligo pathogenesis finding
  • A combined strategy of RRA, WGCNA, LASSO, SVM-RFE and RF can screen disease biomarkers method
  • CIBERSORT was used for the first time to characterize 22 immune cell subsets in vitiligo tissue method
  • The four-biomarker panel discriminates vitiligo in an independent validation set finding
Experimental setups
Assay System Perturbation Readout Platform
Gene expression microarray analysis (DEG identification via limma) Human vitiligo lesional vs normal skin (GSE53146, GSE65127, GSE75819) none (disease vs normal observational) Differentially expressed genes (|log2FC|>1, p<0.05) GEO microarray datasets; RMA normalization
Robust rank aggregation (RRA) meta-analysis Merged human vitiligo skin datasets none Robust DEGs (FC>1, p<0.05) RRA R package
GO/KEGG/GSEA functional enrichment Robust DEG gene set none Enriched biological pathways/terms (FDR<0.25, p<0.05) clusterProfiler R package; c2.cp.kegg.v7.2.symbols.gmt
WGCNA + machine learning (LASSO, SVM-RFE, RF) biomarker screening Merged human vitiligo skin expression matrix none Candidate biomarker genes, hub module genes WGCNA, e1071, randomForest R packages
ROC/AUC validation GSE90880 verification dataset none Diagnostic AUC of combined biomarkers pROC R package
CIBERSORT immune deconvolution + correlation analysis Human vitiligo lesional vs normal skin none 22 immune cell fractions and Spearman correlation with biomarkers CIBERSORT; corrplot, ggstatsplot, ggplot2 R packages
Key results
  • 131 robust DEGs identified (89 upregulated, 42 downregulated) 131 genes
  • Four-biomarker panel (TYR, TYRP1, DCT, LARP7) discriminates vitiligo in validation set AUC=0.942
  • TYR positively correlated with activated dendritic cells r=0.644, p<0.01
  • TYR negatively correlated with macrophages M2 r=-0.387, p<0.01
  • DCT positively correlated with Tregs r=0.316, p=0.02
  • LARP7 negatively correlated with macrophages M2 r=-0.398, p<0.01
  • GSE53146 yielded 847 DEGs; GSE65127 73 DEGs; GSE75819 1064 DEGs 847/73/1064 genes
  • Enrichment in melanogenesis, oxidative phosphorylation, cell cycle, tyrosine metabolism, proteasome pathways
Key statistics
  • other AUC = 0.942 (ROC validation of combined 4 biomarkers in GSE90880)
  • correlation r = 0.644, p < 0.01 (TYR vs dendritic cells activated)
  • correlation r = -0.387, p < 0.01 (TYR vs macrophages M2)
  • correlation r = 0.316, p = 0.02 (DCT vs Tregs)
  • correlation r = 0.354, p < 0.01 (LARP7 vs macrophages M1)
  • correlation r = -0.398, p < 0.01 (LARP7 vs macrophages M2)
  • count 131 robust DEGs (89 up, 42 down) (RRA-integrated DEGs across three datasets)
  • count 22 immune cell subsets (CIBERSORT immune infiltration analysis)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a bioinformatics/re-analysis study of three public vitiligo gene-expression microarray datasets (GSE53146, GSE65127, GSE75819) with one validation set (GSE90880). Differential expression was computed per dataset with limma and integrated across datasets via robust rank aggregation (RRA); candidate biomarkers were screened by intersecting WGCNA modules with three machine-learning feature-selection methods (LASSO, SVM-RFE, random forest) and evaluated by ROC/AUC, and immune cell infiltration was estimated with CIBERSORT and related to biomarkers via Spearman correlation. Results are reported mainly as gene counts, fold-change/p-value thresholds, AUC, and correlation coefficients with p-values rather than as effect estimates with confidence intervals.

Replicationbiological Sample sizeper-dataset sample counts given (5/5, 10/10, 15/15; validation GSE90880); no power/sample-size calculation described Groupslesional vs normal skin Pairingunclear Randomization/blindingna Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionFDR (Benjamini-Hochberg-type) used for GSEA (FDR < 0.25); limma typically reports adjusted p-values though correction is not explicitly stated for the DEG p < 0.05 cutoff
Statistical tests used
Test Applied to n Assumptions
limma differential expression (moderated statistics) DEGs in each microarray dataset (GSE53146, GSE65127, GSE75819) GSE53146: 5 lesional/5 normal; GSE65127: 10/10; GSE75819: 15/15 not stated
Robust rank aggregation (RRA) integration of ranked DEG lists across the three datasets to obtain 131 robust DEGs three datasets not stated
GO/KEGG over-representation and GSEA (clusterProfiler) functional enrichment of robust DEGs 131 robust DEGs na
LASSO logistic regression feature selection of vitiligo-related genes (14 genes) not stated
SVM-RFE (e1071) feature selection of vitiligo-related genes (98 genes) not stated
Random forest (decision-tree ensemble) feature selection of vitiligo-related genes (12 genes) not stated
WGCNA coexpression module analysis identification of disease-correlated module (turquoise) from merged dataset approximate scale-free topology; soft-threshold power 5
ROC/AUC (pROC) validation of combined four-biomarker model in GSE90880 (AUC = 0.942) na
CIBERSORT deconvolution estimation of 22 immune cell subsets, lesional vs normal p < 0.05 filter for the infiltration matrix na
Spearman correlation biomarker (TYR, TYRP1, DCT, LARP7) vs immune cell fractions (Figure 7) not stated
Approaches that could also have been used
  • DEGs were called using p < 0.05 (nominal) together with |log2FC| > 1.
    Could also: An adjusted-p (e.g., Benjamini-Hochberg FDR) threshold could also be applied to the per-gene limma results. — Adjusted p-values control the false discovery rate across the many genes tested, which is a common convention for genome-wide expression screens and would make the DEG list directly comparable across datasets.
  • Biomarker discrimination was summarized with a single AUC point estimate (0.942) on the validation set.
    Could also: A 95% confidence interval for the AUC (e.g., via DeLong or bootstrap, both available in pROC) could also be reported. — An interval communicates the precision of the AUC, which is informative given modest sample sizes in the GEO datasets.
  • Biomarker–immune cell associations were assessed with Spearman correlations reported as nominal p-values.
    Could also: A multiplicity correction (e.g., Benjamini-Hochberg) across the family of biomarker × cell-type correlations could also be applied. — With four biomarkers across many immune subsets, family-wise or FDR adjustment would account for the number of correlations examined.
  • Three feature-selection algorithms (LASSO, SVM-RFE, RF) were combined by intersecting their selected gene sets with WGCNA modules.
    Could also: A resampling/cross-validation or nested-CV stability assessment of the selected features could also be performed. — Repeated resampling characterizes how stable the selected biomarkers are to data perturbation, complementing the single-pass intersection.
  • Group sizes per dataset are reported but no sample-size or power justification is described.
    Could also: A brief statement of available n and its implications, or a sensitivity analysis, could also accompany the analysis. — Documenting the basis for n helps readers interpret the strength of evidence for small-cohort transcriptomic comparisons.
  • CIBERSORT-estimated fractions were filtered at p < 0.05 and compared between groups, with results shown via violin plots.
    Could also: The specific between-group test for each cell type (e.g., Wilcoxon rank-sum) and its dispersion summary could also be stated explicitly. — Naming the comparison test and reporting spread (SD/IQR/CI) makes the immune-infiltration differences reproducible and easier to interpret.
Software: R/limma · R/RRA (RobustRankAggreg) · R/clusterProfiler · R/WGCNA · R/e1071 (SVM-RFE) · R/pROC · CIBERSORT · R/ggplot2, corrplot, ggstatsplot · Reference gene set c2.cp.kegg.v7.2.symbols.gmt v7.2

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
37
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE53 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE53146 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE65127 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE75819 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE90880 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

Downstream reach in the literature

20 downstream papers · 5 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

GSE53 GEO reused by 3 papers in the literature
Most-cited downstream papers:
GSE90880 GEO reused by 3 papers in the literature
Most-cited downstream papers:

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34107850

Paper: Zhang J et al. (2021) Bioengineered. "Identification of TYR, TYRP1, DCT and LARP7 as related biomarkers and immune infiltration characteristics of vitiligo via comprehensive strategies." PMID 34107850 / PMC8806433.

Code link reality (P16 note)

Registry code_url = https://github.com/IndrajeetPatil/ggstatsplot. This is a generic third-party R plotting package (correlation/violin plots), not the authors' analysis code. The paper ships no analysis code. Per the brief, applying a third-party tool to the paper's data is equally valid — so we reproduce the described pipeline (standard packages: GEOquery, limma) on the paper's public GEO data, following the paper's stated parameters.

Pipeline-derived results (the paper's "comprehensive strategies")

  1. DEG identification per dataset via limma, threshold p<0.05 & |log2FC|>1. → reported counts per dataset. [IN SCOPE — primary, cleanly specified]
  2. RRA (RobustRankAggreg) across the 3 datasets → 131 robust DEGs (89 up, 42 down). [IN SCOPE — secondary; method named but params loose]
  3. WGCNA (soft-power 5, MEDissThres 0.25, 10 modules, turquoise module). [OUT — 20%; needs full expression + trait matrices, fragile to reproduce 1:1]
  4. GO/KEGG/GSEA enrichment (FDR<0.25, p<0.05). [OUT — qualitative result lists]
  5. Biomarker selection via overlap of LASSO + SVM-RFE + RF + WGCNA → TYR, TYRP1, DCT, LARP7. [OUT — 20%; stochastic ML, seeds unspecified]
  6. CIBERSORT immune infiltration (22 cell types) + biomarker–immune correlations. [OUT — 20%; depends on upstream gene set + LM22 signature run]
  7. Validation ROC on GSE90880, AUC=0.942. [OUT — depends on step 5 model]

What we attempt (80/20)

  • Primary: reproduce the per-dataset limma DEG counts, focus on the room's accession GSE65127 (73 DEGs: 9 up, 64 down); also GSE53146 (847: 412/435) and GSE75819 (1064: 777/287) as additional clean data points (all small, cheap).
  • Secondary (if cheap): RRA across the three → 131 robust DEGs (89/42).
  • Not attempted: WGCNA, ML biomarker selection, CIBERSORT, validation ROC — these are the hard ~20% (stochastic / under-specified seeds & inputs). We do NOT chase them; partial reproduction is a valid outcome.

Data

  • GSE65127 (primary), GSE53146, GSE75819 — all public GEO microarray series. Downloaded via GEOquery inside the «our HPC» compute job («infra» cwd).
DEG_GSE65127_total
Reported
73
Reproduced
67 (lesional vs all-other-sites) / 171 (lesional vs healthy 10v10)
partial
DEG_GSE65127_down
Reported
64
Reproduced
67
within tolerance
DEG_GSE65127_up
Reported
9
Reproduced
0 (10v30) / 44 (10v10)
did not match
DEG_GSE53146_total
Reported
847
Reproduced
1000
partial
DEG_GSE75819_total
Reported
1064
Reproduced
1118
within tolerance
RRA_robust_total
Reported
131
Reproduced
96
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 55/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

The limma DEG pipeline reproduces the scale and direction of every per-dataset result (GSE75819 within ~5%, GSE65127 down-count near-exact at 67 vs 64), so there is no fabrication concern — every reported value lies within the range of defensible parameter choices, with the paper's GSE65127 total of 73 sitting between our 67 (lesional-vs-all) and 171 (lesional-vs-healthy). The non-1:1 deviations are predominantly on the authors'/data side: the paper omits the control-group composition, probe->gene collapsing, and p-type, and the registry code link is a generic plotting package rather than the analysis code. Severity is moderate (magnitude and direction hold), and the central biomarker/immune-infiltration conclusion was the deliberately-skipped stochastic 20%, so it is neither confirmed nor refuted — overall a solid partial reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

85 k
tokens (I/O) · 4.3 M incl. cache
12 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
2
HPC jobs
hummel
machine