Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Screening of Diagnostic Biomarkers and Immune Infiltration Characteristics Linking Rheumatoid Arthritis and Rosacea Based on Bioinformatics Analysis.

J Inflamm Res · 2024
L1 48/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
48/100
Reproducibility score
1.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 6% of all assessed papers rank 1092 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reconstruct the pipeline and the data; sample-level identification reproduced the paper EXACTLY (GSE12021 9Ctrl/12RA, GSE55457 10Ctrl/13RA dropping OA, GSE65914 20Ctrl/38rosacea). The limma DE step reproduces well for the single-dataset rosacea claim (954 vs 1024 reported DEGs, ~93%, same up>down direction = within-tol). The RA combined-dataset DEG count does NOT reproduce: 720 vs reported 4304 (~6x lower, same down>up direction = mismatch). The discrepancy concentrates in the RA combine/ComBat/normalization step, which the paper underspecifies and which is NOT derivable from the shipped repo: the pipeline input merge.txt is absent and normalize.R's hardcoded group vectors are template residue from an unrelated 'RIF' (recurrent implantation failure) study (54 samples, wrong labels) -- flagged for human review (possible undocumented preprocessing vs error; not asserted fabrication). NOT ATTEMPTED: CIBERSORT immune-fraction values/Fig6 (qualitative, tooling threading bug, dropped per 80/20); WGCNA + cytoHubba/MCODE hub genes (14/17) and 277 co-expressed DEGs (Cytoscape GUI / non-pipeline); ROC of CXCL10/CCL27, ELISA, qRT-PCR, flow cytometry (wet-lab). All heavy compute on «our HPC» SLURM; repo pinned @3329b58; R4.3.3/GEOquery2.70/limma3.58/sva3.50. Large matrices kept on «infra» with SHA256 recorded in data/data.json.

💻 Code ↗ 🗄 Data: GSE12021

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 48
    assessed: 2026-06-14 ⛓ 9b7504813542
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Although rosacea is associated with an increased risk of rheumatoid arthritis (RA), the molecular mechanisms linking the two diseases are unknown; this study aims to uncover shared molecular regulatory networks, diagnostic biomarkers, and immune cell infiltration patterns common to RA and rosacea.

Core claims
  • 277 co-expressed differentially expressed genes are shared between RA and rosacea and are enriched in immune processes and chemokine-mediated signaling pathways finding
  • CXCL10 and CCL27 are potential diagnostic biomarkers linking RA and rosacea, with elevated serum levels in patients finding
  • Macrophages are the most abundant immune cells in RA, whereas dendritic cells are the most abundant in rosacea finding
  • CXCL10 may regulate macrophage function/density and CCL27 may regulate dendritic cell function/density in these diseases mechanism
  • A combined bioinformatics pipeline (limma DEGs, GO/KEGG, PPI/cytoHubba, WGCNA, CIBERSORT) identifies overlapped hub genes connecting RA and rosacea method
  • 14 hub genes overlapped in RA and 17 hub genes overlapped in rosacea between cytoHubba and WGCNA finding
  • CXCL10, CCR7, SAA1, and CCR5 are high-frequency hub genes in the PPI network finding
Experimental setups
Assay System Perturbation Readout Platform
Microarray gene expression / DEG analysis (limma) RA synovial membrane tissue (GSE12021, GSE55457) and rosacea skin tissue (GSE65914) none (disease vs control) differentially expressed genes (|log2FC|>=1, adj P<0.05) Affymetrix GPL96 (RA) and GPL570 (rosacea)
Immune cell deconvolution (CIBERSORT) RA and rosacea tissue expression profiles none abundance of 22 immune cell types (LM22 matrix) LM22 leukocyte gene signature matrix
WGCNA co-expression network analysis RA and rosacea expression datasets none module-trait correlation, hub modules
ELISA serum from RA and rosacea patients none CXCL10 and CCL27 serum concentration (absorbance at 450 nm) Invitrogen KAC2361
Flow cytometry (FCM) PBMC-derived macrophages and dendritic cells from human blood none immune cell densities/markers (CD45, CD86, CD206, CD14, CXCL10, HLA-DR, CD1a, CCL27) CytoFLEX (Beckman Coulter); FlowJo v10.07
qRT-PCR PBMC-isolated macrophages (treated with rHs CXCL10) and DC cells (treated with rHs CCL27) recombinant CXCL10 (0/2 ng/mL) on macrophages; recombinant CCL27 (0/2 ng/mL) on DCs, 24 h relative mRNA expression of IL-1β, IL-6, IFN-α, TNF-α (normalized to ACTB, 2−ΔΔCt) Applied Biosystems 7500 Real-Time PCR System; SYBR Green (TaKaRa)
ROC analysis RA and rosacea expression/biomarker data none AUC of CXCL10 and CCL27 as diagnostic models pROC R package
PPI network / hub gene analysis 277 co-expressed DEGs none hub genes and modules (cytoHubba, MCODE) STRING; Cytoscape v3.8.1
Key results
  • 277 co-expressed DEGs identified as shared between RA and rosacea 277 genes
  • RA combined dataset yielded 4304 DEGs (1961 up, 2343 down) 4304 DEGs
  • Rosacea dataset yielded 1024 DEGs (636 up, 388 down) 1024 DEGs
  • CXCL10 and CCL27 showed higher serum levels in RA/rosacea patients and good diagnostic value AUC>0.7
  • Turquoise module most correlated with RA correlation coefficient=0.79, p=7.9e-9
  • Pink module most correlated with rosacea correlation coefficient=0.80, p=2.54e-12
  • 14 overlapping hub genes in RA and 17 in rosacea between cytoHubba and WGCNA 14 and 17 genes
  • Macrophages most abundant in RA and dendritic cells most abundant in rosacea by CIBERSORT
Key statistics
  • correlation correlation coefficient = 0.79, p = 7.9e-9 (turquoise module correlation with RA (WGCNA))
  • correlation correlation coefficient = 0.80, p = 2.54e-12 (pink module correlation with rosacea (WGCNA))
  • correlation r = 0.79 (module membership vs gene significance, turquoise module RA)
  • correlation r = 0.80 (module membership vs gene significance, pink module rosacea)
  • count 4304 DEGs (1961 up, 2343 down) (RA combined dataset DEGs (GSE12021+GSE55457 samples))
  • count 1024 DEGs (636 up, 388 down) (rosacea dataset DEGs (38 rosacea + 20 control))
  • count 277 co-expressed DEGs (shared DEGs between RA and rosacea)
  • other AUC > 0.7 (threshold for good ROC model fit for CXCL10/CCL27)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics study downloaded public microarray datasets (GEO) for rheumatoid arthritis and rosacea, applied SVA batch correction and limma-based differential expression analysis to identify co-expressed DEGs, then used PPI network hub-gene scoring (cytoHubba, 5 topological patterns) and WGCNA module-trait correlation to converge on overlapping hub genes. Immune infiltration was estimated with CIBERSORT and correlated with hub genes via Pearson coefficients. Candidate biomarkers (CXCL10, CCL27) were evaluated with ROC/AUC and validated in a clinical cohort using ELISA, qRT-PCR, and flow cytometry, with Student's t-test or Kruskal-Wallis H-test applied depending on distributional assumptions.

Replicationbiological Sample sizeGEO datasets: GSE12021 (9 controls, 12 RA), GSE55457 (10 controls, 13 RA), GSE65914 (20 controls, 38 rosacea); clinical validation cohort size not stated in available text GroupsRA vs normal controls; rosacea vs normal controls; disease patients vs healthy controls in clinical validation cohort Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionAdjusted P-value (limma default: Benjamini-Hochberg FDR) for transcriptome-wide DEG identification; Dunn's test post-hoc correction following Kruskal-Wallis for multi-group pairwise comparisons in clinical validation
Statistical tests used
Test Applied to n Assumptions
limma moderated t-statistic (empirical Bayes linear model) DEG identification: RA combined dataset (GSE12021 + GSE55457) and rosacea dataset (GSE65914) vs respective normal controls GSE12021: 12 RA + 9 controls; GSE55457: 13 RA + 10 controls; GSE65914: 38 rosacea + 20 controls not stated
Spearman correlation coefficient WGCNA module-trait relationship heatmaps for RA and rosacea (turquoise module r=0.79, p=7.9e-9; pink module r=0.80, p=2.54e-12) null not stated
Pearson correlation coefficient (r) Correlation of overlapped hub genes with 22 CIBERSORT leukocyte gene signature matrix immune cell types; filtered at |r|>0.4 and P<0.05 null not stated
CIBERSORT (nu-support vector regression deconvolution with LM22 matrix) Immune cell fraction estimation across all RA and rosacea samples; samples filtered at CIBERSORT P<0.05 null not stated
ROC curve / AUC (pROC package, AUC>0.7 threshold) Diagnostic biomarker performance evaluation of CXCL10 and CCL27 null na
Student's t-test (two-group, presumably two-sided) Clinical validation cohort: normally distributed continuous variables between disease and control groups (ELISA concentrations, FCM data) null stated
Kruskal-Wallis H-test with Dunn's multiple-comparisons post-hoc test Clinical validation cohort: non-normally distributed continuous variables between groups null stated
Approaches that could also have been used
  • Pearson correlation was used to link hub gene expression to CIBERSORT immune cell proportion estimates
    Could also: Spearman rank correlation could also be applied to these associations — Pearson assumes bivariate normality; deconvolution-derived cell proportion estimates are frequently skewed or zero-inflated, making a rank-based approach a natural complement; both are standard in immune infiltration correlation analyses
  • CIBERSORT with the LM22 matrix was the sole method used for bulk-tissue immune cell deconvolution
    Could also: Additional deconvolution tools such as xCell, MCP-counter, TIMER2.0, or EPIC could also be applied — Different algorithms use distinct reference signatures and optimization frameworks; reporting concordance across two or more methods is one approach to characterizing the robustness of immune infiltration estimates
  • Hub genes were prioritized by intersecting cytoHubba rankings (5 topological scoring patterns) with WGCNA key module membership
    Could also: Other network centrality metrics (betweenness centrality, eigenvector centrality, PageRank) or community detection algorithms (Louvain, InfoMap) could also be used for hub gene identification — Degree-based and betweenness-based centrality capture different structural roles within a network; reporting sensitivity of the final hub gene list to the choice of scoring method is one approach to assessing prioritization robustness
  • ROC/AUC alone was used to characterize the diagnostic utility of CXCL10 and CCL27
    Could also: Decision curve analysis (DCA) or calibration analysis could also accompany ROC evaluation — AUC summarizes discrimination across all possible thresholds equally, whereas DCA estimates net clinical benefit at clinically relevant decision thresholds; calibration assesses agreement between predicted and observed event rates — together these give a more complete characterization of biomarker clinical value
  • Batch effects across the two RA GEO datasets were addressed using SVA (surrogate variable analysis) before combined limma analysis
    Could also: Explicit batch-variable correction methods such as ComBat (within the SVA package) or limma's removeBatchEffect function could also be used, treating known dataset origin as a batch covariate — SVA estimates latent confounders without requiring known batch labels; ComBat and removeBatchEffect allow direct modeling of known batch membership; applying and comparing both approaches is one way to assess sensitivity of the DEG list to the batch correction strategy
  • Multiple clinical outcomes (ELISA for CXCL10 and CCL27, multiple FCM markers, qRT-PCR for several genes) were each compared between groups without a stated multiplicity adjustment across this family of tests
    Could also: A Benjamini-Hochberg FDR or Bonferroni correction applied across the full family of clinical validation comparisons could also be used — Specifying a correction across the set of simultaneously tested outcomes defines the family-wise or false-discovery error rate being controlled; this is a commonly recommended additional step when multiple endpoints are evaluated in a validation cohort
Software: R/limma 4.2.3 · R/sva · R/WGCNA · R/clusterProfiler · R/pROC · R/Hmisc 4.4.1 · R/corrplot 0.84 · SPSS 23.0 · Cytoscape with cytoHubba and MCODE plugins 3.8.1 · CIBERSORT · FlowJo 10.07 · STRING database

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
3
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE12021 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE55457 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

159 downstream papers · 2 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39104909

Paper: Wang et al. 2024, J Inflamm Res. "Screening of Diagnostic Biomarkers and Immune Infiltration Characteristics Linking Rheumatoid Arthritis and Rosacea …" PMCID PMC11299729. DOI 10.2147/jir.s467760.

Repo: https://github.com/doctorxia203/FII-2023 (commit pinned at run time). Contents (5 files, 51 KB):

  • normalize.Rsva::ComBat batch correction of a pre-merged matrix GSE12021/merge.txt.
  • Diff.Rlimma differential expression (RA vs Ctrl), thresholds |logFC|>1 & adj.P<0.05.
  • Cibersort.R + sourcecibersort.R — CIBERSORT v1.03 (Newman 2015), uses LM22.txt.
  • LM22.txt — 547-gene × 22-cell-type leukocyte signature matrix (shipped).

Datasets (per Methods + Data Sharing Statement)

  • RA: GSE12021 (GPL96: 9 normal + 12 RA = 21) + GSE55457 (GPL96: 10 normal + 13 RA = 23) → combined RA dataset, 44 samples (25 RA, 19 control). Batch-corrected with sva::ComBat.
  • Rosacea: GSE65914 (GPL570: 20 control + 38 rosacea = 58). (Intro text inconsistently cites "GSE6591"; the Data Sharing Statement and sample counts identify GSE65914.)

IN SCOPE — pipeline-derived, code-backed

id result pipeline reported location
C1 RA combined DEGs: 4304 total (1961 up, 2343 down), |log2FC|≥1 & adjP<0.05 normalize.R (ComBat) + Diff.R (limma) Results §"Identification and Enrichment Analysis of DEGs"
C2 Rosacea DEGs: 1024 total (636 up, 388 down) same limma pipeline on GSE65914 same paragraph
C3 CIBERSORT 22-cell immune fractions; qualitative: macrophages most abundant in RA, DCs in rosacea Cibersort.R + LM22 Results §"Immune Infiltration" / Fig 6

Primary target: C1 — directly produced by the shipped normalize.R + Diff.R, a single clear quantitative claim, light compute. C2 is the same pipeline on a different GEO accession (secondary). C3 is qualitative.

OUT OF SCOPE (not attempted / why)

  • Probe-merge input not shipped. merge.txt (the gene×sample matrix feeding the pipeline) is NOT in the repo; it must be reconstructed from GEO (series matrices + GPL96 probe→symbol collapse by mean, per Methods). This reconstruction is the main reproducibility gap and a source of expected numeric drift.
  • normalize.R group vectors are template residue. Shipped batchType=c(rep(1,48),rep(2,6)) and modType=c(rep("Ctrl",24),rep("RIF",24)…) (54 samples, "RIF"=recurrent-implantation- failure labels from an unrelated study) do NOT match the described 44-sample RA design. Faithful reproduction requires correcting these to the documented 21+23 split — recorded as a necessary deviation, not run verbatim.
  • WGCNA, cytoHubba/MCODE PPI hub genes (Cytoscape, GUI, manual) — non-pipeline / not scriptable here.
  • 277 co-expressed DEGs, 14/17 hub genes — depend on WGCNA + Cytoscape steps above.
  • ELISA, qRT-PCR, flow cytometry, ROC of CXCL10/CCL27 — wet-lab, out of scope.

Compute plan

All compute on «our HPC» (SLURM, partition=std). Conda env built inside the compute job (internet on compute nodes): R + Bioconductor limma, sva, GEOquery, hgu133a.db, e1071, preprocessCore. Data downloaded to «infra» work dir inside the job.

Figures / tables: Fig 6
C1_total_RA_DEGs
Reported
4304 (1961 up / 2343 down)
Reproduced
720 (339 up / 381 down)
did not match
C2_total_Rosacea_DEGs
Reported
1024 (636 up / 388 down)
Reproduced
954 (592 up / 362 down)
within tolerance
C3_immune_cell_ranking
Reported
Macrophages (RA) / Dendritic cells (rosacea) most abundant
Reproduced
not completed (CIBERSORT preprocessCore/e1071 env failure)
m.public.grade.error

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 48/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🟡7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

The rosacea DEG claim reproduces well (954 vs 1024, ~93%, same up>down direction), but the primary RA combined-dataset claim of 4304 DEGs is not derivable from the shipped artifacts — a faithful re-implementation of the described limma+ComBat pipeline yields 720 (~6x lower). The gap localizes to the underspecified RA merge/ComBat/normalization step: the input merge.txt is not deposited and normalize.R hardcodes 'RIF' template residue from an unrelated recurrent-implantation-failure study (54 samples), so the defect sits on the authors'/data-availability side, not our method. Severity is high on the primary count, but the central biomarker/immune-infiltration conclusions were not testable here (CIBERSORT env failure; WGCNA/Cytoscape out of scope), so q7 is limited rather than refuted. Net: a critical reproducibility discrepancy with a real code/data red flag, but not asserted as fabrication.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

120.2 k
tokens (I/O) · 8.4 M incl. cache
14 min
runtime · 0.02 CPU-h
1.3 GB
peak RAM
3
HPC jobs
hummel
machine