Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Effect of PAIP1 on the metastatic potential and prognostic significance in oral squamous cell carcinoma.

Int J Oral Sci · 2022
L1 10/100 PQI 70
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
10/100
Reproducibility score
3.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 0% of all assessed papers rank 1169 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to attempt, but the one in-scope pipeline claim does NOT reproduce. The cited code repo (kunalchawlaa/TCGA-Oral-Cancer) is a trivial TCGA clinical-RNA merge utility that by its own README must NOT be used for differential expression and computes no statistic, so the only checkable pipeline result is Fig 1a: PAIP1 (213754_s_at) differential expression in GSE30784 via GEO2R (=limma). Reproduced 1:1 with GEOquery+limma on «our HPC». Reported log2FC=1.217 / -log10P=28.329 is NOT derivable from GSE30784: every PAIP1 probe (5) under every statistic (limma/Welch/MWU/linear-FC) gives log2FC0.4-0.6 and -log10P<=12.85; the reported -log10P=28.329 (P5e-29) is mathematically implausible for the named 90-sample (or full 212-sample) comparison, and the legend's '45 cancer' is wrong (GSE30784 has 167 OSCC, 17 dysplasia, 45 normal). Flagged as possible-fabrication / mislabeled-source; the value format (log2FC + -log10 Mann-Whitney P) is most consistent with a pooled multi-cohort tool such as TNMplot.com, not the cited GEO series. NOT ATTEMPTED (out of scope): TCGA/CCLE copy-number-mRNA correlations (Fig S1, no shipped stats code), KM-plotter survival (external web tool, no numbers in text), all wet-lab. No completeness claim; numeric reproduction is deterministic and re-runnable.

💻 Code ↗ 🗄 Data: GSE30784

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 10
    assessed: 2026-06-15 ⛓ 09ee9362fe15
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether PAIP1 is frequently overexpressed in oral squamous cell carcinoma (OSCC), correlates with clinicopathological features such as lymph node metastasis and prognosis, and functionally promotes OSCC metastatic potential.

Core claims
  • PAIP1 mRNA and protein levels are upregulated in OSCC/HNSCC compared to normal tissue. finding
  • High PAIP1 expression correlates with advanced stage, lymph node metastasis (LNM), and worse pattern of invasion, indicating poor prognosis. finding
  • PAIP1 knockdown attenuates colony formation, migration, and invasion of OSCC cells. finding
  • PAIP1 promotes OSCC invasiveness via increased MMP9 activity (not MMP2 or EMT molecules in this model). mechanism
  • PAIP1 regulates SRC phosphorylation (Tyr416/419), with PAIP1 and pSRC positively correlated in patient samples. mechanism
  • PAIP1 copy number is significantly associated with mRNA levels in HNSCC, indicating amplification-driven overexpression. finding
  • PAIP1 expression is higher at the invasive tumor front (ITF) than the inner tumor mass (ITM). finding
  • PAIP1 may serve as an independent prognostic marker and therapeutic target in OSCC with LNM. resource
Experimental setups
Assay System Perturbation Readout Platform
in silico differential expression / mRNA analysis HNSCC/OSCC patient tissue (GEO GSE30784, GSE37991, GSE78060) none PAIP1 mRNA expression (log2) GEO microarray reporters 213754_s_at and ILMN_1776398; GEO2R
copy number vs mRNA correlation analysis HNSCC tumors and cell lines (TCGA, CCLE) none PAIP1 copy number vs mRNA correlation TCGA/CCLE databases (cBioPortal)
RNA-seq expression analysis HNSCC custom cohort tissue (TCGA GDC) none PAIP1 mRNA (FPKM-UQ, log2) TCGA GDC; Python 3.9/Pandas/Jupyter
proteomic analysis HNSCC/OSCC patient tissue none PAIP1 and pSRC (peptide LIEDNEyTAR, Tyr419) protein levels CPTAC Proteomic Data Commons
immunohistochemistry (IHC) OSCC patient tissue (tumor vs normal; ITM vs ITF; serial sections) none semi-quantitative PAIP1 and pSRC (Tyr419) IHC scores
siRNA knockdown + colony formation, migration, invasion assays OSCC cell lines HN22 and SCC-9 PAIP1 knockdown (siRNA) colony number, migratory and invasive ability
gelatin zymography OSCC cell lines HN22 and SCC-9 PAIP1 knockdown (siRNA) MMP9 enzymatic activity
Western blot OSCC cell lines HN22 and SCC-9 PAIP1 knockdown (siRNA) SRC (Tyr416) phosphorylation
Key results
  • PAIP1 transcript significantly upregulated in cancer vs normal in GSE30784 (volcano plot) log2 fold change 1.217, -log10(P) 28.329
  • PAIP1 copy number significantly associated with mRNA levels in HNSCC (TCGA and CCLE) R=0.81 (TCGA), R=0.707 (CCLE)
  • PAIP1 expression higher at invasive tumor front (ITF) than inner tumor mass (ITM) P<0.0001
  • Increased PAIP1 expression associated with increased risk of non-cohesive worst pattern of invasion (WPOI) relative risk 1.38 (95% CI 1.06–1.8)
  • PAIP1 knockdown decreased colony forming, migration, and invasion in HN22 and SCC-9
  • PAIP1 knockdown significantly decreased MMP9 activity in both cell lines
  • PAIP1 knockdown strongly attenuated SRC (Tyr416) phosphorylation
  • PAIP1 and pSRC (Tyr419) IHC scores positively correlated in OSCC tissue R=0.50, P=0.0001
Key statistics
  • correlation R=0.81 (PAIP1 copy number vs mRNA in HNSCC (TCGA))
  • correlation R=0.707 (PAIP1 copy number vs mRNA in HNSCC (CCLE))
  • fold_change log2 fold change 1.217 (PAIP1 in 45 normal vs 45 cancer samples (GSE30784))
  • pvalue -log10(P)=28.329 (PAIP1 differential expression in GSE30784)
  • correlation R=0.50, P=0.0001 (PAIP1 vs pSRC (Tyr419) IHC scores, Pearson correlation)
  • correlation R=0.724, P=0.003 (PAIP1 vs pSRC peptide (LIEDNEyTAR, Tyr419) protein, paired normal/tumor (CPTAC))
  • other relative risk 1.38 (95% CI 1.06–1.8) (high PAIP1 expression and non-cohesive WPOI risk)
  • pvalue P<0.0001 (ITF vs ITM PAIP1 IHC scores)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper investigated PAIP1 expression in oral squamous cell carcinoma using a multi-platform in silico approach (GEO, TCGA, CPTAC), semi-quantitative IHC on patient tissue, and siRNA knockdown in two OSCC cell lines. Group comparisons were made with Student's t-tests and one-way ANOVA; continuous associations were quantified with Pearson correlation; clinicopathological and histopathological risk was summarized via multivariable relative-risk forest plots. In vitro functional assay data were reported as mean ± SD of triplicates, and survival was evaluated by Kaplan-Meier analysis using KM Plotter. Results were reported with a P < 0.05 threshold throughout, with selected exact P- and R-values given for correlation analyses.

Replicationmixed Sample size45 normal and 45 cancer samples stated for GSE30784; in vitro experiments described as triplicates (likely technical); sample sizes for IHC cohort and remaining GEO/TCGA/CPTAC subsets not explicitly stated in main text GroupsNormal vs. OSCC tissue (multiple cohorts); LNM-positive vs. LNM-negative; TNM stages I–IV; WPOI types I–V (cohesive vs. non-cohesive); PAIP1-knockdown vs. scramble-control OSCC cell lines; high vs. low PAIP1 for survival Pairingmixed Randomization/blindingnot stated DispersionSD Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Student's t-test (two-group comparison) PAIP1 expression in normal vs. tumor across GEO datasets (GSE30784, GSE37991, GSE78060), TCGA, CPTAC, and IHC cohorts; also in vitro colony formation, migration, invasion, MMP9 activity, and pSRC western blot 45 normal and 45 cancer samples for GSE30784; in vitro comparisons described as triplicates; cohort n for IHC and other GEO/TCGA datasets not explicitly stated in main text not stated
One-way ANOVA PAIP1 expression across multiple ordered groups (e.g., TNM stages I–IV; WPOI types I–V) not stated not stated
Pearson correlation PAIP1 vs. pSRC IHC scores in OSCC tissue (R = 0.50, P = 0.0001); PAIP1 vs. pSRC protein from CPTAC (R = 0.724, P = 0.003); PAIP1 copy number vs. mRNA in TCGA (R = 0.81) and CCLE (R = 0.707); multivariable Pearson correlation heat map of PAIP1, pSRC, WPOI, and nodal metastasis (Fig. 5c) not stated not stated
Multivariable relative-risk analysis (forest plot) Clinicopathological variables vs. PAIP1 expression (Fig. 2b); histopathological variables including WPOI vs. PAIP1 (Fig. 3d); RR for non-cohesive WPOI = 1.38, 95% CI 1.06–1.80 not stated not stated
Kaplan-Meier survival analysis (log-rank, via KM Plotter) Overall survival stratified by high vs. low PAIP1 expression in OSCC patients (Fig. S3) not stated not stated
Differential expression analysis (GEO2R, likely limma-based moderated t-statistic) Volcano plot of 45 normal vs. 45 cancer samples from GSE30784; PAIP1 reported as log2(FC) = 1.217, −log10(P) = 28.329 45 normal, 45 cancer (GSE30784) not stated
Approaches that could also have been used
  • IHC semi-quantitative scores and public-database expression values were compared between groups using Student's t-test
    Could also: Mann-Whitney U test (Wilcoxon rank-sum) — IHC scores are ordinal and semi-quantitative and may not follow a normal distribution; a non-parametric rank-based test requires no distributional assumption and is commonly chosen for this data type
  • Multiple pairwise group comparisons (across stages, WPOI types, LNM groups, and multiple independent datasets) were each evaluated at P < 0.05 without a stated correction for the family of tests
    Could also: Benjamini-Hochberg FDR correction or Bonferroni correction applied across the family of comparisons — When many tests share a common hypothesis space, correction methods such as FDR or Bonferroni are widely used to characterize the expected proportion of false positives across the full test family
  • Pearson correlation was used to quantify the association between PAIP1 IHC scores and pSRC IHC scores
    Could also: Spearman rank correlation — Spearman correlation assumes neither bivariate normality nor linearity and is often preferred for ordinal or semi-quantitative IHC data; it is also less sensitive to extreme values that may arise in small scored samples
  • Overall survival by PAIP1 expression level was evaluated with univariate Kaplan-Meier analysis via KM Plotter
    Could also: Multivariable Cox proportional hazards regression with PAIP1 as a covariate alongside stage, LNM, and WPOI — Cox regression provides a hazard ratio with confidence interval and allows simultaneous covariate adjustment for the prognostic variables identified in the paper's own multivariable analysis, complementing the univariate KM approach
  • In vitro functional assays (colony formation, migration, invasion, MMP9, pSRC) were performed in triplicates and summarized as mean ± SD
    Could also: Independent biological replicates (distinct cell passages or experiments) reported with individual data points overlaid on summary statistics — When n is small, individual data points alongside mean ± SD allow readers to directly assess spread and distributional shape; distinguishing biological from technical replicates is also recommended by major reporting guidelines (e.g., Nature Methods) for in vitro studies
  • Relative risk was reported in the multivariable forest plots for clinicopathological and histopathological variables
    Could also: Odds ratios from binary logistic regression with explicit reporting of covariates, reference categories, and model fit — Odds ratios (from logistic regression) are the conventional effect measure for binary outcomes such as LNM presence and are more readily compared across the OSCC literature; explicit model specification also facilitates replication and cross-study synthesis
Software: GEO2R · Python 3.9 / Pandas / Jupyter notebook 3.9 · KM Plotter

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
12
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

1and PDBe in Introduction (http://purl.org/orb/Introduction)
no other assessed paper uses this yet
GSE30784 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE37991 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE78060 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

175 downstream papers · 3 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35153296

Paper: Swarup et al. 2022, Int J Oral Sci — "Effect of PAIP1 on the metastatic potential and prognostic significance in oral squamous cell carcinoma." DOI 10.1038/s41368-022-00162-8. PMCID PMC8841500.

Code: https://github.com/kunalchawlaa/TCGA-Oral-Cancer — Python merge_info.py, a trivial TCGA clinical↔RNA merge utility. Its README explicitly says it "can only be used for single RNA profiling and should be avoided for differential expression analysis." It computes no statistic — it joins a clinical Excel with an RNA-value Excel. So it produces no reported numeric result to compare against.

Data: GEO GSE30784 (Affymetrix HG-U133 Plus 2.0, GPL570) — oral normal/dysplasia/OSCC microarray (Chen et al.).

In scope (pipeline-derived, attempted)

id result location pipeline
C1 PAIP1 (probe 213754_s_at) differential expression in GSE30784: log2FC = 1.217, −log10(P) = 28.329 Fig. 1a legend GEO2R (limma) on GSE30784, "45 normal and 45 cancer samples"

Methods quote: "We performed the differential expression analysis of GSE30784 … for the reporter id 213754_s_at … GEO2R was used to confirm normalization and analysis." Fig 1a legend: "…45 normal and 45 cancer samples from GSE30784 … PAIP1 log2(fold change) 1.217, −log10 (P-value) 28.329."

This is the one clearly-specified, low-hanging pipeline output: public GEO data, named tool (GEO2R = limma), named probe, exact numbers. We reproduce it 1:1 with GEOquery + limma on «our HPC» (the GEO2R engine).

Out of scope (not a comparable pipeline output / external tool / wet-lab)

  • TCGA HNSC analyses (Fig S1c/d: copy-number↔mRNA R=0.81 TCGA / R=0.707 CCLE; stage/LNM/WPOI boxplots Fig 2): the authors' repo only merges TCGA tables; the correlation/grouping stats are done ad-hoc (Jupyter/Pandas, not shipped). Heavy TCGA + CCLE downloads; no shipped code computing R. Down-ranked, not attempted.
  • Survival (KM plotter) — external web tool (kmplot.com), no HR/p reported in text; manual, not a reproducible local pipeline. Out of scope.
  • All wet-lab (IHC, migration/invasion assays, qPCR, Western) — out of scope.

Deviations / ambiguities to flag

  • "45 normal and 45 cancer" but GSE30784 = 45 normal / 17 dysplasia / 167 OSCC. Which 45 OSCC were used for Fig 1a is unspecified → the p-value (n-dependent) may not match exactly; the log2FC (mean difference) is the robust target.
  • "−log10(P)" — unclear whether raw or adjusted P; we report both.
Figures / tables: Fig. 1a
C1
Reported
PAIP1 probe 213754_s_at in GSE30784 (Fig 1a, GEO2R): log2FC=1.217, -log10(P)=28.329, '45 normal and 45 cancer samples'
Reproduced
log2FC=0.494 (GEO2R/limma diff-of-log2-means; 0.591 as linear-mean ratio); max -log10P=12.85 (Welch t); limma -log10P=7.17; MWU -log10P=8.50. 45 normal vs 167 OSCC (true GSE30784 composition). Robust across all 5 PAIP1 probes and 45-45 subsets.
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 10/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🟡7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

The one in-scope, public-data pipeline claim (Fig 1a PAIP1 213754_s_at DE in GSE30784) does not reproduce: reported log2FC=1.217 / -log10P=28.329 vs reproduced log2FC≈0.49-0.59 and -log10P≤12.85, robust across all 5 PAIP1 probes and 4 statistics. The reported p-value (5e-29) is mathematically impossible for the cited ≤212-sample comparison and the legend's '45 cancer' contradicts the series' true 167 OSCC, pointing to a mislabeled/pooled source (TNMplot.com-style) rather than GSE30784 — this is on the authors' side (value not derivable from shared data, fabrication-suspect). The qualitative conclusion that PAIP1 is significantly overexpressed in OSCC does survive (limma P7e-8), so the central biology is limited-confirmed even though the specific reported figures are not reproducible.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

98.7 k
tokens (I/O) · 5.4 M incl. cache
12 min
runtime · 0.01 CPU-h
1.5 GB
peak RAM
2
HPC jobs
hummel
machine