Qualitative Transcriptional Signature for the Pathological Diagnosis of Pancreatic Cancer.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the CORE result 1:1 on «our HPC». The repo (xlucpu/PCsig @ 96df2c0) ships the authors' own REO classifier functions; the paper publishes the final 12-gene-pair signature (Table 2) and per-dataset validation metrics (Table 3). I applied the published 12 pairs via the authors' valuate.performance() to the GEO validation cohorts, fetching them with GEOquery (GEO download on front1) and running the scoring as SLURM «job» (COMPLETED). STRONGEST EVIDENCE: GSE19650 -- the only GEO validation set with BOTH classes -- reproduces BOTH reported metrics to rounding: sensitivity 100% (9/9) and specificity 85.7%=6/7 (paper 85.8%). GSE50827 sensitivity 94.2% (97/103) vs 95.1% (one sample off); GSE43288 specificity exact 100% (3/3). This confirms the signature is genuine and the code works -- no fabrication signal. DISCREPANCIES, all explained: GSE62165 reproduced HIGHER sens (100% vs 92.3%, single-channel data, attributable to unspecified probe->symbol collapse on borderline samples); GSE43288 sens 75% (n=4, with 3/17 genes absent on that platform). REAL BLOCKERS: GSE21501 is a confirmed Agilent two-color array (47% negative values = log-ratios vs a common reference), which confounds within-sample REO gene-vs-gene ordering (reproduced 87.1%; also 132 GEO samples != paper's 102); GSE71729's deposited GENE_NAME annotation resolves 0/17 signature genes (custom UNC array), so the signature cannot be mapped from this source. NOT ATTEMPTED: de-novo signature DISCOVERY (cross-platform merge recipe not specified) and the 3 non-GEO Table-3 cohorts. Net: a faithful, auditable PARTIAL reproduction with a clean dual-metric 1:1 anchor (GSE19650) and fully-documented, non-fabrication explanations for every deviation.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 63assessed: 2026-06-16 ⛓ c2477a588509
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetWithin-sample relative expression orderings (REOs) of genes can generate a robust, platform-independent qualitative transcriptional signature that accurately discriminates pancreatic cancer (PC) tissues and cancer-adjacent normal tissues from non-PC pancreatitis and healthy pancreatic tissues, including in biopsy specimens.
- ★ A 12-gene-pair (17-gene) REO-based signature discriminates PC and cancer-adjacent normal tissue from non-tumor (healthy/pancreatitis) pancreatic tissue. finding
- ★ The signature achieves 96.7% geometric mean sensitivity/specificity and AUC 0.978 across 1,007 PC and 257 non-tumor samples from 9 external validation datasets. finding
- ★ The signature achieves 100% diagnostic accuracy on 20 endoscopic biopsy specimens. finding
- ★ REO-based signatures are insensitive to batch effects, platform differences, tumor cell proportion, RNA degradation, and amplification bias, making them suitable for small/imperfect biopsy samples. mechanism
- RGPs (reversal gene pairs) were defined as gene pairs with consistent REO patterns in >85% of tumor samples and a reversed pattern in >85% of non-tumor samples in training data. method
- A majority-voting classification rule using the top-ranked gene pairs was used to classify samples as tumor or non-tumor. method
- Several signature genes (LAMC2, CST6, S100P, CDH3) have established prior roles in PC carcinogenesis, invasion, and metastasis. mechanism
- Full analysis source code and data are provided via a GitHub repository (PCsig). resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Microarray gene expression profiling | Human pancreatic tissue (normal, pancreatitis, PC) — training cohort, 5 merged datasets (GSE101462, GSE71989, GSE91035, E-MEXP-1121, E-MTAB-1791) | none | Genome-wide mRNA expression used to compute within-sample REOs and identify reversal gene pairs | Illumina GPL10558, Affymetrix GPL570/GPL96, Agilent GPL22763, Illumina WG6 BeadChip v3; RMA normalization via R package 'affy' |
| Microarray gene expression profiling | Human pancreatic tissue, tuning cohort GSE41368 (6 normal, 6 PC) | none | Filtering of candidate RGPs to establish candidate signature | Affymetrix Human Gene 1.0 ST Array |
| Microarray/RNA-seq gene expression profiling | Human pancreatic tissue, 9 external validation datasets (surgically resected, mostly) | none | REO signature classification accuracy, sensitivity, specificity, AUC | Illumina, Affymetrix, Agilent (various GPLs) |
| Microarray gene expression profiling | Human pancreatic tissue from endoscopic biopsy (GSE43288) | none | Diagnostic accuracy of 12-gene-pair REO signature | Affymetrix GPL96 |
| RNA-sequencing | Normal pancreas tissue, GTEx (autopsy specimens) | none | Classification of tissue as normal by REO signature | Illumina TrueSeq RNA sequencing |
| RNA-sequencing (HTSeq-counts) | Pancreatic adenocarcinoma tumor tissue, TCGA-PAAD (surgical resection) | none | FPKM/TPM-derived expression, REO signature sensitivity | Illumina HiSeq RNASeqV2 |
| ROC analysis | Pooled training, tuning, and validation pancreatic tissue samples | none | AUC, sensitivity, and specificity of the 12-gene-pair signature | R package 'pROC' |
- – Forward selection identified 12 gene pairs (from 20 candidate RGPs) giving the highest training classification accuracy geometric mean sensitivity/specificity 93.79% at k=12
- ▲ 12-gene-pair signature validated across 1,007 PC and 257 non-tumor samples from 9 databases geometric mean sensitivity/specificity 96.7%; AUC 0.978 (95% CI 0.947–0.994)
- – Signature correctly classified endoscopic biopsy specimens (17 PC, 3 normal) 100% diagnostic accuracy
- – GTEx autopsy normal pancreatic samples correctly classified as normal 248/248 (100%)
- – Microarray surgically resected tumor tissues correctly identified as tumor 96.07% of 842 tumor tissues
- – TCGA surgical PC samples correctly identified as tumor sensitivity 93.4% (171 samples)
- – E-MTAB-6134 surgical PC samples correctly identified as tumor sensitivity 99% (306 samples)
- other geometric mean sensitivity/specificity = 96.7% (external validation across 9 datasets (1,007 PC + 257 non-tumor samples))
- other AUC = 0.978 (95% CI, 0.947–0.994) (pooled ROC analysis of validation datasets)
- other geometric mean sensitivity/specificity = 93.79% (training cohort accuracy at k=12 gene pairs)
- other diagnostic accuracy = 100% (endoscopic biopsy samples (GSE43288, 17 PC + 3 normal))
- other specificity = 100% (GTEx normal pancreatic autopsy samples)
- other sensitivity = 99% (E-MTAB-6134 surgical PC samples)
- other sensitivity = 93.4% (TCGA surgical PC samples)
- other specificity = 85.8% (GSE19650 non-tumor samples)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper develops a qualitative binary classifier for pancreatic cancer (PC) based on within-sample relative expression orderings (REOs) of gene pairs across multiple public microarray and RNA-seq datasets. Candidate reversal gene pairs (RGPs) were identified by applying a fixed 85% prevalence threshold in training data, filtered in a tuning dataset, and ranked by a geometric-mean rank-difference metric; a forward selection procedure with a majority-voting rule chose the optimal number of pairs (k=12). Diagnostic performance was assessed in nine independent external validation datasets using ROC analysis, AUC, sensitivity, and specificity, with no formal hypothesis tests or p-values reported.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Majority-voting rule (threshold: >50% of gene pairs vote 'tumor') | Training and all validation datasets — primary classification rule for each sample | 269 PC + 146 non-tumor samples in training; 1,007 PC + 257 non-tumor in external validation | not stated |
| ROC analysis and AUC calculation (R package pROC) | Pooled training + tuning + validation cohorts (Figure 3); AUC = 0.978, 95% CI 0.947–0.994 | 1,007 PC tissues and 257 non-tumor samples (external validation cohort) | not stated |
| Geometric mean of sensitivity and specificity (summary accuracy metric) | Training cohort (to select k=12 pairs, Figure 2) and overall validation (reported as 96.7%) | training: 269 PC + 146 non-tumor; validation: 1,007 PC + 257 non-tumor | na |
| Forward stepwise selection (one RGP added at a time, evaluated by geometric mean accuracy) | Training dataset — used to determine optimal number of gene pairs (k=1 to 20) | 269 PC + 146 non-tumor samples in training | not stated |
-
Sensitivity and specificity in each validation dataset were reported as single point estimates at the median voting threshold↳ Could also: Confidence intervals for sensitivity and specificity (e.g., exact binomial or bootstrap 95% CIs) could also be reported alongside point estimates — Per-dataset CIs would convey the precision of each performance estimate, which is especially informative for smaller datasets such as GSE43288 (n=20) or GSE19650 (n=21), where point estimates can vary substantially by chance
-
The 85% prevalence threshold for defining a 'stable' REO in each tissue class was fixed a priori↳ Could also: A sensitivity analysis varying this threshold (e.g., 80%, 85%, 90%) could also be performed — Reporting performance across a range of threshold values would help a reader understand how sensitive the resulting signature is to this specific design choice and where stability begins to break down
-
Forward stepwise selection was used to choose the number of gene pairs, evaluated by geometric mean of sensitivity and specificity on the same training data↳ Could also: Regularized selection methods (e.g., LASSO logistic regression applied to the binary REO features) could also be used to simultaneously select features and estimate their weights — Regularized methods incorporate a penalty that accounts for the large number of candidate features (269 RGPs), potentially reducing the risk of overfitting during the selection step; they also naturally provide a continuous risk score rather than a hard majority-vote boundary
-
Overall diagnostic performance was summarized by the geometric mean of sensitivity and specificity and by AUC↳ Could also: The Matthews Correlation Coefficient (MCC) or F1 score could also summarize binary classification performance — MCC and F1 explicitly account for class imbalance; because non-tumor samples are considerably fewer than tumor samples in several datasets, these metrics can complement the geometric mean by conveying performance on the minority class more directly
-
Validation relied on nine pre-existing external datasets drawn from public repositories↳ Could also: Bootstrap or repeated k-fold cross-validation within the combined training cohort could also estimate an internal optimism-corrected performance measure — Internal validation with optimism correction (e.g., .632+ bootstrap) quantifies the degree to which the signature may be optimistically tuned to the training distribution, complementing external validation and making the reported accuracy estimates easier to interpret
-
The AUC confidence interval was reported only for the pooled validation cohort; individual dataset AUCs were not reported↳ Could also: Per-dataset AUCs with CIs (e.g., via DeLong's non-parametric method) could also be reported — Per-dataset AUCs would reveal whether performance is consistent across platforms, tissue-preparation methods (FF vs. FFPE), and sampling strategies (biopsy vs. surgical resection), providing a clearer picture of generalizability
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 33173782 (PCsig REO qualitative signature)
Paper: Zhou et al. 2020, Front Mol Biosci 7:569842. "Qualitative Transcriptional
Signature for the Pathological Diagnosis of Pancreatic Cancer."
Code: https://github.com/xlucpu/PCsig @ commit 96df2c0704df528234ce0ff04931ca1eb8f1f3ef
Method family: REO (Relative Expression Ordering) — within-sample rank comparison
of gene pairs; classification by majority vote of "reversal gene pairs" (RGPs).
What the repo ships
Five authors' own R functions (no driver script, no data, no README):
slec_stable_pair_tum.R/slec_stable_pair_nor.R— select stable gene pairs whose REO direction holds in ≥ cutoff fraction of samples (cutoff is a parameter).diff.rank.R— rank-difference helper for ranking candidate pairs.identify.optimal.sig.R— greedy forward selection of pairs; computes Sens/Spec/Acc.valuate.performance.R— applies a fixed signature pair-list to an expression matrix, classifies each sample by majority vote, returns Sens/Spec ROC points + AUC (tdROC::calc.AUC).
In scope (pipeline-derived, attempted)
The cleanest faithful test of the authors' code on the paper's own data: take the
published 12-pair signature (Table 2) and run the authors' valuate.performance()
on each validation GEO dataset to reproduce the per-dataset Table 3 metrics.
Classification rule (paper + valuate.performance.R agree): a sample is "tumor" if
> half of the 12 pairs have GeneA > GeneB (i.e. ≥ 7/12; even-length cutoff =
ceiling(0.5·12)+1 = 7). Pairs in Table 2 are oriented GeneA>GeneB in tumor.
Target claims = Table 3 per-dataset sensitivity / specificity, for the GEO-hosted validation cohorts obtainable via GEOquery:
- GSE50827 (103 PC) — Sens 95.1%
- GSE19650 (7 normal, 9 PC) — Sens 100%, Spec 85.8% ← only GEO set with both classes
- GSE62165 (13 normal-adjacent, 118 PC) — Sens 92.3%
- GSE43288 (biopsy; 3 normal, 4 PC, 13 precursor) — Sens/Spec 100%
- GSE21501 (102 PC) — Sens 94.7%
- GSE71729 (46 normal-adjacent, 145 PC) — Sens 95.3%
Out of scope / not attempted (with reason)
- De-novo signature discovery (the 18.3M-pair stable-RGP search over 5 merged training datasets to derive the 12 pairs). Out of immediate scope because cross- platform merging of GSE101462+GSE71989+GSE91035+E-MEXP-1121+E-MTAB-1791 is under-specified (no probe-collapse / batch / gene-universe recipe shipped). We instead verify the published signature, which is the stronger auditable anchor. May extend.
- E-MTAB-6134 (Sens 99%) — ArrayExpress, not GEOquery; deferred.
- GTEx (Spec 100%) and TCGA (Sens 93.4%) — bulk portals; deferred (heavier fetch).
- All wet-lab / pathology-review / IHC content — non-computational, out of scope.
Reproduction choices that may affect exact match (documented)
- Probe→gene-symbol collapse: GEO series matrices are probe-level. The paper does not specify the collapse rule. We collapse multi-probe genes by the probe with the largest mean expression (standard in REO/Guan-lab pipelines). Alternative = mean; flagged.
- REO is rank-based within sample, so collapse choice and platform annotation are the main sources of small deviation vs reported %.
- TrueClass labels derived from each GEO series' sample phenotype fields.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The published 12-pair REO signature reproduces its strongest anchor (GSE19650: 100% sens, 85.7%≈85.8% spec) to rounding and matches GSE50827/GSE43288 closely, so the core diagnostic claim is supported with no fabrication signal. Remaining deviations sit on the input/preprocessing side — an unspecified probe→symbol collapse (our method choice) and two-color-array log-ratio data unsuitable for REO (data format) — not in the classifier logic. Severity is moderate (a few % per dataset, direction preserved) but two cohorts and the pooled AUC=0.978 could not be re-derived, so this is a solid partial reproduction (yellow) rather than a 1:1 confirmation.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.