Proteogenomic analysis prioritises functional single nucleotide variants in cancer samples.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
1:1 PARTIAL reproduction of the proteogenomic variant-calling pipeline on fully public data - clean independent RE-RUN after requeue (prior ROOM_RESULT lacked qc_room/datasets blocks; «infra» workdir reclaimed). ALL compute ran as SLURM jobs on «our HPC» COMPUTE nodes (no login-node compute): 2218574 RNA align+call (n109, 59m, read1=read2=113,399,653 exact, 90.11% mapped), 2218575 SnpEff annotate, 2218576 WGS overlap. Pipeline: (RNA) HISAT2/GRCh37 -> bcftools call -mv -> filter DP>=2,MQ>=10 -> SnpEff GRCh37.75; (WGS) authors' deposited Jurkat GATK PASS calls (Zenodo 400615) -> RefSeq-exon overlap -> SnpEff -> RNA overlap. RESULT: 5 of 6 attempted pipeline claims reproduce WITHIN ~10% despite a fully substituted toolchain (Tophat1.4->HISAT2, samtools-mpileup->bcftools, ANNOVAR->SnpEff): RNA non-synonymous 8233 vs 7284 (+13%), RNA stop-gain 225 vs 211 (+7%), WGS exonic 73774 vs 79601 (-7.3%), WGS nonsyn/stopgain 11320 vs 10852 (+4.3%), and the RNA/WGS overlap FRACTION 65.1% vs 64.1% (near-exact). The ONE mismatch is the raw RNA total-SNV count (389062 vs 42471, ~9x), loosely specified and dominated by low-confidence non-coding calls; even DP>=10 gives 132824 (3x) - modern bcftools call -mv is far more permissive than the 2013 samtools mpileup. Values reproduce the prior run essentially bit-for-bit (C1 389062 vs 389064; C2/C3/C5/C6/C7 identical) -> deterministic, genuine. NO fabrication indicators: the annotation-defined numbers are independently re-derivable from public data to within ~10% and the cross-dataset overlap fraction matches almost exactly. NOT ATTEMPTED: SAAV peptide counts (378/1048) + downstream iBAQ/phospho (un-released custom DB-build script + multi-GB MaxQuant; docs_insufficient), and gffread C4 (no pinned numeric value). The earlier screening drop (no_code, 'gffread is just a converter') is OVERTURNED per rule P16. All grades provisional; a human auditor decides.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 73assessed: 2026-06-16 ⛓ 8e5f2d7cc211
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether mass spectrometry-based proteomics (and phosphoproteomics) data can be used to prioritise and functionally annotate single nucleotide variants (SNVs) detected by RNA-seq/WGS in cancer cells, using Jurkat leukaemia cells as a model.
- ★ A customised SAAV peptide database built from RNA-seq/WGS variant calls can be used to search proteomics data and detect single amino acid variant (SAAV)-containing peptides at the protein level method
- ★ Deeper (ultra-deep) proteomics data substantially increases the number of SAAV-containing peptides identified compared to standard-depth proteomics finding
- ★ RNA-seq-derived SNV databases outperform WGS-derived SNV databases for SAAV peptide detection because RNA-seq coverage is enriched in highly expressed, abundantly translated transcripts finding
- ★ Including SAAV-containing peptides in protein quantification (iBAQ) modestly but significantly improves correlation between protein abundance and mRNA expression finding
- ★ A stop-gain SAAV truncates YTHDC1 at G589, removing an arginine-rich RNA-binding region, identifying it as a candidate somatic driver mutation affecting splicing in Jurkat cells finding
- ★ Phosphoproteomics data can directly identify SNVs that create novel phosphorylation sites or kinase recognition motifs, not just enable detection of existing phosphopeptides finding
- ★ A novel somatic SAAV (A85T) creates a new phosphorylation site at T86 in the splicing factor SF3B1, a gene frequently mutated in leukaemia finding
- SNVs detected as SAAV peptides by mass spectrometry have significantly higher RNA-seq read depth than SNVs not detected, indicating detection bias toward highly expressed genes finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA-seq | Jurkat cell line | none | SNV calls / gene expression (mRNA levels) | — |
| Whole genome sequencing (WGS) | Jurkat cell line | none | SNV calls, variant allele frequency, read coverage | — |
| Proteomics (LC-MS/MS, Sheynkman dataset) | Jurkat cell line | none | SAAV-containing peptide identification, protein abundance (iBAQ) | — |
| Ultra-deep proteomics (LC-MS/MS, Mertins dataset) | Jurkat cell line | none | SAAV-containing peptide identification, protein abundance | — |
| Phosphoproteomics (LC-MS/MS, Mertins dataset) | Jurkat cell line | none | SAAV-containing phosphopeptide identification, phosphorylation site/motif creation | MaxQuant |
| Manual tandem mass spectrum inspection | Jurkat cell line | none | Confirmation of YTHDC1 truncated peptide and SF3B1 phosphopeptide identity | — |
- ▲ 1,048 SAAV-containing peptides (929 unique SAAVs) identified in the ultra-deep Mertins dataset vs 378 peptides (358 unique SAAVs) in the Sheynkman dataset ~3-fold increase in unique SAAVs (920 vs 328 comparable set)
- ▼ WGS-based custom SAAV database identified fewer SAAVs in Mertins proteomics data than the RNA-seq-based database 686 vs 930 SAAVs
- ▼ SNVs unique to the RNA-seq database had significantly lower WGS read coverage (median 3 reads) than SNVs common to both databases p<0.0001, Mann-Whitney test
- ▲ Correlation between updated iBAQ protein abundance (including SAAV peptides) and mRNA expression increased for the 318 SAAV-containing proteins R=0.6978 to R=0.7154, p<0.0001
- – A stop-gain SAAV truncating YTHDC1 at G589 was detected in both proteomics datasets; YTHDC1 shows significant enrichment of stop-gain mutations in COSMIC 14.1% (24/170) stop-gain, p<0.0001, chi-square test
- – 462 SAAV-containing phosphopeptides identified; 24 (5.2%) resulted from SAAV creation of a phosphorylation site or motif 462/41,698 spectra (1.11%); 24/462 (5.2%)
- – SF3B1 A85T SAAV creates a novel phosphorylation site at T86, confirmed by manual spectrum examination
- ▲ SNVs detected as SAAV peptides had significantly higher RNA-seq read depth than undetected SNVs median 194 vs 42 reads, p<0.001, Mann-Whitney test
- correlation R=0.6632 (abundance correlation between reference and corresponding SAAV-containing peptides)
- correlation R=0.6825 to R=0.6830 (p<0.01) (protein-mRNA correlation before/after including SAAV peptides, all proteins)
- correlation R=0.6978 to R=0.7154 (p<0.0001) (protein-mRNA correlation before/after including SAAV peptides, SAAV-containing proteins only)
- pvalue p<0.0001, Mann-Whitney test (WGS coverage of RNA-seq-unique vs shared SAAV-associated SNVs)
- pvalue p<0.0001, Fisher's exact test (reference peptide detection rate: heterozygous vs homozygous SAAV peptides)
- count 42,471 SNVs total; 7,284 non-synonymous; 211 stop-gain (RNA-seq variant calling in Jurkat cells)
- fold_change ~500% (over 2 million) more spectra (Mertins dataset depth vs Sheynkman dataset depth)
- count 1,010 SAAVs detected by MS represent 13.9% (1,010/7,284) of non-synonymous/nonsense SNVs (proportion of RNA-seq SNVs confirmed at protein level)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This computational proteogenomics study integrates RNA-seq, whole-genome sequencing, and mass spectrometry datasets from the Jurkat cell line to identify and validate single amino acid variant (SAAV) containing peptides and novel phosphorylation events. The primary analyses are descriptive (counts, overlaps, percentages) with inferential statistics used selectively to validate biological expectations (e.g., that higher RNA-seq coverage predicts MS detectability, and that heterozygous loci more often show both alleles). Results are reported chiefly as counts, proportions, and Pearson correlation coefficients, with p-values given as thresholds rather than exact values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Mann-Whitney U (one-tailed direction implied by context) | Comparison of WGS read coverage for SAAVs unique to RNA-seq database vs. SAAVs common to both RNA-seq and WGS databases (Figure 2C) | Not explicitly stated per group; derived from overlap analysis of 930 RNA-seq SAAVs vs 686 WGS SAAVs | not stated |
| Mann-Whitney U | Comparison of RNA-seq read depth of SNVs with identified SAAV peptides vs. unidentified SNVs (Figure 3B); reported medians 194 vs. 42 reads | 1,010 identified SAAVs vs. remaining SNVs from 7,284 non-synonymous/nonsense total | not stated |
| Fisher's exact test | Comparison of proportion of SAAVs with detectable reference peptide: heterozygous (355/516) vs. homozygous (7/274) | 516 heterozygous and 274 homozygous SAAV-containing peptides | not stated |
| Williams' (1959) t-test for correlated correlations | Comparison of Pearson R before vs. after inclusion of SAAV peptides in iBAQ quantification, all 6,881 proteins (R = 0.6823 vs. 0.6830) | n = 6,881 proteins | not stated |
| Williams' (1959) t-test for correlated correlations | Same correlation comparison restricted to 318 proteins containing SAAVs (R = 0.6978 vs. 0.7154) | n = 318 proteins | not stated |
| Chi-square test | Comparison of proportion of stop-gain mutations in YTHDC1 (24/170, 14.1%) vs. genome-wide background across all COSMIC genes | 170 COSMIC samples with point mutations in YTHDC1; genome-wide background not numerically specified | not stated |
-
Multiple inferential tests (Mann-Whitney, Fisher's exact, chi-square, Williams' t-test) were applied across distinct comparisons without a multiplicity correction↳ Could also: A Bonferroni or Benjamini-Hochberg FDR correction could be applied across the family of six tests — With six independent tests, the probability of at least one spurious finding under the null grows; applying even a Bonferroni correction (α/6 ≈ 0.008) would still leave all reported comparisons significant given p < 0.0001 for most, while making the error-rate assumption explicit
-
Correlation coefficients before and after SAAV inclusion were compared using Williams' (1959) t-test, which tests two correlations sharing one variable in the same sample↳ Could also: Steiger's z-test (1980) is another widely used method for comparing two dependent correlations from the same sample; both are acceptable and give similar conclusions — Reporting which test was chosen and why (e.g., shared-variable structure) aids reproducibility; both approaches are standard, and noting the choice explicitly helps readers select the same method for replication
-
RNA-seq read depth distributions for identified vs. unidentified SNVs were compared with a Mann-Whitney U test and summarised with medians and 10–90th percentile bars↳ Could also: A receiver-operating characteristic (ROC) analysis or logistic regression with read depth as a predictor of MS detectability could also quantify the relationship — These approaches would additionally provide an effect-size measure (AUC or odds ratio with CI) describing how predictive read depth is of SAAV detection, complementing the significance test
-
The proportion of stop-gain mutations in YTHDC1 was compared with the genome-wide COSMIC background using a chi-square test↳ Could also: A one-sample binomial test (or Fisher's exact test if small expected counts exist) could also be used, directly testing whether the observed count of 24/170 exceeds the genome-wide stop-gain rate as a fixed reference proportion — The binomial test does not rely on the large-sample normal approximation of chi-square and is exact; it is also more transparent when the reference rate is treated as a known constant rather than estimated from data
-
Pearson correlation (R) was used to summarise the relationship between protein abundance (iBAQ) and mRNA expression↳ Could also: Spearman rank correlation could also be used, or Pearson on log-transformed values, since mRNA and protein abundance data are typically right-skewed — Spearman correlation is less sensitive to outliers and distributional assumptions, which is relevant when abundance values span many orders of magnitude as in proteomics/RNA-seq data
-
The entire analysis is based on a single cell line (Jurkat), with no biological replication across independent samples or cell lines↳ Could also: Extending the pipeline to additional cell lines or patient-derived samples, even as a small validation set, would also be a standard approach in proteogenomics studies — A multi-sample design would allow assessment of how generalizable the SAAV detection rate and phosphorylation findings are beyond Jurkat, and would enable formal statistical comparisons across biological units
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-29221171
Paper: Ma S, Menon R, Poulos RC, Wong JWH. Proteogenomic analysis prioritises functional single nucleotide variants in cancer samples. Oncotarget 2017;8(56):95841-95852. PMID 29221171 · PMCID PMC5707065 · DOI 10.18632/oncotarget.21339.
Overview of the pipeline (from Methods, PMC full text)
The study builds a sample-specific variant protein database for the Jurkat leukaemia T-cell line, then searches public Jurkat mass-spectrometry proteomics data against it to find peptides carrying single-amino-acid variants (SAAVs), i.e. SNVs that are expressed at the protein level — the "functional" prioritisation.
Pipeline stages and the tools named in the paper:
- RNA-seq variant calling — Jurkat RNA-seq (GEO GSE45428; SRA runs
SRR791578–SRR791586). Tophat v1.4.0 alignment →
samtools mpileupvariant calling. Filter: drop SNVs with <2 supporting reads or mapping quality <10. → reported 42,471 total SNVs; 7,284 non-synonymous; 211 stop-gain. - WGS variant calling — Jurkat WGS (SRA SRP101994; Zenodo 400615), ~60×. Variant calls filtered "PASS", overlapped with RefSeq exons. → reported 79,601 exonic SNVs; 10,852 non-synonymous/stop-gain; overlap with RNA-seq SNVs 4,674 (64.1 %).
- Reference proteome construction — gffread translates the RefSeq GTF transcripts into protein sequences (the named GitHub code artifact, github.com/gpertea/gffread). ANNOVAR annotates SNVs against RefSeq.
- SAAV database construction — a custom Python script maps SAAVs from the ANNOVAR-annotated SNVs and emits theoretical tryptic peptides → variant DB.
- MS database search — MaxQuant v1.3.5.8 two-stage search (stage 1: RefSeq
proteome+contaminants; stage 2: unidentified spectra vs SAAV DB), 10 ppm,
FDR 0.01, carbamidomethyl(C) fixed; against two public Jurkat MS datasets:
- Sheynkman deep proteome — PeptideAtlas PASS00215 (~0.5M spectra)
- Mertins ultra-deep + phospho — MassIVE MSV000078509 (~2.5M + ~0.85M spectra) → reported SAAV peptides: 378 peptides / 358 unique SAAVs (Sheynkman); 1,048 peptides / 929 unique SAAVs (Mertins); WGS-search 686 SAAVs. Phospho: 462 SAAV phosphopeptide spectra; 24 create phospho-sites; 5 somatic.
In scope (pipeline-derived, attemptable)
| # | Reported result | Pipeline | Data | Feasibility |
|---|---|---|---|---|
| C1 | 42,471 total RNA-seq SNVs | align + samtools mpileup + filter | GSE45428 (public, SRA) | PRIMARY TARGET — fully public, runnable on «our HPC» |
| C2 | 7,284 non-synonymous RNA-seq SNVs | + RefSeq annotation (ANNOVAR/VEP) | GSE45428 + RefSeq | feasible (annotator substitution) |
| C3 | 211 stop-gain RNA-seq SNVs | + RefSeq annotation | GSE45428 + RefSeq | feasible |
| C4 | gffread translates RefSeq GTF → N protein seqs | gffread (named repo) | RefSeq GTF+genome | feasible, fast (the literal code artifact, P16) |
| C5 | 79,601 exonic WGS SNVs | WGS variant calling | SRP101994 / Zenodo 400615 | heavier (60× WGS ~900M reads); attempt after C1–C3 |
Out of scope (not a clean pipeline reproduction here)
- SAAV peptide counts (378/1048) and phospho counts (462/24/5): depend on the
custom (un-shipped) Python SAAV-DB script + MaxQuant search of multi-GB raw MS
data. The mapping script is not released (no repo/supplement for it), so the
database-build step cannot be reproduced faithfully —
docs_insufficientfor that stage. Recorded but not graded as 1:1. - iBAQ correlation improvements (R 0.6823→0.6830 etc.), Williams' t-test p-values: downstream of the un-shipped DB-build + MS search.
- Candidate-mutation biology (YTHDC1 G589*, SF3B1 A85T): manual/curatorial, literature + COSMIC/ExAC lookups — out of scope (non-pipeline).
Notes on faithfulness / substitutions (declared up front)
- Tophat v1.4.0 (2012) is not installable from current bioconda (only 2.1.1). Variant counts are dominated by reference build + mpileup params, not the
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.