Enhanced protein isoform characterization through long-read proteogenomics.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓Any deviation was negligible
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough and REPRODUCED 1:1 for the core results. The paper's end-to-end pipeline (Long-Read-Proteogenomics Nextflow v1.0.0) deposits complete per-step outputs on Zenodo; recomputing the printed numbers directly from those shipped tables reproduces C1 (transcript classes), C2 (45,068 protein candidates + full pFSM/pNIC/pNNC split + gene count), C3 (hybrid DB composition) and C5 (4,503 identical groups) BIT-EXACTLY, and C4 (99.19% peptide recovery) within rounding — 5/8 claims confirmed against fabrication. The 3 partials are honest scope limits, not mismatches: C6's headline '14' is a manual-validation subset (pipeline emits ~43 candidates), and C7/C8's Rescue&Resolve Case counts are computed inside the modified MetaMorpheus binary and not carried in any deposited table (a Tier-C re-run with raw mzML + the MM container would be needed). A subtle audit gotcha worth flagging: C1/C2 match only against the FILTERED deposited files (5'-deg-filtered SQANTI; filtered protein classification) — the unfiltered files give larger counts. No fabrication detected; the deposited evidence fully backs every number we could recompute.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-19 ⛓ 140bc2ddbb1a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether integrating sample-matched long-read RNA-seq with MS-based proteomics can overcome the reference-database mismatch problem in protein inference, thereby enabling improved detection and characterization of protein isoform diversity.
- ★ A long-read proteogenomics pipeline integrating PacBio long-read RNA-seq with MS-based proteomics enhances isoform-resolved protein characterization method
- ★ SQANTI Protein, a new classification scheme comparing full-length predicted protein isoforms (N-terminus, splice junctions, C-terminus) to reference isoforms, was introduced method
- ★ A novel protein inference algorithm, 'Rescue & Resolve,' incorporates long-read transcript abundance to enable detection of isoforms typically discarded due to insufficient peptide support method
- ★ PacBio-derived protein isoform models diverge substantially from the GENCODE reference for most genes, most commonly via 'Partial Overlap' finding
- ★ Novel protein isoforms (e.g., from retained introns and skipped exons) can be discovered via MS using a sample-specific long-read-derived database finding
- ★ An open-source Nextflow pipeline implementing the full long-read proteogenomics workflow was released resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Long-read RNA sequencing (Iso-Seq) | Jurkat T-lymphocyte cell line | none | full-length transcript isoform sequences and abundance (CPM) | PacBio |
| Transcript isoform classification (SQANTI3) | Jurkat-derived long-read transcripts vs GENCODE v35 reference | none | novelty classification (FSM/NIC/NNC) | SQANTI3 |
| ORF prediction | Jurkat-derived full-length transcripts | none | candidate ORF coding-potential score and best ORF call | CPAT |
| LC-MS/MS (DDA) bottom-up proteomics | Jurkat T-lymphocyte cell line tryptic digest, 28 high-pH RPLC fractions | none | peptide- and protein-level identifications at 1% FDR | MetaMorpheus |
- – 43,865 transcript isoforms were full splice matches (FSM) to GENCODE; 75,491 were novel (43,075 NIC, 32,416 NNC)
- – In 13.93% (1274) of genes, the most abundant transcript isoform is novel 13.93%
- – 16,331 (24%) of predicted ORFs were pFSM protein isoforms; 28,737 (41%) were novel (7642 pNIC, 21,095 pNNC) 41%
- – Less than 5% of genes in the high-confidence space have PacBio-derived isoform models exactly matching the reference database <5%
- – 'Partial Overlap' was the most frequent database-sample discordance (69% of genes); 'Superset' 8.9% and 'Distinct' 3.1% 69%
- – Hybrid database (PacBio-Hybrid) comprises 35,119 PacBio-derived protein entries from 6653 high-confidence genes and 48,413 GENCODE entries for the remaining 13,276 genes
- – For a third of genes, the most abundant protein isoform was not the GENCODE reference isoform, and 42.5% (1215) of those were entirely novel 42.5%
- – 371 transcript-level ISMs were actually protein-level FSMs; 4086 known protein isoforms (25% of pFSMs) originated from novel transcripts with novel splicing confined to UTRs 25%
- count 43,865 FSM / 75,491 novel transcripts (43,075 NIC, 32,416 NNC) (PacBio transcript classification vs GENCODE)
- fold_change 13.93% (1274 genes) (genes where most abundant transcript isoform is novel)
- count 16,331 pFSM (24%); 28,737 novel protein isoforms (41%: 7642 pNIC, 21,095 pNNC) (SQANTI Protein classification of ORFs)
- other <5% exact match; 69% Partial Overlap; 8.9% Superset; 3.1% Distinct (database-sample discordance categories in high-confidence gene space)
- count 35,119 PacBio-derived entries (6653 genes) + 48,413 GENCODE entries (13,276 genes) (composition of PacBio-Hybrid protein database)
- other 45,068 protein isoforms from 10,348 genes (total protein isoforms considered for database generation after SQANTI Protein annotation)
- other 42.5% (1215 isoforms) (novel isoforms among cases where most abundant protein isoform was non-reference)
- fold_change 371 transcript ISMs reclassified as pFSM; 4086 pFSMs (25%) from novel transcripts (discordance between transcript-level and protein-level classification)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a pipeline development and proof-of-concept paper integrating PacBio long-read RNA-seq with LC-MS/MS proteomics in a single Jurkat T-lymphocyte cell line. The statistical approach is primarily descriptive and bioinformatics-based: transcript and protein isoforms are classified by frequency counts and percentages, ORFs are ranked by CPAT coding scores with threshold-based filtering, and MS identifications are controlled at a 1% false discovery rate (FDR) via MetaMorpheus. No inferential hypothesis tests comparing experimental groups are reported in the provided text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| FDR control at 1% (method not further specified; computed by MetaMorpheus) | Peptide- and protein-level identifications from LC-MS/MS spectra searched against PacBio-Hybrid or GENCODE databases | 28 high-pH RPLC fractions from one Jurkat tryptic digest | not stated |
| CPAT coding-score thresholding (score > 0.9) for ambiguous ORF selection | ORF prediction from long-read transcripts; used to flag transcripts with multiple high-scoring ORF candidates | 12,787 transcripts with two or more ORFs above threshold (9% of all transcripts) | not stated |
| Saturation-discovery curve (rarefaction-style visual assessment) | Confirmation that unique gene and isoform counts plateau with sequencing depth | — | not stated |
-
The proteogenomics demonstration was performed on a single biological sample (one Jurkat cell line preparation) with technical fractionation (28 LC-MS fractions).↳ Could also: Including multiple independent biological replicates (e.g., n ≥ 3 independent cell line passages) and computing inter-replicate reproducibility metrics (e.g., Pearson or Spearman correlation of peptide intensities, coefficient of variation across replicates). — Biological replication would allow quantification of how consistently the long-read proteogenomics pipeline recovers isoform identifications across independent samples, which is a standard expectation for method validation studies and supports downstream differential expression analyses.
-
MS identifications were controlled at a 1% FDR using MetaMorpheus; the specific FDR estimation procedure (e.g., target-decoy competition, PEP) is not described in the provided text.↳ Could also: Reporting the specific FDR estimation method (e.g., separate target-decoy competition vs. concatenated, or posterior error probability (PEP) via Percolator) and the decoy strategy used. — FDR estimates can vary substantially depending on the decoy strategy and estimation model; explicit reporting is standard in proteomics benchmarking papers and aids reproducibility and cross-study comparison.
-
Transcript sampling adequacy was assessed visually using saturation-discovery curves showing a plateau in gene and isoform counts.↳ Could also: Applying formal rarefaction statistics or fitting an asymptotic model (e.g., Michaelis-Menten or Chao estimator) to the discovery curve to quantify the estimated fraction of the transcriptome captured and derive a confidence interval around it. — A model-based estimate would provide a quantitative bound on transcriptome completeness rather than a visual plateau assessment, which is especially informative when comparing coverage across datasets or sequencing depths.
-
The filtering threshold for low-abundance transcripts was set at approximately 3 CPM, described as empirically derived from technical limitations of the PacBio run.↳ Could also: Applying a data-adaptive threshold derived from the observed count distribution (e.g., the inflection point of a ranked abundance plot, or a mixture-model-based low-count filter as used in RNA-seq tools such as edgeR's filterByExpr). — A distribution-based threshold can be calibrated to the specific dataset's noise floor and is more reproducible across datasets with different sequencing depths than a fixed CPM cutoff.
-
Isoform classification results (pFSM, pNIC, pNNC) and database-sample concordance categories (Match, Subset, Superset, etc.) are reported as counts and percentages only.↳ Could also: Reporting bootstrap confidence intervals or, if replicates existed, inter-replicate ranges around key proportions (e.g., the fraction of genes in the Partial Overlap category). — Uncertainty estimates around classification proportions would convey how stable these numbers are with respect to sequencing depth or technical variability, strengthening the generalizability of the classification framework.
-
ORF prediction relied on CPAT coding scores with a fixed threshold (> 0.9) to identify ambiguous cases, followed by rule-based ranking.↳ Could also: Benchmarking ORF predictions against an independent tool (e.g., TransDecoder or RNAsamba) and reporting concordance rates, or using an ensemble voting approach across multiple callers. — Comparing predictions from two or more independent ORF callers provides an empirical estimate of prediction uncertainty and highlights cases where the choice of tool may influence the final database composition.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35241129
Paper: Miller RM et al. "Enhanced protein isoform characterization through long-read proteogenomics." Genome Biol 2022. PMID 35241129 · PMCID PMC8892804 · DOI 10.1186/s13059-022-02624-y.
Primary pipeline (authors' own):
github.com/sheynkman-lab/Long-Read-Proteogenomics — a Nextflow workflow,
release v1.0.0 (2022-01-30) matches the publication. (The BRIEF's
smith-chem-wisc/MetaMorpheus is one component — the modified MetaMorpheus
branch LongReadProteogenomics used for the MS search; the end-to-end pipeline
is the Nextflow workflow above.) Repo is MIT-licensed, public (archived
read-only 2026-05-26, still fully readable/clonable).
Deposited workflow outputs (open, CC-BY-4.0):
- Zenodo
10.5281/zenodo.5987905and10.5281/zenodo.5920920— "Workflow Results" (~9.2–9.3 GB): SQANTI3 classification tables, ORF/CPAT outputs, protein classification + filtering, hybrid/GENCODE/UniProt protein databases, MetaMorpheus search results vs all three DBs, peptide analyses,jurkat.flncBAM chunks. - Zenodo
10.5281/zenodo.5703754— input/raw MS + long-read data. - Zenodo
10.5281/zenodo.5234651— pipeline test data (testprofile).
Datasets the paper relies on (all profiled in data/dataset_profile.json)
| accession | repo | content |
|---|---|---|
| PRJNA783347 (SRX13222302/03) | SRA | PacBio IsoSeq long-read RNA-seq, Jurkat |
| GSE45428 | GEO | Illumina short-read RNA-seq (~38.8M PE 150bp) |
| MSV000083304 | MassIVE | multi-protease LC-MS/MS (validation) |
| PASS00215 | PeptideAtlas | trypsin 28-fraction LC-MS/MS (primary search) |
| PXD012272 | ProteomeXchange | umbrella for the MS deposit |
| zenodo 5987905/5920920/5703754/5234651 | Zenodo | workflow outputs / inputs / test |
IN SCOPE — pipeline-derived numeric results (what we attempt)
Each is a deterministic count/aggregate produced by a named pipeline module, so each is recomputable from the deposited tables and/or by re-running the module.
| id | reported result | producing pipeline step / module |
|---|---|---|
| C1 | SQANTI transcript classes: 43,865 FSM / 75,491 novel / 43,075 NIC / 32,416 NNC | sqanti3 (SQANTI3 v1.3 vs GENCODE v35) → classification.txt |
| C2 | protein isoform candidates 45,068 from 10,348 genes; 16,331 pFSM(24%) / 7,642 pNIC(11%) / 21,095 pNNC(30%); 28,737 novel(41%) | orf_calling + protein_classification (SQANTI Protein) |
| C3 | PacBio-Hybrid DB: 35,119 PacBio entries from 6,653 genes + 48,413 GENCODE entries (remaining 13,276 genes) | refine_orf_database + protein_filter + make_hybrid_database |
| C4 | PacBio-Hybrid recovers 99% peptide + 99% gene IDs vs GENCODE search | metamorpheus (MM 0.0.316 branch) + peptide_analysis |
| C5 | only 41% (4,503) of protein isoform groups identical between PacBio-Hybrid and GENCODE | protein_groups_compare |
| C6 | 14 novel peptides (manually validated) | peptide_novelty_analysis (candidate generation) |
| C7 | Rescue: 355 protein groups rescued (343 Case 1 / 12 Case 2) | protein_inference (Rescue & Resolve, modified MetaMorpheus) |
| C8 | Resolve: 2,600 indistinguishable (Case 3); 1,434 resolved to dominant isoform | protein_inference (Rescue & Resolve) |
Reproduction strategy (honest, tiered):
- Tier A — recompute from deposited outputs (fabrication check). Recompute C1–C3, C5–C8 by parsing the deposited SQANTI classification, SQANTI-Protein classification, protein DB FASTAs, protein-group comparison and rescue/resolve tables. Confirms the paper's printed numbers are backed 1:1 by the deposited pipeline outputs. Heavy parses run on «our HPC»/«infra».
- Tier B — independently re-run a deterministic module on «our HPC». e.g. count FASTA entries of the deposited hybrid DB (C3) and re-derive pFSM/pNIC/pNNC from the deposited SQANTI-Protein input (C2) using the repo's own scripts, not the authors' summary file.
- Tier C — end-to-end pipeline execution. Run the Nextflow
testprofile (`conf/test_without
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This study is not yet reproduced: status is partial/preliminary with completed_utc=null, agreement.json grade is null, and every claim's reproduced value is blank because the «our HPC» compute job didn't run. The deviation, such as it is, sits entirely on our side (compute pending) — not on the authors: code is public+MIT and all data accessions (Zenodo CC-BY, MassIVE/PRIDE/GEO/SRA) resolve, and the claims are deterministic row-counts (e.g. FSM 43865, 45068 isoforms, 41%/4503, 355 rescued) that are checkable 1:1. No fabrication or derivability concern is evident; the grade is held at yellow purely because nothing could be compared yet, not because of any substantive discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.