Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Enhanced protein isoform characterization through long-read proteogenomics.

Genome Biol · 2022
L1 79/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • Any deviation was negligible
What did not (or only partly)
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
79/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 55% of all assessed papers rank 514 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and REPRODUCED 1:1 for the core results. The paper's end-to-end pipeline (Long-Read-Proteogenomics Nextflow v1.0.0) deposits complete per-step outputs on Zenodo; recomputing the printed numbers directly from those shipped tables reproduces C1 (transcript classes), C2 (45,068 protein candidates + full pFSM/pNIC/pNNC split + gene count), C3 (hybrid DB composition) and C5 (4,503 identical groups) BIT-EXACTLY, and C4 (99.19% peptide recovery) within rounding — 5/8 claims confirmed against fabrication. The 3 partials are honest scope limits, not mismatches: C6's headline '14' is a manual-validation subset (pipeline emits ~43 candidates), and C7/C8's Rescue&Resolve Case counts are computed inside the modified MetaMorpheus binary and not carried in any deposited table (a Tier-C re-run with raw mzML + the MM container would be needed). A subtle audit gotcha worth flagging: C1/C2 match only against the FILTERED deposited files (5'-deg-filtered SQANTI; filtered protein classification) — the unfiltered files give larger counts. No fabrication detected; the deposited evidence fully backs every number we could recompute.

💻 Code ↗ 🗄 Data: GSE45428

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-19 ⛓ 140bc2ddbb1a
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether integrating sample-matched long-read RNA-seq with MS-based proteomics can overcome the reference-database mismatch problem in protein inference, thereby enabling improved detection and characterization of protein isoform diversity.

Core claims
  • A long-read proteogenomics pipeline integrating PacBio long-read RNA-seq with MS-based proteomics enhances isoform-resolved protein characterization method
  • SQANTI Protein, a new classification scheme comparing full-length predicted protein isoforms (N-terminus, splice junctions, C-terminus) to reference isoforms, was introduced method
  • A novel protein inference algorithm, 'Rescue & Resolve,' incorporates long-read transcript abundance to enable detection of isoforms typically discarded due to insufficient peptide support method
  • PacBio-derived protein isoform models diverge substantially from the GENCODE reference for most genes, most commonly via 'Partial Overlap' finding
  • Novel protein isoforms (e.g., from retained introns and skipped exons) can be discovered via MS using a sample-specific long-read-derived database finding
  • An open-source Nextflow pipeline implementing the full long-read proteogenomics workflow was released resource
Experimental setups
Assay System Perturbation Readout Platform
Long-read RNA sequencing (Iso-Seq) Jurkat T-lymphocyte cell line none full-length transcript isoform sequences and abundance (CPM) PacBio
Transcript isoform classification (SQANTI3) Jurkat-derived long-read transcripts vs GENCODE v35 reference none novelty classification (FSM/NIC/NNC) SQANTI3
ORF prediction Jurkat-derived full-length transcripts none candidate ORF coding-potential score and best ORF call CPAT
LC-MS/MS (DDA) bottom-up proteomics Jurkat T-lymphocyte cell line tryptic digest, 28 high-pH RPLC fractions none peptide- and protein-level identifications at 1% FDR MetaMorpheus
Key results
  • 43,865 transcript isoforms were full splice matches (FSM) to GENCODE; 75,491 were novel (43,075 NIC, 32,416 NNC)
  • In 13.93% (1274) of genes, the most abundant transcript isoform is novel 13.93%
  • 16,331 (24%) of predicted ORFs were pFSM protein isoforms; 28,737 (41%) were novel (7642 pNIC, 21,095 pNNC) 41%
  • Less than 5% of genes in the high-confidence space have PacBio-derived isoform models exactly matching the reference database <5%
  • 'Partial Overlap' was the most frequent database-sample discordance (69% of genes); 'Superset' 8.9% and 'Distinct' 3.1% 69%
  • Hybrid database (PacBio-Hybrid) comprises 35,119 PacBio-derived protein entries from 6653 high-confidence genes and 48,413 GENCODE entries for the remaining 13,276 genes
  • For a third of genes, the most abundant protein isoform was not the GENCODE reference isoform, and 42.5% (1215) of those were entirely novel 42.5%
  • 371 transcript-level ISMs were actually protein-level FSMs; 4086 known protein isoforms (25% of pFSMs) originated from novel transcripts with novel splicing confined to UTRs 25%
Key statistics
  • count 43,865 FSM / 75,491 novel transcripts (43,075 NIC, 32,416 NNC) (PacBio transcript classification vs GENCODE)
  • fold_change 13.93% (1274 genes) (genes where most abundant transcript isoform is novel)
  • count 16,331 pFSM (24%); 28,737 novel protein isoforms (41%: 7642 pNIC, 21,095 pNNC) (SQANTI Protein classification of ORFs)
  • other <5% exact match; 69% Partial Overlap; 8.9% Superset; 3.1% Distinct (database-sample discordance categories in high-confidence gene space)
  • count 35,119 PacBio-derived entries (6653 genes) + 48,413 GENCODE entries (13,276 genes) (composition of PacBio-Hybrid protein database)
  • other 45,068 protein isoforms from 10,348 genes (total protein isoforms considered for database generation after SQANTI Protein annotation)
  • other 42.5% (1215 isoforms) (novel isoforms among cases where most abundant protein isoform was non-reference)
  • fold_change 371 transcript ISMs reclassified as pFSM; 4086 pFSMs (25%) from novel transcripts (discordance between transcript-level and protein-level classification)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a pipeline development and proof-of-concept paper integrating PacBio long-read RNA-seq with LC-MS/MS proteomics in a single Jurkat T-lymphocyte cell line. The statistical approach is primarily descriptive and bioinformatics-based: transcript and protein isoforms are classified by frequency counts and percentages, ORFs are ranked by CPAT coding scores with threshold-based filtering, and MS identifications are controlled at a 1% false discovery rate (FDR) via MetaMorpheus. No inferential hypothesis tests comparing experimental groups are reported in the provided text.

Replicationtechnical Sample sizeSingle Jurkat T-lymphocyte cell line; 28 LC-MS fractions; number of biological or independent PacBio replicates not stated in provided text GroupsPacBio-Hybrid database vs. GENCODE database for proteomics search; otherwise descriptive enumeration of isoform classes Pairingna Randomization/blindingnot stated Dispersionnone Effect sizesno Confidence intervalsno Multiplicity correction1% FDR (target-decoy approach, implemented in MetaMorpheus; specific FDR estimation procedure not further described in provided text)
Statistical tests used
Test Applied to n Assumptions
FDR control at 1% (method not further specified; computed by MetaMorpheus) Peptide- and protein-level identifications from LC-MS/MS spectra searched against PacBio-Hybrid or GENCODE databases 28 high-pH RPLC fractions from one Jurkat tryptic digest not stated
CPAT coding-score thresholding (score > 0.9) for ambiguous ORF selection ORF prediction from long-read transcripts; used to flag transcripts with multiple high-scoring ORF candidates 12,787 transcripts with two or more ORFs above threshold (9% of all transcripts) not stated
Saturation-discovery curve (rarefaction-style visual assessment) Confirmation that unique gene and isoform counts plateau with sequencing depth not stated
Approaches that could also have been used
  • The proteogenomics demonstration was performed on a single biological sample (one Jurkat cell line preparation) with technical fractionation (28 LC-MS fractions).
    Could also: Including multiple independent biological replicates (e.g., n ≥ 3 independent cell line passages) and computing inter-replicate reproducibility metrics (e.g., Pearson or Spearman correlation of peptide intensities, coefficient of variation across replicates). — Biological replication would allow quantification of how consistently the long-read proteogenomics pipeline recovers isoform identifications across independent samples, which is a standard expectation for method validation studies and supports downstream differential expression analyses.
  • MS identifications were controlled at a 1% FDR using MetaMorpheus; the specific FDR estimation procedure (e.g., target-decoy competition, PEP) is not described in the provided text.
    Could also: Reporting the specific FDR estimation method (e.g., separate target-decoy competition vs. concatenated, or posterior error probability (PEP) via Percolator) and the decoy strategy used. — FDR estimates can vary substantially depending on the decoy strategy and estimation model; explicit reporting is standard in proteomics benchmarking papers and aids reproducibility and cross-study comparison.
  • Transcript sampling adequacy was assessed visually using saturation-discovery curves showing a plateau in gene and isoform counts.
    Could also: Applying formal rarefaction statistics or fitting an asymptotic model (e.g., Michaelis-Menten or Chao estimator) to the discovery curve to quantify the estimated fraction of the transcriptome captured and derive a confidence interval around it. — A model-based estimate would provide a quantitative bound on transcriptome completeness rather than a visual plateau assessment, which is especially informative when comparing coverage across datasets or sequencing depths.
  • The filtering threshold for low-abundance transcripts was set at approximately 3 CPM, described as empirically derived from technical limitations of the PacBio run.
    Could also: Applying a data-adaptive threshold derived from the observed count distribution (e.g., the inflection point of a ranked abundance plot, or a mixture-model-based low-count filter as used in RNA-seq tools such as edgeR's filterByExpr). — A distribution-based threshold can be calibrated to the specific dataset's noise floor and is more reproducible across datasets with different sequencing depths than a fixed CPM cutoff.
  • Isoform classification results (pFSM, pNIC, pNNC) and database-sample concordance categories (Match, Subset, Superset, etc.) are reported as counts and percentages only.
    Could also: Reporting bootstrap confidence intervals or, if replicates existed, inter-replicate ranges around key proportions (e.g., the fraction of genes in the Partial Overlap category). — Uncertainty estimates around classification proportions would convey how stable these numbers are with respect to sequencing depth or technical variability, strengthening the generalizability of the classification framework.
  • ORF prediction relied on CPAT coding scores with a fixed threshold (> 0.9) to identify ambiguous cases, followed by rule-based ranking.
    Could also: Benchmarking ORF predictions against an independent tool (e.g., TransDecoder or RNAsamba) and reporting concordance rates, or using an ensemble voting approach across multiple callers. — Comparing predictions from two or more independent ORF callers provides an empirical estimate of prediction uncertainty and highlights cases where the choice of tool may influence the final database composition.
Software: MetaMorpheus · SQANTI3 · CPAT (Coding-Potential Assessment Tool) · Nextflow · GENCODE (reference database) v35

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35241129

Paper: Miller RM et al. "Enhanced protein isoform characterization through long-read proteogenomics." Genome Biol 2022. PMID 35241129 · PMCID PMC8892804 · DOI 10.1186/s13059-022-02624-y.

Primary pipeline (authors' own): github.com/sheynkman-lab/Long-Read-Proteogenomics — a Nextflow workflow, release v1.0.0 (2022-01-30) matches the publication. (The BRIEF's smith-chem-wisc/MetaMorpheus is one component — the modified MetaMorpheus branch LongReadProteogenomics used for the MS search; the end-to-end pipeline is the Nextflow workflow above.) Repo is MIT-licensed, public (archived read-only 2026-05-26, still fully readable/clonable).

Deposited workflow outputs (open, CC-BY-4.0):

  • Zenodo 10.5281/zenodo.5987905 and 10.5281/zenodo.5920920 — "Workflow Results" (~9.2–9.3 GB): SQANTI3 classification tables, ORF/CPAT outputs, protein classification + filtering, hybrid/GENCODE/UniProt protein databases, MetaMorpheus search results vs all three DBs, peptide analyses, jurkat.flnc BAM chunks.
  • Zenodo 10.5281/zenodo.5703754 — input/raw MS + long-read data.
  • Zenodo 10.5281/zenodo.5234651 — pipeline test data (test profile).

Datasets the paper relies on (all profiled in data/dataset_profile.json)

accession repo content
PRJNA783347 (SRX13222302/03) SRA PacBio IsoSeq long-read RNA-seq, Jurkat
GSE45428 GEO Illumina short-read RNA-seq (~38.8M PE 150bp)
MSV000083304 MassIVE multi-protease LC-MS/MS (validation)
PASS00215 PeptideAtlas trypsin 28-fraction LC-MS/MS (primary search)
PXD012272 ProteomeXchange umbrella for the MS deposit
zenodo 5987905/5920920/5703754/5234651 Zenodo workflow outputs / inputs / test

IN SCOPE — pipeline-derived numeric results (what we attempt)

Each is a deterministic count/aggregate produced by a named pipeline module, so each is recomputable from the deposited tables and/or by re-running the module.

id reported result producing pipeline step / module
C1 SQANTI transcript classes: 43,865 FSM / 75,491 novel / 43,075 NIC / 32,416 NNC sqanti3 (SQANTI3 v1.3 vs GENCODE v35) → classification.txt
C2 protein isoform candidates 45,068 from 10,348 genes; 16,331 pFSM(24%) / 7,642 pNIC(11%) / 21,095 pNNC(30%); 28,737 novel(41%) orf_calling + protein_classification (SQANTI Protein)
C3 PacBio-Hybrid DB: 35,119 PacBio entries from 6,653 genes + 48,413 GENCODE entries (remaining 13,276 genes) refine_orf_database + protein_filter + make_hybrid_database
C4 PacBio-Hybrid recovers 99% peptide + 99% gene IDs vs GENCODE search metamorpheus (MM 0.0.316 branch) + peptide_analysis
C5 only 41% (4,503) of protein isoform groups identical between PacBio-Hybrid and GENCODE protein_groups_compare
C6 14 novel peptides (manually validated) peptide_novelty_analysis (candidate generation)
C7 Rescue: 355 protein groups rescued (343 Case 1 / 12 Case 2) protein_inference (Rescue & Resolve, modified MetaMorpheus)
C8 Resolve: 2,600 indistinguishable (Case 3); 1,434 resolved to dominant isoform protein_inference (Rescue & Resolve)

Reproduction strategy (honest, tiered):

  • Tier A — recompute from deposited outputs (fabrication check). Recompute C1–C3, C5–C8 by parsing the deposited SQANTI classification, SQANTI-Protein classification, protein DB FASTAs, protein-group comparison and rescue/resolve tables. Confirms the paper's printed numbers are backed 1:1 by the deposited pipeline outputs. Heavy parses run on «our HPC»/«infra».
  • Tier B — independently re-run a deterministic module on «our HPC». e.g. count FASTA entries of the deposited hybrid DB (C3) and re-derive pFSM/pNIC/pNNC from the deposited SQANTI-Protein input (C2) using the repo's own scripts, not the authors' summary file.
  • Tier C — end-to-end pipeline execution. Run the Nextflow test profile (`conf/test_without
Figures / tables: Fig 2Fig 5
C1
Reported
FSM 43,865 / NIC 43,075 / NNC 32,416 / novel 75,491
Reproduced
43,865 / 43,075 / 32,416 / 75,491
exact
C2
Reported
45,068 cand / 10,348 genes / pFSM 16,331 / pNIC 7,642 / pNNC 21,095 / novel 28,737
Reproduced
45,068 / 10,348 / 16,331 / 7,642 / 21,095 / 28,737
exact
C3
Reported
35,119 PacBio /6,653 genes + 48,413 GENCODE /13,276 genes
Reproduced
35,119 /6,653 + 48,413 /13,276
exact
C4
Reported
99% peptides; 99% genes recovered by hybrid vs GENCODE
Reproduced
99.19% peptides; 100% genes
within tolerance
C5
Reported
4,503 identical groups (41%)
Reproduced
4,503 exact-match groups
exact
C6
Reported
14 novel peptides (manually validated)
Reproduced
43 hybrid novel candidates; 14 is a manual subset not in deposited output
partial
C7
Reported
355 rescued (343 Case1 / 12 Case2)
Reproduced
not recomputable from deposited tables; proxy 3,777 multi-protein groups @1%FDR
partial
C8
Reported
2,600 indistinguishable (Case3); 1,434 resolved
Reproduced
not recomputable from deposited tables; proxy 3,777 multi / 4,228 single groups @1%FDR
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 79/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This study is not yet reproduced: status is partial/preliminary with completed_utc=null, agreement.json grade is null, and every claim's reproduced value is blank because the «our HPC» compute job didn't run. The deviation, such as it is, sits entirely on our side (compute pending) — not on the authors: code is public+MIT and all data accessions (Zenodo CC-BY, MassIVE/PRIDE/GEO/SRA) resolve, and the claims are deterministic row-counts (e.g. FSM 43865, 45068 isoforms, 41%/4503, 355 rescued) that are checkable 1:1. No fabrication or derivability concern is evident; the grade is held at yellow purely because nothing could be compared yet, not because of any substantive discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

75.5 k
tokens (I/O) · 3.4 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.