Implementing the reuse of public DIA proteomics datasets: from the PRIDE database to Expression Atlas.
The main results reproduced, with only marginal, non-material deviations.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
IN PROGRESS. Data descriptor reanalysing 10 public human SWATH-MS datasets via a Nextflow/OpenSWATH/PyProphet/TRIC/MSstats pipeline, deposited to Expression Atlas (E-PROT-59..73). Reproduction route is strong and well-described: the repo ships per-dataset MSstats .rda intermediates plus the authors' own counting script (.CMDL_MSstats_rda_count.R) whose embedded comments give exact expected protein/peptide counts (PXD004873=3530, PXD004691=2872, PXD014943=5946, PXD003497=2754, PXD004589=3703, PXD014194=2239). Plan: download .rda to «infra», run MSstats count on «our HPC», compare to paper Fig.2/Table. Full upstream re-run (raw .wiff, ~19,000 CPU core-h) is OUT of scope. Currently BLOCKED on «our HPC» VPN tunnel (TCP connect to jump host stalls) — waiting for central fix, retrying. Offline artifacts (scope.md, claims.tsv) written.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-19 ⛓ 71caa4246130
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-19
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan public Data Independent Acquisition (DIA/SWATH-MS) proteomics datasets be systematically reanalysed with a generic, automated open pipeline and the resulting protein expression robustly integrated into Expression Atlas? The paper tests whether well-annotated, iRT-spiked public DIA data can be reused in a generic context and reproduce the original studies' results.
- ★ An open, containerised, Nextflow-orchestrated reanalysis pipeline combining metadata annotation, SWATH-MS analysis, statistical analysis, and Expression Atlas integration was developed for public DIA data. method
- ★ This is the first time public DIA data has been systematically reanalysed and integrated with gene expression information in Expression Atlas. resource
- ★ Ten public human DIA datasets (1,278 SWATH-MS runs) from PRIDE were reanalysed using a generic spectral library and integrated into Expression Atlas. finding
- ★ The reanalysis reproduces the global measurement-variance characteristics (CV) and protein abundances of the original studies, demonstrating robustness. finding
- ★ A generic Combined Assay Library (CAL) plus known iRT peptide spike-in enables protein quantification yielding protein counts in the expected range for mammalian tissues/cells at 1% protein FDR. finding
- The pipeline applies a consistency filter removing proteins/groups with >50% of target features missing within MS runs across a study group. method
- A 1% global protein FDR threshold is a good trade-off; 0.1% FDR proved impractical (plasma dataset analysis failed to complete). method
- Per-MS-run correlations between reanalysed and original results were less pronounced, highlighting inherent difficulties of DIA reanalysis. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| SWATH-MS (DIA) proteomics reanalysis | 10 public human datasets (cancer tissues/cell lines including prostate, breast, hepatocellular, lymphoma, NCI-60 cell lines, kidney, liver S9, plasma) | none (reanalysis of public data) | number of detected/quantified proteins and protein intensities | OpenSWATH/pyProphet; Nextflow; CAL spectral library (SWATHAtlas SAL00031) |
| DDA-based spectral library generation (target library) | pooled/representative human samples | none | peptide spectral library for targeted DIA extraction | Combined Assay Library (CAL); plasma-targeted library for Plasma dataset |
| Differential expression statistical analysis | datasets PXD000672, PXD004691, PXD014943 | none (condition comparison per original design) | number of differentially expressed proteins | R / MSstats (median normalisation, top3 inference) |
| Technical reproducibility / coefficient of variation analysis | PXD014194 (breast cancer), PXD004873 (hepatocellular), PXD003497 (prostate regions) technical replicates | none | median CV of protein quantitative values | — |
| Correlation analysis of protein abundances | technical replicate pairs of PXD003497, PXD004873, PXD014194 | none | Pearson correlation (R) of log2 protein intensities | — |
- – Reanalysed 10 public DIA datasets comprising 1,278 SWATH-MS runs 1,278 runs
- – Total processing ~19,000 CPU core hours (~2,200 per dataset, ~15 per SWATH-MS run) 19,000 CPU core hours
- ▲ NCI-60 cell-line dataset PXD003539 yielded the most detected proteins at 1% FDR 7,097 proteins
- ▲ Lymphoma dataset PXD014943 yielded high protein count at 1% FDR 5,946 proteins
- ▲ Liver S9 dataset PXD010912 detected proteins at 1% FDR 4,224 proteins
- ▼ Plasma dataset PXD001064 detected few proteins, as expected 207 proteins
- ▲ Technical-replicate correlation between reanalysed and equivalent original results was high R = 0.9–0.98, p ≤ 0.001
- – Median CV of reanalysed data closely matched original data, overall below 21% e.g. 16.8% vs 19.0% (PXD014194)
- correlation R = 0.9 to 0.98 (p ≤ 0.001) (Pearson correlation, reanalysis vs original protein intensities for technical replicates)
- correlation R = 0.52 to 0.84 (per-MS-run reanalysis-versus-original protein expression correlation)
- count 1,278 (total SWATH-MS runs reanalysed across 10 datasets)
- count 2,754 / 2,872 / 3,703 proteins (prostate cancer datasets PXD003497 / PXD004691 / PXD004589 at 1% FDR)
- count 3,530 proteins (hepatocellular cancer PXD004873 at 1% FDR)
- count 2,239 proteins (breast cancer tumour dataset PXD014194 at 1% FDR)
- mean 16.8 vs 19.0; 7.14 vs 6.05; 20.8 vs 20.3 (median CV% reanalysed vs original (PXD014194, PXD004873, PXD003497))
- count ~19,000 CPU core hours total (~2,200 per dataset, ~15 per run) (computational cost of reanalysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a methods/resource paper describing a reanalysis pipeline for public SWATH-MS (DIA) proteomics datasets, with results validated against the original publications rather than a single hypothesis-driven experiment. Protein identification confidence was controlled via FDR/q-value thresholds (pyProphet), reproducibility of technical replicates was assessed with coefficients of variation (CV) and Pearson correlation, and downstream differential expression analysis for three datasets was performed with R/MSstats using default settings (except median normalisation and 'top3' protein inference). Results were reported primarily as protein counts, median CVs, and Pearson correlation coefficients with an associated p-value threshold, compared descriptively to the original studies' reported values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson product-moment correlation | correlation of reanalysed vs. originally published protein abundances across technical replicate MS-run pairs (Fig. 4; Supplementary Fig. S3) | technical replicate MS-run pairs per dataset (exact n not stated in text) | not stated |
| MSstats differential expression analysis (default linear-model-based group comparisons) | differential expression analysis for three datasets (PXD000672, PXD004691, PXD014943), Fig. 5 | not explicitly stated (based on MS runs per study group) | not stated |
| FDR/q-value thresholding (pyProphet global protein-level q-values) | protein identification confidence across all ten reanalysed datasets at 5%, 1%, and 0.1% FDR (Fig. 2) | all detected protein groups per dataset | not stated |
-
Concordance between reanalysed and originally published protein abundances was assessed using Pearson correlation coefficients.↳ Could also: Bland-Altman agreement analysis or a concordance correlation coefficient (e.g., Lin's CCC) — These approaches directly quantify agreement (including systematic bias) between two measurement methods, complementing a correlation coefficient which mainly captures linear association.
-
Reproducibility of technical replicates was summarised using the coefficient of variation (CV) and displayed as violin plots.↳ Could also: Reporting SD or a 95% confidence interval alongside CV — Absolute dispersion measures (SD, CI) give readers a directly comparable sense of variability magnitude in the original measurement units, complementing the normalised CV metric.
-
Differential expression analysis for three datasets was performed with MSstats using its default linear-model-based statistical framework.↳ Could also: limma or DEqMS with empirical Bayes variance moderation — Empirical Bayes moderation borrows information across proteins to stabilise variance estimates, which can be particularly useful when the number of replicates per group is small, as is common in proteomics studies.
-
Protein identification FDR was controlled at the pyProphet global protein level (5%, 1%, 0.1% thresholds).↳ Could also: Explicitly reporting the multiple-testing correction (e.g., Benjamini-Hochberg FDR) applied to the MSstats differential expression p-values themselves — Since identification FDR and differential-expression testing are separate statistical stages, stating the correction method used specifically for the differential expression p-values would make the control of false positives at that stage fully transparent.
-
The correlation between reanalysed and original MS-run-paired protein values was tested with a single overall p-value threshold (p ≤ 0.001).↳ Could also: A paired non-parametric test such as the Wilcoxon signed-rank test on matched values — A paired non-parametric comparison can test for systematic differences between reanalysis and original results without relying on the normality assumptions underlying Pearson correlation.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35701420
Paper: Walzer et al. (2022) Implementing the reuse of public DIA proteomics datasets: from the PRIDE database to Expression Atlas. Sci Data 9:335. DOI 10.1038/s41597-022-01380-9 · PMCID PMC9197839. Repo: https://github.com/PRIDE-reanalysis/DIA-reanalysis (authors' own code — P16 own-repo).
What kind of paper
A data descriptor. The authors built a Nextflow pipeline that re-analyses 10 public human SWATH-MS / DIA proteomics datasets from PRIDE with a single harmonised workflow, and deposits the harmonised quantitative results into EBI Expression Atlas (E-PROT-59,60,66,67,68,69,70,71,72,73). There is essentially no wet-lab work — all reported numbers are pipeline-derived, so almost the whole paper is in scope in principle. The cost barrier is compute volume, not method opacity.
Pipeline (as described in Methods + repo)
Upstream (nextflows/pridereswath_upstream.nf): wiffConverter 0.7.0 → Yamato QC
1.0.4 → OpenSWATH 2.4.0 (git 868546e) → PyProphet 2.0.dev1 (git ddcedac) →
TRIC/msproteomicstools 0.8.0 (git eeed765). Spectral library = Combined Assay
Library SAL00031 (1,164,312 transitions / 139,449 peptides / 10,316 proteins).
Downstream (nextflows/pridereswath_downstream.nf): TRIC tsv + annotation →
DIA_downstream_process.R (MSstats dataProcess, summaryMethod=TMP,
featureSubset="top3", normalization handled downstream, MBimpute=TRUE) → .rda
→ report / protein counting (.CMDL_MSstats_rda_count.R, Protein_numbers.R).
Consistency filter = keep proteins with <50% missing features per study group.
Expression-Atlas integration maps UniProtKB→Ensembl (MyGene.info, Ensembl 99).
Datasets reanalysed (all human SWATH-MS)
PXD004873(HCC,76 runs,E-PROT-69) · PXD000672(kidney,48,E-PROT-59) · PXD004691(prostate,224,E-PROT-68) · PXD014943(DLBCL,113,E-PROT-67) · PXD003497(prostate,60,E-PROT-66) · PXD004589(prostate,210,E-PROT-70) · PXD014194(breast,145,E-PROT-72) · PXD003539(NCI-60,120,E-PROT-73) · PXD001064(plasma,240,E-PROT-60) · PXD010912(liver S9,42,E-PROT-71). Total 1,278 runs.
IN SCOPE — pipeline-derived results we attempt
- Per-dataset reanalysed protein counts at 1% FDR (top3) — the headline
Table / Fig.2 numbers (Total distinct proteins): PXD004873=3530, PXD004691=2872,
PXD014943=5946, PXD003497=2754, PXD004589=3703, PXD014194=2239, plus PXD000672,
PXD003539, PXD001064, PXD010912. Reproduction route: the repo ships the
per-dataset MSstats
.rdaintermediates (EBI S3) —load()→ count distinctRunlevelData$Proteinlevels..CMDL_MSstats_rda_count.Ris the authors' own counting script and its embedded comments give the exact expected values (3530, 2872, 5946, 2754, 3703, 2239 …) — a self-contained 1:1 target. - Consistency-filtered protein counts (<50% missing per group) — also emitted
by the same script (e.g. PXD004873 → 3392). Same
.rda, same route. - Expression Atlas end-product protein counts (E-PROT-xx) — download the deposited baseline protein-expression matrices, count proteins, cross-check they equal the reanalysed counts (e.g. E-PROT-69 ↔ 3530). Zero-compute, verifies the deposit delivers what the paper promises.
- Peptide / iRT counts per dataset — same
.rda(ProcessedData$PEPTIDElevels; iRT peptides). Expected values embedded in the count script.
HARDER (attempt after the floor)
- Technical-replicate Pearson R (0.90–0.98) and median CV (<21%) —
DIA_postprocess_correlation.R,DIA_postprocess_variation.R. Need normalised exports; feasible from the.rda. - Differential-expression overlaps (e.g. PXD014943 118 vs 97, overlap 21;
PXD000672 262 vs 613, overlap 59; FC correlation R=0.52 / 0.84) —
DIA_postprocess_differential.R; needs original supplementary tables (shipped ininputs/original_results_from_supplementaries/).
OUT OF SCOPE (not attempted)
- Full upstream re-run (raw .wiff → OpenSWATH → PyProphet → TRIC): ~19,000 CPU core-hours over
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.