Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

An integrative proteomics method identifies a regulator of translation during stem cell maintenance and differentiation.

Nat Commun · 2021
L1 92/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
92/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 83% of all assessed papers rank 179 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1 on the auditable pipeline outputs). ProteoTracker (Sabatier et al., Nat Commun 2021, PMID 34772928). 5 of 6 in-scope pipeline claims reproduced EXACTLY: passed-protein counts C1=7778 (two independent deposits), C2=9565, C3=5478; C4 microarray direction (SBDS lower in iPSC than hFF) confirmed both from the deposited TAC table AND by an independent raw-CEL oligo-RMA+limma run on «our HPC» (SLURM «job», logFC +1.60, adjP 4e-7); C6 the ProteoTracker Fisher/sector statistical core reproduced to machine precision (max diff 3.3e-16 over 155,560 rows; passed/failed 90702/64858 exact; >75% two-transition fraction = 78.6%). C5 (155/9565 SBDS-KD) PARTIAL: the permutation-FDR code that defines the 155 is not shipped, so the exact count is not derivable from the deposit. No fabrication indicators - every reproduced number is exactly recoverable from the shipped data/code. NOT attempted (out of scope): wet-lab assays (polysome profiling, Seahorse, qRT-PCR, KD phenotypes), full MaxQuant re-search from raw .raw (the deposited search result is the audited pipeline product), and full TPP melting-curve refit. The pre-filter 'quantified' totals (9509/11451/9000) live only in raw MS, not in any deposit.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 92
    assessed: 2026-06-21 ⛓ 4cb003f35570
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether combining protein thermal stability/solubility measurements with expression proteomics in a single integrated assay (PISA-Express) can reveal pathway- and complex-level protein property changes—particularly in ribosome biogenesis—that regulate pluripotent stem cell maintenance and differentiation, and whether modulating ribosome maturation via SBDS can be used to control stemness in vitro.

Core claims
  • PISA-Express is a method that simultaneously measures protein expression and thermal stability (solubility) changes using only two samples per replicate per cell type method
  • ProteoTracker is a web-based visualization tool that maps protein trajectories across cell-type transitions using Sankey diagrams without dimension reduction resource
  • Modulating ribosome maturation through the SBDS protein can be used to manipulate cell stemness in vitro finding
  • Protein thermal stability distinguishes cell types independently from protein expression finding
  • Fold-changes in protein stability and expression are not significantly correlated for individual proteins, indicating these are independent analytical dimensions finding
  • Ribosome-related proteins show the largest and most significantly enriched trajectory changes (highest mean distance, top GO enrichment) during PSC differentiation finding
  • Metabolic pathway reprogramming (e.g., glycolysis, oxidative phosphorylation) precedes DNA repair and chromatin remodeling changes during early PSC differentiation finding
  • PISA-Express requires only 2 samples per replicate versus 10 samples typically required for standard thermal proteome profiling (TPP) method
Experimental setups
Assay System Perturbation Readout Platform
PISA (proteome-wide integral solubility alteration) / thermal proteome profiling with TMT multiplexing hi12 iPSC, H9 ESC, EB, hFF, and RKO cells none (comparison across cell types/differentiation states) protein thermal stability (Sm parameter) from melting-curve-like solubility across a narrow temperature range TMT10/TMT11 labeling, LC-MS/MS
Expression proteomics (37°C fraction of PISA-Express samples) hi12 iPSC, H9 ESC, EB, hFF, and RKO cells none (comparison across cell types/differentiation states) protein expression (Exp parameter, TMT reporter ion ratio to pooled linker sample) TMT10/TMT11 labeling, LC-MS/MS
Expression proteomics with SDS-based extraction for increased depth/coverage iPSC and ESC differentiating into embryoid bodies (EBs) none (differentiation time course) protein expression, expanded proteome coverage (11,451 proteins quantified) LC-MS/MS
Expression proteomics HT29 cells and iPSC(hi12)-derived neurons, plus hFF, hi12 iPSC, and EB none (cell type comparison) protein expression LC-MS/MS
Cellular reprogramming human foreskin fibroblasts (hFF) reprogrammed into iPSCs reprogramming factors generation and validation of iPSC line
Key results
  • 9509 proteins quantified for both stability and expression across three biological replicates; 7778 passed selection criteria (≥2 unique peptides, no missing values)
  • iPSC protein stability and expression correlated strongly with ESC but not with hFF, RKO, or EB r=0.85 (stability), r=0.86 (expression)
  • No significant correlation observed between fold-changes in stability versus expression for individual proteins
  • Protein trajectory transition group B-to-A (stabilized/downregulated to stabilized/upregulated) showed the highest mean trajectory distance and strongest GO enrichment, all top terms related to ribosome
  • More than 75% of the quantified proteome significantly changed in expression or stability during the iPSC→EB→hFF transitions >75%
  • ESC (female) had the lowest number of stability/expression outliers versus iPSC; hFF had similar stability outliers but more expression outliers than RKO versus iPSC
  • Validation dataset quantified 11,451 proteins (9565 passing criteria); additional HT29/neuron datasets covered 9000 proteins (5478 passing criteria)
Key statistics
  • correlation r=0.85 (iPSC vs ESC protein thermal stability (Sm) correlation)
  • correlation r=0.86 (iPSC vs ESC protein expression (Exp) correlation)
  • count 9509 proteins (proteins with stability and expression data across 3 replicates)
  • count 7778 proteins (proteins passing selection criteria (≥2 unique peptides, no missing values))
  • count 11,451 proteins quantified; 9565 passed criteria (SDS-based validation expression dataset during iPSC/ESC to EB differentiation)
  • count 9000 proteins quantified; 5478 passed criteria (HT29 and iPSC-derived neuron expression datasets)
  • pvalue <0.05 (combined Fisher p-value threshold for assigning significant sectors A–D to protein trajectories)
  • count 25 PTTGs (total protein trajectory transition groups analyzed by GO enrichment)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper introduces PISA-Express, a multiplexed proteomics method that simultaneously quantifies protein expression (Exp) and thermal stability (Sm, via the PISA assay) across five human cell types using TMT10/TMT11 labeling with n = 3 biological replicates per cell type. Cell-type differences were characterized by PCA of both dimensions, fold-change calculations relative to iPSC, and Fisher's combined p-value formula to jointly assess stability and expression changes for each protein. Altered proteins were grouped into trajectory sectors and submitted to Gene Ontology (GO) enrichment analysis, with results visualized using Sankey diagrams in the ProteoTracker web tool.

Replicationbiological Sample sizen = 3 biologically independent samples per cell type; stated in Fig. 1a legend; no formal power calculation mentioned Groupshi12 iPSC vs. H9 ESC, EB, hFF, RKO; additionally iPSC→EB→hFF trajectory analysis Pairingunpaired Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Principal Component Analysis (PCA) Separation of five cell types by protein stability (Sm) and expression (Exp) — Fig. 2a, b 7778 proteins, n=3 biological replicates per cell type not stated
Fisher's combined p-value formula Combining per-protein p-values for stability and expression changes to assign protein trajectories to sectors A–D vs. E; threshold p < 0.05 7778 proteins not stated
Pearson correlation coefficient Correlation of iPSC protein stability and expression values against ESC, hFF, RKO, EB — Supplementary Fig. 4c, d not stated (dataset-wide) not stated
Standard-deviation-based outlier detection (>3 mean SDs) Identifying proteins with statistically extreme fold changes in Sm and Exp across cell types — Fig. 2c 7778 proteins not stated
Gene Ontology (GO) enrichment analysis Each protein trajectory transition group (PTTG) — 25 groups total — tested against background of all quantified proteins variable per PTTG; background = all quantified proteins not stated
Fold change (log2 FC) calculation Per-protein Sm and Exp ratios relative to iPSC across all cell-type comparisons — Fig. 2c, d, 3d–g average of 3 biological replicates na
Approaches that could also have been used
  • Protein outliers were defined by exceeding three mean standard deviations of fold change across the dataset
    Could also: A moderated t-test or empirical Bayes approach (e.g., limma) could be applied per protein across the three replicates, producing per-protein p-values and FDR-adjusted q-values — With only n = 3 replicates, per-protein variance estimates are noisy; limma-style shrinkage borrows strength across proteins to stabilize variance, which can improve sensitivity and specificity compared to a global SD threshold
  • Fisher's formula was used to combine separate p-values for thermal stability and expression changes into a single per-protein score
    Could also: Stouffer's Z-score method or the Cauchy combination test could also combine the two p-values — Fisher's method assumes independence between the combined p-values; Stouffer's allows weighting by sample size or reliability of each dimension, and the Cauchy combination is robust to dependence between the two tests, which may be relevant given that stability and expression are measured on the same sample
  • No multiple-testing correction was described for the 7778 per-protein Fisher combined p-values tested at a single p < 0.05 threshold
    Could also: Benjamini–Hochberg FDR correction could be applied across all protein-level tests — Controlling the false discovery rate is standard practice when testing thousands of hypotheses simultaneously; it provides a transparent bound on the expected proportion of false positives among proteins assigned to sectors A–D
  • GO enrichment analysis was applied to each of 25 protein trajectory transition groups (PTTGs) without a stated correction for multiple GO terms or multiple PTTGs
    Could also: FDR correction (e.g., Benjamini–Hochberg) within each PTTG, or a combined family-wise correction across all PTTGs, could be applied — Testing many GO terms across 25 groups generates a large multiple-comparison family; an explicit FDR threshold would help distinguish enrichments that are likely to replicate from those that could arise by chance
  • Dispersion in violin plots was summarized with median and interquartile range (25th/75th percentile box, 5th/95th percentile whiskers)
    Could also: Mean ± SD or 95% bootstrap confidence intervals on fold changes could additionally be reported — With n = 3 replicates per cell type, the 5th–95th percentile range reflects the full protein distribution rather than replicate uncertainty; reporting replicate-level dispersion (SD or CI) would more directly communicate reproducibility of the measurements
  • Cell-type separation was visualized and interpreted qualitatively from PCA plots without formal statistical testing of group separation
    Could also: Permutation-based multivariate tests such as PERMANOVA (adonis) on the full protein matrix could formally test whether cell types are significantly separated in multivariate stability or expression space — PERMANOVA provides a p-value for group separation that accounts for the multivariate structure and sample size, complementing the visual interpretation of PCA plots
Software: ProteoTracker (web tool, genexplain.com) · LC-MS/MS (instrument platform, specific software not named)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34772928 (ProteoTracker / Sabatier et al., Nat Commun 2021)

Title: An integrative proteomics method identifies a regulator of translation during stem cell maintenance and differentiation. PMID 34772928 · PMCID PMC8590018 · DOI 10.1038/s41467-021-26879-4

Data & code deposits (verbatim from Data/Code availability)

  • GEO GSE135409 — microarray RNA (Affymetrix HTA 2.0), hFF vs iPSC, 6 samples. OPEN.
  • ProteomeXchange/PRIDE PXD018453 — PISA-Express MS data.
  • ProteomeXchange/PRIDE PXD014830 — TPP data (hi11/12/13, hFF, RKO in cells; hi12+hFF lysate TPP; ribosomal protein expression hi12+RKO; protein-expression hFF/hi12/HT29/ Neurons/EBs/RKO; SBDS-KD protein expression in hFF; SBDS-KD after EB induction in hi12).
  • ProteomeXchange/PRIDE PXD015874 — protein-expression of EB induction in H9 & HS980.
  • Zenodo 10.5281/zenodo.5018241 — TPP statistics (derived output, TPP R pkg v3.13).
  • Code: GitHub RZlab/ProteoTracker + Zenodo 10.5281/zenodo.5549677 (Shiny web interface).
  • Supplementary Datasets 1–4 + Source Data (xlsx) attached to the paper.

What the computational pipelines are

  1. MaxQuant 1.6.2.3 (Andromeda) TMT10/11, UniProt human 2019_05 (73,910 entries), FDR 0.01, ≥7-aa peptides, 2 missed cleavages → protein identification/quantification.
  2. Selection filter: keep proteins with ≥2 unique peptides and no missing values in any replicate/cell type → the "passed" protein sets.
  3. PISA-Express: per protein compute stability parameter Sm and expression FC Exp vs iPSC; significance per sector via Fisher combined p-value < 0.05 → ProteoTracker Sankey sectors A–E.
  4. TPP R package v3.13 (Franken et al.) → thermal stability / melting behaviour.
  5. Microarray DE: CEL → (TAC/RMA) → nonpaired one-way ANOVA + Benjamini–Hochberg, adjusted-p < 0.01, fold-change ≥ 2.0.
  6. ProteoTracker Shiny app (R) — visualization of Sm/Exp trajectories.

IN SCOPE — pipeline-derived results we attempt

ID Reported result (paper loc) Pipeline Reproduction approach Weight
C1 9509 proteins quantified; 7778 passed (≥2 uniq pep, no missing) — Results "PISA-Express", Suppl. Data 1 MaxQuant + filter Count rows in deposited Suppl. Data 1 / PRIDE proteinGroups; re-apply filter core
C2 11,451 quantified; 9565 passed (SDS expt) — Results, Suppl. Data 2 MaxQuant + filter Count + filter Suppl. Data 2 core
C3 5478 passed across all cell types (~9000 quantified) — Results, Suppl. Data 2 MaxQuant + filter Count + filter Suppl. Data 2 core
C4 Microarray: SBDS mRNA lower in iPSC than hFF (Suppl. Data 4); DEG at adj-p<0.01 & FC≥2 RMA + ANOVA/BH Full raw→DE on GSE135409 CEL files on «our HPC»; report DEG count + SBDS direction/FC core (only public raw→result chain)
C5 155 / 9565 proteins significantly altered in SBDS-KD vs control; 2 false positives in 57,390 permutations stat test + permutation Recompute from deposited SBDS-KD expression table harder
C6 ProteoTracker sector A–E assignment via Fisher combined p<0.05 (Fig 2e); count proteins per sector ProteoTracker R logic Re-run sector assignment from repo's shipped data / PISA tables harder

OUT OF SCOPE — wet-lab / manual / instrument (not attempted)

  • Polysome profiling (Fig 4a–b); Seahorse metabolic flux; qRT-PCR.
  • SBDS-knockdown phenotypes: ~50% protein-synthesis reduction (Fig 5d), OCT4/NANOG mRNA upregulation — wet-lab measurements.
  • TMT labelling, cell culture, EB induction, electroporation, MS acquisition.
  • Full MaxQuant re-search from raw .raw files: technically possible on «our HPC» but extremely heavy (TMT, GB-scale raw, version/DB-sensitive). We reproduce the protein- count outputs by auditing the deposited MaxQuant output instead (the deposited search result IS the pipeline product); a full re-search is noted as a possible stretch goal.
Figures / tables: Fig 2
C1
Reported
9509 proteins quantified; 7778 passed selection (>=2 unique peptides, no missing values) - main PISA-Express, 5 cell types
Reproduced
7778 passed (EXACT) - Suppl. Data 1 (MOESM4 Table 1) has 7778 rows AND repo ProteoTracker_data.tsv has 7778 rows; 9509 'quantified' is a pre-filter MaxQuant count not present in any deposited table (raw re-search out of scope)
exact
C2
Reported
11,451 proteins quantified; 9565 passed (SDS/EB-induction expression experiment)
Reproduced
9565 passed (EXACT) - Suppl. Data 2 Table 2-4 has 9565 rows; 11,451 'quantified' is pre-filter, not in deposit
exact
C3
Reported
~9000 proteins quantified; 5478 passed across all cell types
Reproduced
5478 passed (EXACT) - Suppl. Data 2 Table 2-2 has 5478 rows
exact
C4
Reported
Microarray (GSE135409, Affymetrix HTA 2.0): SBDS mRNA is present in lower quantities in iPSCs than in hFF
Reproduced
CONFIRMED two ways. Deposited Suppl. Data 4 (TAC): SBDS=TC07001467.hg.1, hFF=12.94 vs iPSC=9.96 log2, FC 7.9x, FDR 1.15e-4. Independent CEL->RMA->limma on «our HPC» (oligo+limma): hFF=9.25 vs iPSC=7.64 log2, logFC +1.60, P=7.4e-9, adjP=4.0e-7 -> iPSC lower. Direction+significance reproduced (absolute scale differs: TAC SST-RMA vs plain RMA)
exact
C5
Reported
155 of 9565 proteins significantly altered in SBDS-KD; only 2 in 57,390 permutations
Reproduced
PARTIAL - threshold sweep on Suppl. Data 2 Table 2-4 'p.value SBDS siRNA vs Scrambled siRNA': p<0.05->1232, p<0.01->428, p<0.001->90, p<0.05&FC>=1.5->51, BH<0.1->70; no simple threshold yields exactly 155. The permutation-FDR method defining the 155 (and the 57,390-permutation FP test) is not shipped in the repo, so the exact count is not reproducible from the deposit
partial
C6
Reported
ProteoTracker: Fisher combined p<0.05 sector A-E assignment; >75% of proteome significantly changed expression or stability during the two transitions (Fig 2)
Reproduced
EXACT - recomputed Fisher combined p (sumlog/df=4) from the two shipped p-values for all 155,560 protein x comparison rows; max abs diff vs shipped fisher column = 3.3e-16 (machine epsilon). passed/failed at p<0.05 = 90702/64858, matches shipped exactly. Two-transition (hFF->iPSC->EB) union of significant proteins = 6116/7778 = 78.6% (>75%, matches paper). Sector A-E assignment reproduced per server.R lines 356-360
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 92/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

298.7 k
tokens (I/O) · 19.6 M incl. cache
62 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.