Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Spatially resolved phosphoproteomics reveals fibroblast growth factor receptor recycling-driven regulation of autophagy and survival.

Nat Commun · 2022
L1 87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1 on the core). Watson & Ferguson 2022 APEX2 spatial phosphoproteomics. The Zenodo deposit is a mirror of the GitHub repo (no separate raw data; raw MS on PRIDE, out of scope). Re-ran the authors' R scripts 1-5 on the shipped MaxQuant tables on «our HPC» (conda R 4.5.3). 8/13 claims EXACT, 4 within-tol (~98%), 1 consistent, 1 blocked. The significant phosphosite set + hierarchical clusters reproduce bit-for-bit (1653 sites; RE 732 / RAB11 532 / other 389; id-set Jaccard 1.000), as do global-upregulated sites (477), total sites (10771), the mTOR/autophagy network (73 nodes/158 edges), and the paper's Fig 4f overlap (107). Paper Fig 4e overlaps reproduce at ~98% (945 vs 961 sites; 577 vs 588 proteins; 733 vs 743 RE; the recycling-cluster PERCENTAGE 77.6% vs 77.4% is near-exact). The only non-trivial gap is the proximal-protein significance set (count exact at 3303, but ~14% identity drift), traced entirely to QRILC imputation stochasticity + a newer imputeLCMD; this propagates to the Fig 4e counts. Found & fixed a fatal undefined-variable bug in script 1 (lfq_apex_normGFP->lfq_normGFP). NOT attempted: raw MaxQuant processing (PRIDE), the manual Perseus volcano (its output ships, so downstream IS reproduced), and script 6 (BLOCKED - inputs not deposited; figure-only). No value appears fabricated.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.7197969

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 87
    assessed: 2026-06-22 ⛓ 1488669e1713
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates which FGFR2b signalling partners are recruited in close proximity to recycling endosomes during FGF10-induced receptor recycling, and how this spatially restricted signalling shapes downstream cellular responses such as autophagy and survival.

Core claims
  • A spatially resolved phosphoproteomics (SRP) approach combining APEX2-driven proximity biotinylation with phosphopeptide enrichment was developed to identify FGFR2b signalling partners near recycling endosomes. method
  • FGF10-stimulated FGFR2b activates mTOR-dependent signalling and ULK1 at the recycling endosomes, leading to autophagy suppression and cell survival. finding
  • Inhibiting FGFR2b trafficking (via DnRAB11 or DnDNM2) alters the global phosphoproteome without changing FGFR2b or ERK1/2 activation. finding
  • mTOR signalling is specifically enriched among proteins in the recycling-dependent (RAB11-dependent) phosphorylation response cluster, common to HeLa-FGFR2b and T47D cells. finding
  • The FGFR2b recycling route, rather than mere cytoplasmic presence of the receptor, regulates specific branches of downstream signalling. finding
  • APEX2 tagging of FGFR2b, RAB11, or GFP does not alter FGFR2b trafficking or FGF10-induced signalling activation. method
  • RAB11-APEX2-based pulldown enriches known recycling endosome markers (RAB25, RCP) and HA-FGFR2b, validating proximity to recycling endosomes. finding
  • FGF10 induces FGFR2b recycling to the plasma membrane via RAB11-positive recycling endosomes, whereas FGF7 induces FGFR2b degradation. finding
Experimental setups
Assay System Perturbation Readout Platform
Confocal immunofluorescence microscopy (colocalization) HeLa_FGFR2bST cells expressing wtRAB11-, DnRAB11-, or DnDNM2-eGFP FGF10 stimulation (0, 40, 120 min); dominant-negative RAB11/Dynamin2 FGFR2b colocalization with EEA1 and GFP-tagged trafficking markers
Western blot / immunoblot HeLa_FGFR2bST cells expressing GFP, DnRAB11, or DnDNM2 FGF10 stimulation (8, 40 min) FGFR2b and ERK1/2 phosphorylation
MS-based quantitative phosphoproteomics and proteomics HeLa-FGFR2b and T47D cells expressing GFP, DnRAB11, or DnDNM2 FGF10 stimulation (40 min); trafficking-blocking dominant negatives Phosphorylated site abundance, PCA, fuzzy c-means clustering, KEGG pathway enrichment Mass spectrometry
APEX2-based spatially resolved phosphoproteomics (proximity biotinylation + phosphopeptide enrichment) HeLa_FGFR2b-APEX2ST, T47D_FGFR2KO_FGFR2b-APEX2ST, HeLa-FGFR2bST_RAB11-APEX2, HeLa-FGFR2bST_GFP-APEX2 FGF10 stimulation; biotin-phenol + H2O2 labelling (1 min) Biotinylated and phosphorylated proteins after streptavidin pulldown Mass spectrometry
Streptavidin pulldown + immunoblot HeLa-FGFR2bST_RAB11-APEX2 FGF10 stimulation, APEX2 biotinylation RAB25, HA-FGFR2b, RCP levels in pulldown
Confocal-based trafficking/internalisation assay HeLa_FGFR2b-APEX2ST and T47D_FGFR2KO_FGFR2b-APEX2ST FGF7 vs FGF10 stimulation (up to 120 min) FGFR2b subcellular localisation (plasma membrane, cytoplasm, degradation)
Key results
  • DnRAB11 and DnDNM2 expression impaired FGFR2b trafficking but did not alter FGFR2b or ERK1/2 phosphorylation at early time points
  • 7620 phosphorylated sites quantified in HeLa-FGFR2b and 8075 in T47D
  • Fuzzy c-means clustering identified 11 clusters of phosphosites significantly dysregulated across the four conditions ANOVA p<0.0001
  • KEGG pathway over-representation analysis identified mTOR signalling as enriched specifically in the recycling response cluster in both HeLa-FGFR2b and T47D
  • No signalling pathways were specifically enriched upon inhibition of FGFR2b internalisation alone (DnDNM2)
  • FGF10 induced FGFR2b to gradually leave the cell surface, accumulate in cytoplasm, and recycle back to the plasma membrane; FGF7 induced internalisation followed by degradation
  • RAB11-APEX2 pulldown was positive for RAB25, HA-FGFR2b, and RCP, confirming proximity labelling at recycling endosomes
  • Phosphorylated PLCγ and SHC, but not histone H3, were detected in APEX2 pulldowns after FGF10 treatment
Key statistics
  • pvalue ANOVA p<0.0001 (Significance threshold for phosphosites used in fuzzy c-means clustering)
  • count 7620 phosphorylated sites (Sites quantified in HeLa-FGFR2b phosphoproteome)
  • count 8075 phosphorylated sites (Sites quantified in T47D phosphoproteome)
  • pvalue p<0.0005 (one-sided student's t-test) (Colocalization quantification of FGFR2b with GFP-tagged proteins/EEA1)
  • count N=3 independent biological replicates (Confocal colocalization quantification (Fig. 1b))
  • count N≥3 independent biological replicates (Immunoblot analysis (Fig. 1c))
  • other 20 nm (APEX2) vs 10 nm (BioID) labelling radius (Proximity-dependent biotinylation radius comparison)
  • count 11 clusters (Number of phosphosite clusters identified by fuzzy c-means clustering)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combines quantitative MS-based phosphoproteomics/proteomics across engineered cell lines (HeLa-FGFR2b, T47D) with confocal imaging quantification and immunoblotting to characterize FGFR2b trafficking-dependent signalling. Group comparisons in imaging used a one-sided Student's t-test, phosphoproteomic profiles across four conditions were compared by ANOVA followed by fuzzy c-means/t-SNE clustering, and pathway-level enrichment (KEGG) was assessed with Fisher's exact test plus FDR adjustment. Results are reported largely as significance thresholds (e.g., p<0.0005, p<0.0001) alongside median ± SD summaries for imaging quantification.

Replicationbiological Sample sizeN = 3 independent biological replicates (2-5 cells analyzed per N) for co-localization quantification; N ≥ 3 independent biological replicates for immunoblots; no explicit power/sample-size calculation described in this excerpt GroupsFGFR2b trafficking states (wtRAB11, DnRAB11, DnDNM2, GFP control) under FGF10/FGF7 stimulation over time, across HeLa-FGFR2b and T47D cell lines Pairingunclear Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionFDR adjustment
Statistical tests used
Test Applied to n Assumptions
One-sided Student's t-test Quantification of FGFR2b co-localization with GFP-tagged proteins and with EEA1 (Fig. 1b) N = 3 independent biological replicates, 2-5 cells analyzed per N not stated
ANOVA Identification of phosphorylated sites significantly dysregulated across the four HeLa-FGFR2b trafficking conditions prior to fuzzy c-means clustering (Fig. 2c, d) not explicitly stated (based on MS runs across experimental conditions) not stated
Fisher's exact test with FDR adjustment KEGG pathway over-representation analysis comparing membrane, internalisation, and recycling response clusters (Fig. 2f) not stated not stated
Pearson correlation Comparison of cellular proteome across experimental conditions to check that dominant-negative protein expression did not alter the proteome (Supplementary Fig. 2a, j) not stated not stated
Approaches that could also have been used
  • Co-localization quantification (Fig. 1b) was compared using a one-sided Student's t-test on a small number of biological replicates (N=3) with multiple cells per replicate.
    Could also: A mixed-effects (hierarchical) model treating cell as nested within biological replicate — This would explicitly account for non-independence among the 2-5 cells measured per replicate, which can otherwise affect the effective sample size used in the significance calculation.
  • A one-sided t-test was used for the co-localization comparison.
    Could also: A two-sided t-test — A two-sided test would also be a standard choice unless the direction of the expected difference was specified in advance, and can be reported alongside the one-sided result for readers who prefer non-directional hypothesis testing.
  • Phosphorylated sites differentially regulated across the four trafficking conditions were identified using ANOVA prior to fuzzy c-means clustering.
    Could also: A moderated statistics framework such as limma, commonly used in proteomics/phosphoproteomics — Moderated variance estimation can improve stability of significance calls when the number of replicates per condition is small, which is common in MS-based proteomics experiments.
  • KEGG pathway enrichment used Fisher's exact test with FDR adjustment.
    Could also: Bonferroni correction or gene set enrichment analysis (GSEA) — Bonferroni offers a more conservative family-wise error control, while GSEA-type approaches use the full ranked phosphoproteomic dataset rather than a fixed significance cutoff, which can capture pathway-level trends among sub-threshold sites.
  • Imaging quantification results were summarized as median ± SD.
    Could also: Reporting with a 95% confidence interval alongside or instead of SD — A CI directly conveys the precision of the estimated median/mean and is often preferred for communicating uncertainty with small replicate numbers (N=3).
  • Proteome consistency across conditions was assessed using Pearson correlation.
    Could also: Reporting exact p-values or effect-size statistics (e.g., correlation coefficient with CI) alongside the correlation values — This would let readers directly evaluate the strength and precision of the reported similarity between conditions rather than relying on a qualitative description of 'high' correlation.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig 4eFig 4dFig 4fFig 5j
C1_sty_sig
Reported
1653 sig proximal phosphosites
Reproduced
1653 (id-set Jaccard 1.000)
exact
C2_sty_RE
Reported
732 (RE_profile cluster)
Reproduced
732
exact
C3_sty_RAB11
Reported
532 (RAB11_profile cluster)
Reproduced
532
exact
C4_pro_sig
Reported
3303 sig proximal proteins
Reproduced
3303 (gene-set Jaccard 0.862)
within tolerance
C5_global_upreg
Reported
477 FGF10-upregulated global sites @40min
Reproduced
477 (Jaccard 1.000)
exact
C6_total_sty
Reported
10771 total global phosphosites
Reproduced
10771 (Jaccard 1.000)
exact
C7_phos_prots_sites
Reported
961 overlap phosphosites (Fig 4e)
Reproduced
945
within tolerance
C8_phos_prots_proteins
Reported
588 overlap proteins (Fig 4e)
Reproduced
577
within tolerance
C9_RE_fraction
Reported
743 (77.4%) overlap in RE cluster (Fig 4e)
Reproduced
733 (77.6%)
within tolerance
C10_fig4f_overlap
Reported
107 global-recycling overlap (Fig 4f)
Reproduced
107
exact
C11_mtor_network
Reported
73 nodes / 158 edges mTOR network
Reproduced
73 nodes / 158 edges (Jaccard 1.000)
exact
C12_apex_tag_effect
Reported
APEX2 tag does not affect global quantification (Suppl Fig 5j)
Reproduced
ctrl t-test 2 sites; FGF10 ANOVA 109 sites <=0.05 (minority)
m.public.grade.consistent
C13_script6
Reported
endosomal-marker heatmaps (figure)
Reproduced
BLOCKED - input files not deposited
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

249.3 k
tokens (I/O) · 13.4 M incl. cache
46 min
runtime · 0.05 CPU-h
1.8 GB
peak RAM
2
HPC jobs
hummel
machine