Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Cancer cells copy migratory behavior and exchange signaling networks via extracellular vesicles.

EMBO J · 2018
L1 46/100 3/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
46/100
Reproducibility score
1.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 6% of all assessed papers rank 1101 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (P16-valid third-party STAR->HTSeq->DESeq2 pipeline on the paper's own ENA data), INDEPENDENTLY RE-CONFIRMED BIT-FOR-BIT. Fresh «our HPC» «job» (COMPLETED 2026-06-24T12:31:06Z, 1h33m, node n093) re-ran the full pipeline self-contained (conda env-build + GRCm38 r102 ref dl + 72 ENA SE runs -> merge to 18 samples + STAR 2.7.10b index/align + HTSeq 2.0.3 union unstranded + DESeq2 1.50.2) and produced results.json BYTE-IDENTICAL to the prior run AND counts_matrix.tsv with the IDENTICAL sha256 (396d521c673aa0e8...) -> the reproduced values are genuine deterministic pipeline output. The pipeline reproduces DE/detection counts of the CORRECT order of magnitude and qualitative pattern (100K EVs >> 16.5K EVs ~ cells; B16F1>B16F10; detection ~12-15k transcripts), and several values land within ~2-13% (C2 up_raw 103 vs 105; C5b up_raw 1007 vs 1089; C4c/C4d). But exact reported integers (65/105/571; 12450/12802/11696/11527) are NOT reproduced under any single consistent threshold, and C3 (571) is a clear mismatch (nearest 899). Most plausible cause: the paper does not pin its Ensembl annotation release (we used r102, 55,487 features) and used older STAR/HTSeq builds, plus raw-vs-adjusted-P ambiguity -- NOT a fabrication signal; no value looks non-derivable from the shipped data. NOT attempted: all wet-lab/imaging/MS-proteomics results (out of scope); B16F10 EV-vs-cell contrast (C5c/C5d) not emitted by the run script. counts_matrix.tsv kept on «infra» (sha256 396d521c..., 3.35MB, 55487x18).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 46
    assessed: 2026-06-21 ⛓ 00c69c2c1ae7
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Cancer cell subclones with distinct metastatic potential but common clonal origin (B16F1 vs B16F10 melanoma) can phenocopy each other's migratory behavior through exchange of extracellular vesicle (EV) cargo, even though only subtle molecular differences drive their metastatic heterogeneity.

Core claims
  • B16F1 and B16F10 melanoma subclones functionally exchange EVs in vivo, transferring Cre mRNA cargo detectable via a Cre-LoxP color-switch reporter system finding
  • B16F1 cells that take up B16F10-derived EV cargo (eGFP+ reporter+) show higher migration speed than B16F1 cells that did not (DsRed+ reporter+), demonstrating phenocopying of migratory behavior finding
  • There is a large discrepancy between EV transfer efficiency in vitro (EV uptake occurs but <0.01% functional cargo release in co-culture) versus in vivo (functional transfer occurs), underlining the importance of studying EV exchange in the in vivo setting finding
  • EVs shed by melanoma subclones into the tumor microenvironment contain thousands of distinct proteins and RNAs belonging to interconnected signaling networks involved in cellular processes such as migration finding
  • A method combining enzymatic tumor dissociation with differential ultracentrifugation successfully isolates two EV populations (16.5K and 100K) directly from the in vivo tumor microenvironment method
  • The Cre-LoxP reporter system distinguishes functional release of EV luminal cargo into recipient cell cytoplasm from mere EV uptake method
  • 16.5K EV fraction is enriched for larger EVs (≥150 nm) and 100K EV fraction is enriched for smaller EVs (≤150 nm), both also containing melanosome-like structures finding
  • B16F10 cells have higher intrinsic migration speed and metastatic potential than B16F1 cells within the same tumor microenvironment finding
Experimental setups
Assay System Perturbation Readout Platform
intravital microscopy (cell tracking/migration speed) co-injected B16F1/B16F10 tumors in C57BL/6 mice none (mixed clone co-injection) single-cell migration speed and tracks
in vitro EV uptake assay (PKH67 labeling, confocal imaging) B16F1 and B16F10 cell lines addition of labeled 16.5K/100K EVs from reciprocal cell type EV internalization by recipient cells
Cre-LoxP reporter system, in vivo functional EV transfer locally co-injected Cre+ and reporter+ B16F1/B16F10 tumors in mice Cre+ EV donor cells vs reporter+ recipient cells DsRed-to-eGFP color switch (Cre activity) and recipient migration speed
Cre-LoxP reporter system, in vitro co-culture B16F1/B16F10 Cre+ and reporter+ cell lines, 3-week co-culture Cre+ EV donor vs reporter+ recipient co-culture DsRed-to-eGFP color switch (Cre activity)
RT-PCR cells, 16.5K EVs, and 100K EVs from B16F10 Cre+ and B16F1 Cre- tumors none Cre and RPL38 mRNA detection RT-PCR
electron microscopy 16.5K and 100K EV fractions isolated from B16F1/B16F10 tumors none EV size and morphology electron microscopy
total RNA sequencing 16.5K and 100K EVs isolated from B16F1 and B16F10 tumors none transcript detection/counts in EVs
label-free mass spectrometry (proteomics) and Western blot cells and 16.5K/100K EV fractions from B16F1/B16F10 tumors (and subcellular fractions of B16F10 cells) none protein/marker detection (tetraspanins, flotillins, ESCRT, HSPs, calnexin, cytochrome-C, GM130) and normalized MS/MS counts tandem mass spectrometry (MS/MS); Western blot
Key results
  • B16F10 cancer cells have a higher average migration speed than B16F1 cells within the same tumor area 1.5-fold
  • Micrometastases derived from B16F10 cells were found more frequently than from B16F1 cells, confirming differential metastatic potential
  • B16F1 cells that took up B16F10-derived EV cargo (eGFP+ reporter+) migrated faster than B16F1 cells that did not (DsRed+ reporter+); uptake of B16F1-derived EVs did not enhance migration speed of recipients
  • Cre+ EV cargo was functionally transferred in vivo, with reporter+ cells switching color only in tumors also containing Cre+ cells
  • In vitro co-culture of Cre+ and reporter+ cells showed almost no functional Cre-mediated color switch despite confirmed EV uptake <0.01%
  • Both 16.5K and 100K EV populations were taken up by recipient cells of the reciprocal cell type in vitro
  • Thousands of transcripts were detected in EVs across both fractions and both cell lines 11,527-12,802 transcripts
  • Thousands of proteins were detected in EVs across both fractions and both cell lines 3,210-3,333 proteins
Key statistics
  • fold_change 1.5-fold higher migration speed in B16F10 vs B16F1 (average migration speed of B16F10 vs B16F1 cells within same imaging field, 26 positions in 8 mice)
  • count 6 of 9 mice (B16F10 majority) vs 1 of 11 mice (B16F1 majority) with micrometastases (differential metastatic potential confirmed via lung/lymph node/liver micrometastases)
  • other <0.01% Cre-mediated color switch (functional Cre+ EV transfer in 3-week in vitro co-culture vs substantial in vivo transfer)
  • count 12,450 / 12,802 / 11,696 / 11,527 transcripts (transcripts detected in all 3 replicates for B16F1 16.5K/100K and B16F10 16.5K/100K EVs, respectively)
  • count 3,210 / 3,333 / 3,213 / 3,276 proteins (proteins detected in all 3 replicates for B16F1 16.5K/100K and B16F10 16.5K/100K EVs, respectively)
  • other FC ≥ 10 and P ≤ 0.01 (threshold for GO term cellular compartment enrichment of EV proteins over donor cell)
  • count n = 15 positions in 4 mice; n = 16 positions in 4 mice; n = 22 positions in 6 mice (sample sizes for intravital migration speed comparisons in Fig 2C-E)
  • other 3.2% (proportion of 16.5K EV protein(s) identified as truncated in low molecular weight gel band analysis (text truncated))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper uses a mouse xenograft/intravital-microscopy design comparing two melanoma subclones (B16F1, B16F10) and a Cre-LoxP reporter system to quantify EV-mediated transfer, alongside RNA-seq and label-free mass spectrometry to profile EV cargo. Group comparisons of migration speed and Cre+ EV transfer are reported as mean ± SEM (or SD for proteomic markers) with nonparametric tests (Wilcoxon signed-rank, Mann-Whitney) and one Student's t-test noted; sample sizes are given as numbers of imaging positions, mice, or independent experiments/preparations in figure legends rather than via a priori power calculations. Proteomic/transcriptomic enrichment (e.g., GO term analysis) is reported using fold-change and p-value cutoffs (FC ≥ 10, P ≤ 0.01) without an explicitly stated correction method in the excerpted text.

Replicationmixed Sample sizeSample sizes given per figure as numbers of imaging positions, mice, independent experiments, or independent EV preparations (e.g., n = 26 positions in eight mice; n = 3 independent experiments); no power/sample-size calculation described GroupsB16F1 vs B16F10 migration speed; EV-cargo-receiving (eGFP+) vs non-receiving (DsRed+) recipient cells; in vitro vs in vivo Cre+ EV transfer efficiency; EV proteome/transcriptome vs donor cell Pairingpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Wilcoxon signed rank test Fig 1D — relative migration speed of B16F10 vs B16F1 cells within the same imaging field n = 26 positions in eight mice not stated
Wilcoxon signed rank test Fig 2C–E — migration speed of recipient cells (eGFP+ vs DsRed+ reporter+) within the same imaging field n = 15 positions in four mice (C); n = 16 positions in four mice (D); n = 22 positions in six mice (E) not stated
Student's t-test Fig 1J — in vitro co-culture Cre+ EV transfer vs reporter-only control n = 3 independent experiments not stated
Mann-Whitney test Fig 1J — in vitro vs in vivo Cre+ EV transfer efficiency n = 3 independent experiments not stated
Fold-change/p-value threshold (statistical test for enrichment not named in this excerpt) Fig 3E — GO term enrichment for cellular compartment of proteins enriched in EVs vs donor cell three independent EV preparations not stated
Approaches that could also have been used
  • Paired migration-speed comparisons within the same imaging field are analyzed with the Wilcoxon signed-rank test.
    Could also: A paired (two-tailed) t-test, or a linear mixed-effects model with imaging field/mouse as a random effect — A mixed-effects model would also let multiple cells per field and per mouse be modeled explicitly, which can increase precision when repeated measurements are nested within the same animal or field.
  • Several different comparisons across figures (Wilcoxon, Mann-Whitney, t-test) do not describe a shared multiple-comparison correction.
    Could also: A false discovery rate procedure such as Benjamini-Hochberg, or a Bonferroni correction, applied across the family of related comparisons — This would also help control the overall (family-wise or FDR) error rate when several related hypotheses are tested within the same study.
  • GO term enrichment of EV proteins is reported using a fold-change ≥ 10 and P ≤ 0.01 cutoff, without a stated correction method in this excerpt.
    Could also: A hypergeometric/Fisher's exact test with Benjamini-Hochberg FDR correction, as is standard in tools like DAVID, GOseq, or clusterProfiler — This is a widely used approach for enrichment analyses and would also formally account for testing many GO categories simultaneously.
  • Variability for migration-speed data is summarized as mean ± SEM.
    Could also: Reporting SD or a 95% confidence interval alongside or instead of SEM — SD directly conveys the spread of individual observations, and a CI additionally communicates the precision of the estimated group difference, which can be informative with the modest sample sizes used here (e.g., n = 3 to n = 26).
  • Differences are summarized with p-value thresholds (e.g., P ≤ 0.01) rather than exact p-values or CIs for the group comparisons.
    Could also: Reporting exact p-values together with an effect size and its confidence interval — Exact values and CIs would also allow readers to gauge both the magnitude and precision of an effect, beyond a binary significance threshold.
  • RNA-seq and label-free proteomics data (transcript/protein counts across conditions) are described primarily by detection counts and fold-change/enrichment thresholds rather than a named differential-analysis model.
    Could also: Model-based differential expression tools such as DESeq2 or limma (for transcripts) and MSstats or similar for proteomics, typically paired with FDR correction — These approaches would also explicitly model count/intensity distributions and variance across replicates, which can be useful for quantifying and ranking cargo differences between EV populations.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-29907695

Paper: Steenbeek SC, Pham TV, de Ligt J, et al. "Cancer cells copy migratory behavior and exchange signaling networks via extracellular vesicles." EMBO J 2018;37(15):e98357. PMID 29907695 · PMCID PMC6068466 · DOI 10.15252/embj.201798357.

Code (Data/Code availability): RNA-seq processed with the UMC Utrecht in-house pipeline https://github.com/UMCUGenetics/RNASeq (Perl wrapper around standard tools). Data: ENA PRJEB20729 — 72 single-end NextSeq-500 runs = 18 biological samples (2 cell lines × 3 fractions × 3 replicates), 4 lanes/sample. sample_alias recovers the full design (e.g. B16F1_100kEV_2).

The pipeline (Materials & Methods, "RNA sequencing")

FastQC (v0.11.4) QC → STAR (v2.4.2a) align to GRCm38HTSeq-count (v0.6.1), union mode → DESeq2 median-of-ratios normalization → differential expression, selection Log2FC ≥ 1 and P ≤ 0.05 (some panels P ≤ 0.01). This is a fully standard, P16-valid pipeline: applying STAR→HTSeq→DESeq2 to the deposited reads per the described parameters reproduces the reported counts.

IN SCOPE (pipeline-derived, attempted here)

id reported result paper loc how reproduced
C1 DE transcripts B16F1 vs B16F10 cells = 65 Results / Fig EV / text DESeq2 contrast Cell F1 vs F10, count genes |log2FC|≥1 & P≤0.05
C2 DE in 16.5K EVs (F1 vs F10) = 105 Results text DESeq2 contrast 16.5kEV F1 vs F10, same threshold
C3 DE in 100K EVs (F1 vs F10) = 571 Results text DESeq2 contrast 100kEV F1 vs F10, same threshold
C4 transcripts present in all 3 replicates: B16F1 16.5K=12,450; 100K=12,802; B16F10 16.5K=11,696; 100K=11,527 Results text per EV group, genes with count ≥ 1 in all 3 reps
C5 EV-enriched vs donor cells (B16F1): 231 & 249 (16.5K), 1,089 & 1,463 (100K) Results text DESeq2 EV-vs-Cell contrasts, enriched side

Primary targets are C1–C4 (clearest, single thresholds). C5 is a bonus (same DESeq2 framework). Detection threshold and raw-vs-adjusted P are paper-ambiguous; both interpretations are reported (auditability over assertion).

OUT OF SCOPE (not attempted, with reason)

  • All wet-lab / imaging results (intravital microscopy, Cre/LoxP colour switching, migration assays, EV uptake, proteomics MS) — not pipeline-derived.
  • Mass-spectrometry proteomics networks — separate non-RNA pipeline, no shipped code.
  • Exact STAR 2.4.2a / HTSeq 0.6.1 builds + exact Ensembl GTF release — the paper does not pin the annotation version; we use a documented GRCm38 Ensembl release and closest obtainable tool versions, noting the deviation (the hard last ~20%).
  • GO/pathway enrichment category membership — qualitative, no pinnable number.

Compute

All on «our HPC» («infra»), partition std (192-core / 768 GB nodes). Data + index + BAMs stay on «infra»; only small result tables return to «host».

Figures / tables: Fig EV5BFig EV2AFig EV3Fig 5B
C1
Reported
65
Reproduced
raw_p=220/padj=39/up_raw=104
partial
C2
Reported
105
Reproduced
raw_p=847/padj=71/up_raw=103
partial
C3
Reported
571
Reproduced
raw_p=1843/padj=1076/up_raw=899
did not match
C4a
Reported
12450
Reproduced
14115
partial
C4b
Reported
12802
Reproduced
15268
partial
C4c
Reported
11696
Reproduced
12541
partial
C4d
Reported
11527
Reproduced
12136
partial
C5a
Reported
231
Reproduced
raw_p=590/padj=181/up_raw=267
partial
C5b
Reported
1089
Reproduced
raw_p=2192/padj=1558/up_raw=1007
partial
C5c
Reported
249
Reproduced
not_computed (B16F10 EV-vs-cell not emitted)
partial
C5d
Reported
1463
Reproduced
not_computed (B16F10 EV-vs-cell not emitted)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 46/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

Reproduction ran end-to-end on the authors' own ENA data (PRJEB20729, 18/18 samples, grade A) with a P16-valid third-party pipeline, recovering the correct order of magnitude and qualitative pattern (100K EVs >> 16.5K EVs ~ cells; B16F1 > B16F10; ~12-15k transcripts), with several values within ~2-13% (C2, C4c/d, C5b). The deviations sit on the input/preprocessing and methodology side — an unpinned Ensembl annotation release plus tool-version drift and the paper's own raw-vs-adjusted-P ambiguity — not on a fabrication signal; no value looks non-derivable from the shipped data. The lone clear miss is C3 (571 vs nearest 899, ~1.6x), and C5c/C5d were not computed due to genuine paper-side ambiguity. Overall a solid partial reproduction with explainable, technical/underspecification-driven discrepancies.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

622 k
tokens (I/O) · 52.9 M incl. cache
426 min
runtime · 8.25 CPU-h
46.4 GB
peak RAM
1
HPC jobs
hummel
machine