In vivo microscopy reveals macrophage polarization locally promotes coherent microtubule dynamics in migrating cancer cells.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH + 1:1 WHERE INPUTS ARE CONSISTENT. The MT_Dynamics image-analysis pipeline (the authors' own code) is reproducible: re-running the shipped extract_features notebook on the repo's shipped example_data under a Python-3.6 env pinned to python_dependencies.txt regenerated the authors' shipped feature pickles (t2/t3/t4/t6_features) to floating-point identity for 2 of the 3 control EB3 movies (all 17->30 columns, incl. the coherence/cosine-similarity features that define the paper's claim; worst abs diff ~1e-9..1e-11). The 3rd movie (03 '.oib - Series 1') mismatches because its shipped CSV (max_frame=39) trips the code's run_length=5-if-max_frame>40-else-3 branch to run_length=3, whereas its shipped reference pickle (285 tracks) is consistent with run_length=5+a further data difference -- i.e. the shipped CSV is inconsistent with its shipped reference pickle for that one movie (a provenance slip in the example, not a divergence from our process; the two internally-consistent movies are exact). Reproduced feature VALUES are sane: speed 0.53 um/s (paper in-vitro ~0.35, same units/order) and near-zero cell-wide coherence (consistent with the paper's low in-vitro coherence). NOT ATTEMPTED: the paper's headline biological figure numbers (in-vivo vs in-vitro speed/coherence/orientation), because the underlying intravital+in-vitro EB3 movie set is not publicly released (only 3 control example movies ship); flagged in AUDIT.md/claims.tsv as unverifiable-from-public-artifacts (not by itself fabrication). Also: the brief's data accession GSE118828 is wrong (ovarian scRNA-seq, PMID 30383866) -- this imaging paper has no sequencing data. Verdicts are provisional/automated; AUDIT.md guides a human re-check.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-15 ⛓ 44933d58109a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether microtubule (MT) dynamics in cancer cells differ when observed in vivo (via intravital microscopy) compared to standard 2D and 3D in vitro culture models, and whether local signaling from tumor-associated macrophages (TAMs), particularly IL4-polarized macrophages, drives the coherent MT alignment observed in vivo to promote tumor cell elongation and migration.
- ★ Cancer cells in vivo display higher coherent orientation of MT dynamics along their major axis compared to 2D in vitro cultures finding
- ★ The in vivo MT coherence phenotype is distinct from and not fully reproduced by 3D collagen gel cultures finding
- ★ Co-culturing tumor cells with IL4-polarized macrophages reproduces the in vivo MT coherence/orientation phenotype in vitro finding
- ★ MΦ depletion, MT disruption, targeted kinase inhibition, and IL10R blockade (altering MΦ polarization) all reduce MT coherence and/or tumor cell elongation finding
- ★ Developed an integrated imaging pipeline combining intravital microscopy, automated plus-end (EB3) tip tracking, and multivariate statistics to quantify MT dynamics in live xenograft tumors method
- EB3-mApple expression is a relatively non-perturbative reporter of endogenous MT dynamics finding
- ★ TAMs frequently neighbor or wrap around MT-rich tumor cell protrusions near vasculature and collagen fibers finding
- Increased in vivo MT coherence is not entirely explained by differences in cell elongation between conditions finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Intravital microscopy with EB3 plus-end tip tracking (plusTipTracker) | HT1080-EB3-mApple subcutaneous xenograft, dorsal window chamber, nu/nu mice | none | MT track features (growth speed, orientation, coherence, curvature, etc.) | confocal intravital microscopy |
| Intravital microscopy with EB3 plus-end tip tracking | ES2-EB3-mApple ovarian cancer xenograft | none | MT track features | confocal intravital microscopy |
| Live-cell fluorescence microscopy with EB3 tracking | HT1080-EB3-mApple and ES2-EB3-mApple, 2D tissue culture plastic | none | MT track features (matched to in vivo) | same imaging/tracking system as in vivo |
| Confocal microscopy with EB3 tracking + PCA | HT1080-EB3-mApple in 3D collagen I hydrogel | 3D collagen embedding vs 2D/in vivo | MT track features, PCA scores/loadings, cell shape features | confocal microscopy |
| Fluorescence microscopy with EB3 tracking, co-culture | HT1080/ES2-EB3-mApple + bone-marrow-derived or RAW264.7 macrophages | IL4 polarization of macrophages (M2-like, IL4-MΦ) | MT coherence and orientation, effect size (Cohen's D) | — |
| Fluorescence microscopy with EB3 tracking, co-culture + PCA | HT1080/ES2 tumor cells + macrophages | MΦ polarization state: M0 (MCSF only) vs LPS/IFNγ (M1-like) vs IL4 (M2-like) | MT coherence and orientation (PCA composite score) | — |
| Multicolor intravital microscopy (fluorescent reporter + NP + SHG) | HT1080 xenograft in Mertk-GFP/+ NOD.SCID reporter mice | none | Tumor cell proximity to GFP+ TAMs and AngioSPARK+ vasculature; EB3 tracks | 2-photon microscopy / second harmonic generation (SHG) |
| Genomic/RNA-seq correlation analysis | The Cancer Genome Atlas (TCGA) patient cohort | none | MAPRE3 (EB3) alteration/expression correlation with overall survival | — |
- ▲ In vivo HT1080 MT tracks show increased mean cellular coherence vs in vitro (0.10 ± 0.006 vs 0.03 ± 0.002 s.e.m.) 3.2-fold
- ▼ Fewer HT1080 MT tracks angled >45° off the cell major axis in vivo vs in vitro (16.7 ± 0.79% vs 32.6 ± 0.60% s.e.m.) ~2-fold difference
- – 3D collagen culture increased MT orientation (phenocopying in vivo) but 5/14 features, including cellular MT coherence, did not match in vivo behavior 5/14 features mismatched
- – Both in vivo and 3D-cultured cells show a positive PC1 shift (elongation) vs 2D culture, but this does not fully account for the distinct in vivo MT phenotype
- ▲ IL4-MΦ co-culture consistently increased MT coherence and orientation in both HT1080 and ES2 cells, and in both BMDM and RAW264.7 macrophage models
- – TAM to HT1080 tumor cell ratio in xenograft model ~1:4
- – EB3-mApple expression did not correlate with most MT track features (12/14); weak correlation with displacement/path length disappeared after correcting for track duration R2 < 0.25
- – Manual validation of plusTipTracker tracking accuracy showed a low false positive rate 11/300 (<5%)
- fold_change 3.2-fold (increase in mean cellular MT coherence, in vivo vs in vitro HT1080)
- mean 0.10 ± 0.006 (in vivo) vs 0.03 ± 0.002 (in vitro) s.e.m. (mean cellular MT coherence, HT1080)
- mean 32.6 ± 0.60% (in vitro) vs 16.7 ± 0.79% (in vivo) s.e.m. (fraction of MT tracks angled >45° off cell major axis)
- mean 0.35 ± 0.15 µm/s (in vitro) vs 0.38 ± 0.18 µm/s (in vivo) s.d. (average MT growth rate, HT1080)
- correlation R2 < 0.25 (EB3-mApple expression vs MT track displacement/path length)
- count 11/300 false positives (<5%) (manual validation of plusTipTracker tracking accuracy)
- count n = 8126 tracks, 73 cells, 4 tumors (HT1080); n = 2857 tracks, 42 cells, 5 tumors (ES2) (sample sizes for in vivo/in vitro MT track analysis)
- other ~1:4 ratio (relative abundance of TAMs to HT1080 tumor cells in xenograft, based on 80 cells / 4 tumors)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combined intravital microscopy with automated MT plus-end tip tracking (plusTipTracker) to extract 14 features of MT dynamics from individual tracks, then compared those features across in vivo xenograft, 2D in vitro, and 3D collagen culture conditions, as well as across macrophage co-culture and polarization conditions. Primary group comparisons used two-tailed permutation tests, with Benjamini-Hochberg (BH) FDR correction applied in at least two figure panels; a two-tailed t-test was used for cell morphology comparisons. Dimensionality reduction via PCA was used to summarize multivariate MT feature and cell shape profiles across conditions, and Cohen's D effect sizes were reported alongside significance stars to convey the magnitude of differences. Results were summarized with medians (for track distributions) and means ± SEM or SD, with individual cell- or tumor-level data points overlaid.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Two-tailed permutation test with Benjamini-Hochberg FDR correction | In vivo vs. 2D in vitro comparison of 14 MT track features (Fig. 1d, e) | 8126 tracks across 73 cells and 4 tumors (HT1080); 2857 tracks across 42 cells and 5 tumors (ES2) | not stated |
| Two-tailed permutation test | MT orientation and cellular coherence across 2D, 3D collagen, and in vivo conditions (Fig. 2c) | 9448 total tracks from 85 cells | not stated |
| Two-tailed t-test | Cell shape features (elongation, circularity) across 2D, 3D, and in vivo conditions (Fig. 2e) | 106 cells | not stated |
| Two-tailed permutation test with Benjamini-Hochberg FDR correction | Monoculture vs. IL4-MΦ co-culture MT feature comparisons across both cell lines (Fig. 4d) | 22,371 tracks across 164 cells (HT1080); 1424 tracks across 33 cells (ES2) | not stated |
| Principal components analysis (PCA) | MT feature dimensionality reduction across culture conditions (Fig. 2b) and cell shape feature profiles (Fig. 2d); also used to derive a combined MT coherence + orientation score for MΦ polarization comparisons | 5794 tracks middle-95% (MT PCA); 106 cells (shape PCA) | na |
| Cohen's D effect size | Effect size comparison of individual MT features between in vivo vs. in vitro (Fig. 1d) and monoculture vs. co-culture (Fig. 4c) | — | na |
-
Permutation tests were applied treating individual MT tracks as observations, with tracks nested within cells nested within tumors↳ Could also: Linear mixed-effects models (e.g., lme4 in R) with random intercepts for tumor and cell could also be used to formally account for the hierarchical nesting of tracks within cells within tumors — Mixed-effects models explicitly model within-tumor and within-cell correlation, reducing the risk of pseudoreplication when hundreds or thousands of tracks from relatively few biological replicates (4–5 tumors) are pooled; this approach can also estimate how much variance is attributable to each level of the hierarchy
-
Cell shape features across three conditions (2D, 3D, in vivo) were compared using pairwise two-tailed t-tests (Fig. 2e)↳ Could also: A one-way ANOVA with a post-hoc correction (e.g., Tukey HSD or Dunnett's test vs. a reference condition) could also be used when comparing three groups simultaneously — An omnibus ANOVA followed by post-hoc correction controls the family-wise error rate across all pairwise comparisons within the three-group family, whereas separate uncorrected t-tests inflate the Type I error rate as the number of comparisons grows
-
Variability around group means is reported as SEM in several places (e.g., 0.10 ± 0.006 s.e.m.; data are means ± s.e.m.)↳ Could also: Standard deviation (SD) or 95% confidence intervals could also be reported to summarize spread — SEM decreases with larger sample sizes and reflects estimation precision of the mean rather than biological variability across observations; SD or 95% CI more directly convey the spread of the underlying distribution and facilitate reader assessment of effect magnitude relative to variability, particularly when n is small at the biological-replicate level
-
14 MT track features were each tested individually between conditions, with BH correction applied in some but not all panels↳ Could also: A multivariate test (e.g., MANOVA or a permutation-based multivariate ANOVA such as PERMANOVA) could also serve as a single omnibus test across all 14 features before examining individual features — An omnibus multivariate test provides a single, family-wide error-controlled answer to whether the full MT feature profile differs between conditions, which complements or can precede the per-feature tests and reduces the burden of multiple-comparison correction across 14 correlated outcomes
-
PCA was used to reduce MT track features and cell shape features to low-dimensional summaries compared across conditions↳ Could also: Supervised dimensionality reduction such as linear discriminant analysis (LDA) or sparse PCA could also be used, and nonlinear methods such as UMAP could also visualize the feature space — LDA maximizes between-group separation and yields a quantitative discriminability score; UMAP can reveal nonlinear cluster structure that PCA may not capture; either would complement PCA by providing an alternative lens on whether the feature spaces of in vivo, 2D, and 3D conditions are truly distinct
-
Effect sizes were reported as Cohen's D computed on individual track feature distributions↳ Could also: A rank-based effect size such as the rank-biserial correlation, or Glass's delta using a reference-condition SD, could also accompany the nonparametric permutation tests — Cohen's D assumes approximately normal, equal-variance distributions; because MT dynamics data (e.g., coherence, orientation) can be skewed and the permutation test itself is distribution-free, a rank-based effect size is internally consistent with the chosen inferential approach and may more accurately reflect the magnitude of differences in heavy-tailed or bounded distributions
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
In vivo HT1080 microtubule track coherence is ~3.2-fold higher than in vitro 2D culture.imaging ht1080-xenograft-mouse up 2020×1papers★ This paper is the founder (earliest)
-
Microtubule track alignment along the cell major axis is higher in vivo than in vitro in HT1080 xenografts, with roughly half the fraction of off-axis tracks.imaging ht1080-xenograft-mouse up 2020×1papers★ This paper is the founder (earliest)
-
3D collagen culture partially recapitulates in vivo microtubule dynamics in HT1080 cells, increasing MT orientation but failing to reproduce in vivo cellular MT coherence (5 of 14 features mismatched).imaging ht1080 mixed 2020×1papers★ This paper is the founder (earliest)
-
IL4-polarized (M2-like) macrophage co-culture increases microtubule coherence and orientation in neighboring HT1080 and ES2 tumor cells.imaging ht1080 up 2020×1papers★ This paper is the founder (earliest)
-
In vivo and 3D-cultured HT1080 cells both show increased elongation and positive PC1 shift relative to 2D, but cell shape elongation alone does not account for in vivo microtubule dynamics.imaging ht1080 mixed 2020×1papers★ This paper is the founder (earliest)
-
EB3-mApple expression level does not significantly correlate with 12 of 14 microtubule track features in HT1080 cells, validating the reporter system.imaging ht1080 none 2020×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32665556
Paper: Luthria et al. 2020, Nat Commun 11:3521. "In vivo microscopy reveals macrophage polarization locally promotes coherent microtubule dynamics in migrating cancer cells." DOI 10.1038/s41467-020-17147-y · PMCID PMC7360550.
Code: https://github.com/gluthria/MT_Dynamics (commit 61acd79, master).
"Data" in brief (GSE118828): WRONG accession — that GEO series is single-cell
RNA-seq of ovarian cancer (PMID 30383866), unrelated to this imaging paper. This
paper has no sequencing data; its inputs are intravital/confocal microscopy
movies of EB3-comet (microtubule plus-end) dynamics. The mis-enrichment is noted
and ignored.
What the computational pipeline does (in scope)
extract_features_from_track_matrices.ipynb + track_analysis_functions_v1.py
take EB3-comet track matrices (CSV, exported from U-track/plusTipTracker via a
MATLAB step) plus per-cell segmentation masks, and compute a MT-track × features
DataFrame per movie. Feature families (match the paper's Methods):
| Pipeline feature (code) | Paper concept |
|---|---|
calc_motion_features → speed, net_displacement, path length, persistance, curvature |
speed (µm/s), displacement, persistence, curvature |
calc_mass_features → mass mean/std |
comet intensity |
compute_cosine_distance → similarity_20/200/5000 |
coherence = cosine similarity of a track to nearby tracks within a radius (the paper's defining metric) |
getOrientationFromAxis → maj_orient/min_orient |
orientation vs cell major/minor axis |
identifyCellFromTrack, distance-to-axis/centre |
spatial localisation |
Pipeline = Python 3.6 (numpy/pandas/scipy/scikit-image/trackpy/pims). The repo
ships example_data/: 3 "Control EB3" movies' track matrices + masks +
resolution/framerate table, and the authors' own output pickles
(pickle_objs/t2_features.p, t3_features.p, t4_features.p, t6_features.p, plus
*_tracks.p, track_distance_dicts.p, file_names.p).
In-scope reproduction target (what we attempt)
RU-1 (pipeline reproducibility, 1:1): Re-run the shipped notebook on the
shipped example_data under a Python-3.6 env pinned to python_dependencies.txt,
and check it regenerates the authors' shipped feature pickles
(t2/t3/t4/t6_features.p) numerically. This is the cleanest honest reproduction:
their code, their data, their reference outputs. Graded by per-column numeric
agreement after aligning images by file_names.p.
RU-1b (feature-value sanity): report the reproduced speed and coherence
(similarity_20) distributions actually computed on the example control data, in
the paper's units, to show the metric machinery yields sensible values.
Out of scope (not attempted — and why)
- Headline biological numbers — in-vivo HT1080 speed 0.38±0.18 µm/s vs in-vitro
0.35±0.15; cellular coherence 0.10±0.006 vs 0.03±0.002 (3.2-fold); orientation
45° deviation 16.7% vs 32.6% (paper Figs 1–3). These require the full set of intravital + in-vitro EB3 movies, which are NOT publicly released (only 3 control example movies ship). Not reproducible from available data → not attempted. Flagged: these values are not derivable from the shipped data/code alone.
- Wet-lab / imaging: intravital two-photon microscopy, macrophage polarization staining/flow cytometry, drug treatments, co-culture. Out of scope by definition.
- U-track tracking step + MATLAB
create_track_matrices.mlx: upstream of the shipped CSVs; the example CSVs are the documented pipeline input, so we start there. - Downstream stats (PCA, permutation tests, K-L divergence): depend on the full dataset's feature tables; not attempted.
80/20 statement
The low-hanging, clearly-specified output is RU-1: does the shipped pipeli
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Re-running the authors' MT_Dynamics pipeline on their own shipped example_data regenerated the feature/coherence tables to floating-point identity for 2 of 3 control movies (all 17->30 cols, worst ~1e-9), so the computational core — including the coherence similarity_* metric central to the paper — is genuinely reproducible. The single mismatched movie is an authors'-side provenance slip (shipped CSV inconsistent with its shipped reference pickle via a frame-threshold branch), not a divergence from our process. The paper's headline biological numbers cannot be reproduced because the underlying intravital/in-vitro EB3 movie set is unreleased — a data-availability limit, not demonstrated fabrication — and the brief's GEO accession (GSE118828) is wrongly attached to this imaging paper. Net: a solid pipeline-level 1:1 with explainable deviations and an unavoidable ceiling on the central claim.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.