Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

In vivo microscopy reveals macrophage polarization locally promotes coherent microtubule dynamics in migrating cancer cells.

Nat Commun · 2020
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH + 1:1 WHERE INPUTS ARE CONSISTENT. The MT_Dynamics image-analysis pipeline (the authors' own code) is reproducible: re-running the shipped extract_features notebook on the repo's shipped example_data under a Python-3.6 env pinned to python_dependencies.txt regenerated the authors' shipped feature pickles (t2/t3/t4/t6_features) to floating-point identity for 2 of the 3 control EB3 movies (all 17->30 columns, incl. the coherence/cosine-similarity features that define the paper's claim; worst abs diff ~1e-9..1e-11). The 3rd movie (03 '.oib - Series 1') mismatches because its shipped CSV (max_frame=39) trips the code's run_length=5-if-max_frame>40-else-3 branch to run_length=3, whereas its shipped reference pickle (285 tracks) is consistent with run_length=5+a further data difference -- i.e. the shipped CSV is inconsistent with its shipped reference pickle for that one movie (a provenance slip in the example, not a divergence from our process; the two internally-consistent movies are exact). Reproduced feature VALUES are sane: speed 0.53 um/s (paper in-vitro ~0.35, same units/order) and near-zero cell-wide coherence (consistent with the paper's low in-vitro coherence). NOT ATTEMPTED: the paper's headline biological figure numbers (in-vivo vs in-vitro speed/coherence/orientation), because the underlying intravital+in-vitro EB3 movie set is not publicly released (only 3 control example movies ship); flagged in AUDIT.md/claims.tsv as unverifiable-from-public-artifacts (not by itself fabrication). Also: the brief's data accession GSE118828 is wrong (ovarian scRNA-seq, PMID 30383866) -- this imaging paper has no sequencing data. Verdicts are provisional/automated; AUDIT.md guides a human re-check.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-15 ⛓ 44933d58109a
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

How do microtubule (MT) dynamics actually behave in cancer cells in vivo within the tumor microenvironment, and do local intercellular interactions—specifically pro-tumor macrophage signaling—promote coherent MT dynamics that drive tumor cell elongation and migration?

Core claims
  • Cancer cells in vivo display higher coherent orientation and alignment of MT dynamics along their cell major axis than 2D in vitro cultures, and distinct from 3D collagen gel cultures. finding
  • The in vivo MT coherence phenotype is reproduced in vitro when tumor cells are co-cultured with IL4-polarized (M2-like) macrophages. finding
  • Pro-tumor (IL4-polarized) macrophage signaling locally promotes coherent MT dynamics and elongation in neighboring tumor cells, modulating in vivo tumor cell motility/migration. mechanism
  • MΦ depletion, MT disruption, targeted kinase inhibition (EGFR/PI3K), and IL10R blockade (altered MΦ polarization) reduce MT coherence and/or tumor cell elongation. finding
  • An integrated imaging pipeline combining intravital microscopy, EB3-mApple plus-end tip tracking (plusTipTracker), and multivariate statistics quantifies MT dynamics in live xenograft tumors. method
  • MT coherence is a defining feature of in vivo tumor cell dynamics and migration. finding
  • EB3-mApple is a relatively non-perturbative reporter of MT dynamics; MAPRE3 expression/alterations do not correlate with cancer patient survival in TCGA. resource
  • TAMs frequently neighbor or wrap around MT-rich tumor cell protrusions, near vasculature and fibrillar extracellular matrix. finding
Experimental setups
Assay System Perturbation Readout Platform
Intravital (in vivo confocal) microscopy with EB3-mApple plus-end MT tip tracking HT1080 human fibrosarcoma xenografts in nu/nu mice, dorsal window chamber none (baseline imaging) MT growth track features (speed, orientation, coherence, location, persistence, curvature, displacement, path length) plusTipTracker algorithm
Intravital microscopy with EB3-mApple tip tracking ES2 human ovarian cancer xenografts in mice none MT track features, in vivo vs in vitro plusTipTracker
In vitro time-lapse fluorescence microscopy (2D culture) HT1080-EB3-mApple cells on 2D tissue culture plastic none MT track features for in vivo vs in vitro comparison same IVM imaging system; plusTipTracker
Confocal time-lapse microscopy (3D culture) HT1080-EB3-mApple cells in 3D collagen I hydrogel none MT track features and PCA of MT/shape features plusTipTracker; PCA
Multicolor intravital microscopy + 2-photon second harmonic generation HT1080 xenografts in Mertk-GFP/+ NOD.SCID reporter mice none TAM proximity (GFP+) to MT-rich tumor cells, vasculature (AngioSPARK-680 NP), and ECM fibers 2-photon microscopy; AngioSPARK-680 nanoparticle
In vitro macrophage co-culture with EB3 MT tracking HT1080 and ES2 tumor cells co-cultured with bone-marrow-derived IL4-polarized (M2-like) MΦ or RAW264.7-derived IL4-MΦ IL4-MΦ co-culture (24 h) MT coherence and orientation in MΦ-adjacent tumor cells plusTipTracker
In vitro macrophage polarization co-culture with EB3 MT tracking Tumor cells co-cultured with M0-like MΦ (MCSF), IL4-MΦ (M2-like), or LPS/IFNγ-MΦ (M1-like) differential MΦ polarization MT coherence/orientation via PCA single principal component plusTipTracker; PCA
Key results
  • In vivo HT1080 MT tracks showed ~3.2-fold higher mean cellular coherence than in vitro tracks 3.2-fold (0.10±0.006 vs 0.03±0.002 s.e.m.)
  • Nearly twice as many MT tracks were angled >45° off the cell major axis (orientation <0.71) in vitro vs in vivo 32.6±0.60% vs 16.7±0.79% s.e.m.
  • IL4-MΦ co-culture increased MT coherence and orientation in neighboring tumor cells across HT1080 and ES2, and with two MΦ models
  • 3D collagen culture increased MT orientation but failed to reproduce in vivo cellular MT coherence; 5/14 features did not match in vivo 5/14 features mismatched
  • MT tip-tracking false positive rate manually assessed as <5% <5% (11/300)
  • EB3-mApple expression did not significantly correlate with most MT track features (12/14); displacement and path length correlation lost after correcting for track duration R2<0.25
  • Both in vivo and 3D-cultured cells showed positive PC1 shift (decreased circularity, increased elongation) vs 2D, but in vivo MT dynamics not fully explained by elongation
  • TAM:HT1080 tumor cell content ratio roughly 1:4 in the xenograft model ~1:4
Key statistics
  • fold_change 3.2-fold increase in mean cellular coherence in vivo vs in vitro (HT1080 cellular MT coherence, 73 cells, 4 tumors)
  • mean 0.10±0.006 (in vivo) vs 0.03±0.002 (in vitro) s.e.m. (HT1080 mean cellular MT coherence)
  • other 32.6±0.60% (in vitro) vs 16.7±0.79% (in vivo) s.e.m. (% MT tracks angled >45° off cell major axis (orientation <0.71))
  • mean in vitro 0.35±0.15 vs in vivo 0.38±0.18 µm/s s.d. (HT1080 average MT growth rate)
  • count <5% false positive (11/300) (manual MT tracking accuracy check)
  • correlation R2<0.25 (EB3-mApple expression vs displacement/path length track features)
  • count n=2857 tracks across 42 cells and 5 tumors (ES2 xenograft MT track analysis)
  • count HT1080 n=22,371 tracks/164 cells; ES2 n=1424 tracks/33 cells (IL4-MΦ co-culture vs monoculture MT analysis)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combined intravital microscopy with automated MT plus-end tip tracking (plusTipTracker) to extract 14 features of MT dynamics from individual tracks, then compared those features across in vivo xenograft, 2D in vitro, and 3D collagen culture conditions, as well as across macrophage co-culture and polarization conditions. Primary group comparisons used two-tailed permutation tests, with Benjamini-Hochberg (BH) FDR correction applied in at least two figure panels; a two-tailed t-test was used for cell morphology comparisons. Dimensionality reduction via PCA was used to summarize multivariate MT feature and cell shape profiles across conditions, and Cohen's D effect sizes were reported alongside significance stars to convey the magnitude of differences. Results were summarized with medians (for track distributions) and means ± SEM or SD, with individual cell- or tumor-level data points overlaid.

Replicationmixed Sample sizeDescribed as number of MT tracks, cells, and tumors per condition; no formal a priori power analysis stated; xenograft experiments used n=4–5 tumors as the biological unit, with many cells and tracks per tumor GroupsIn vivo xenograft vs. 2D in vitro vs. 3D collagen gel; monoculture vs. IL4-MΦ co-culture; M0 vs. IL4-MΦ vs. LPS/IFNγ-MΦ polarization states; two cancer cell lines (HT1080, ES2) Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR correction
Statistical tests used
Test Applied to n Assumptions
Two-tailed permutation test with Benjamini-Hochberg FDR correction In vivo vs. 2D in vitro comparison of 14 MT track features (Fig. 1d, e) 8126 tracks across 73 cells and 4 tumors (HT1080); 2857 tracks across 42 cells and 5 tumors (ES2) not stated
Two-tailed permutation test MT orientation and cellular coherence across 2D, 3D collagen, and in vivo conditions (Fig. 2c) 9448 total tracks from 85 cells not stated
Two-tailed t-test Cell shape features (elongation, circularity) across 2D, 3D, and in vivo conditions (Fig. 2e) 106 cells not stated
Two-tailed permutation test with Benjamini-Hochberg FDR correction Monoculture vs. IL4-MΦ co-culture MT feature comparisons across both cell lines (Fig. 4d) 22,371 tracks across 164 cells (HT1080); 1424 tracks across 33 cells (ES2) not stated
Principal components analysis (PCA) MT feature dimensionality reduction across culture conditions (Fig. 2b) and cell shape feature profiles (Fig. 2d); also used to derive a combined MT coherence + orientation score for MΦ polarization comparisons 5794 tracks middle-95% (MT PCA); 106 cells (shape PCA) na
Cohen's D effect size Effect size comparison of individual MT features between in vivo vs. in vitro (Fig. 1d) and monoculture vs. co-culture (Fig. 4c) na
Approaches that could also have been used
  • Permutation tests were applied treating individual MT tracks as observations, with tracks nested within cells nested within tumors
    Could also: Linear mixed-effects models (e.g., lme4 in R) with random intercepts for tumor and cell could also be used to formally account for the hierarchical nesting of tracks within cells within tumors — Mixed-effects models explicitly model within-tumor and within-cell correlation, reducing the risk of pseudoreplication when hundreds or thousands of tracks from relatively few biological replicates (4–5 tumors) are pooled; this approach can also estimate how much variance is attributable to each level of the hierarchy
  • Cell shape features across three conditions (2D, 3D, in vivo) were compared using pairwise two-tailed t-tests (Fig. 2e)
    Could also: A one-way ANOVA with a post-hoc correction (e.g., Tukey HSD or Dunnett's test vs. a reference condition) could also be used when comparing three groups simultaneously — An omnibus ANOVA followed by post-hoc correction controls the family-wise error rate across all pairwise comparisons within the three-group family, whereas separate uncorrected t-tests inflate the Type I error rate as the number of comparisons grows
  • Variability around group means is reported as SEM in several places (e.g., 0.10 ± 0.006 s.e.m.; data are means ± s.e.m.)
    Could also: Standard deviation (SD) or 95% confidence intervals could also be reported to summarize spread — SEM decreases with larger sample sizes and reflects estimation precision of the mean rather than biological variability across observations; SD or 95% CI more directly convey the spread of the underlying distribution and facilitate reader assessment of effect magnitude relative to variability, particularly when n is small at the biological-replicate level
  • 14 MT track features were each tested individually between conditions, with BH correction applied in some but not all panels
    Could also: A multivariate test (e.g., MANOVA or a permutation-based multivariate ANOVA such as PERMANOVA) could also serve as a single omnibus test across all 14 features before examining individual features — An omnibus multivariate test provides a single, family-wide error-controlled answer to whether the full MT feature profile differs between conditions, which complements or can precede the per-feature tests and reduces the burden of multiple-comparison correction across 14 correlated outcomes
  • PCA was used to reduce MT track features and cell shape features to low-dimensional summaries compared across conditions
    Could also: Supervised dimensionality reduction such as linear discriminant analysis (LDA) or sparse PCA could also be used, and nonlinear methods such as UMAP could also visualize the feature space — LDA maximizes between-group separation and yields a quantitative discriminability score; UMAP can reveal nonlinear cluster structure that PCA may not capture; either would complement PCA by providing an alternative lens on whether the feature spaces of in vivo, 2D, and 3D conditions are truly distinct
  • Effect sizes were reported as Cohen's D computed on individual track feature distributions
    Could also: A rank-based effect size such as the rank-biserial correlation, or Glass's delta using a reference-condition SD, could also accompany the nonparametric permutation tests — Cohen's D assumes approximately normal, equal-variance distributions; because MT dynamics data (e.g., coherence, orientation) can be skewed and the permutation test itself is distribution-free, a rank-based effect size is internally consistent with the chosen inferential approach and may more accurately reflect the magnitude of differences in heavy-tailed or bounded distributions
Software: plusTipTracker

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
21
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE103322 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
GSE118828 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
5IQR in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
plasmid_54892 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:Addgene_59150 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-32665556

Paper: Luthria et al. 2020, Nat Commun 11:3521. "In vivo microscopy reveals macrophage polarization locally promotes coherent microtubule dynamics in migrating cancer cells." DOI 10.1038/s41467-020-17147-y · PMCID PMC7360550.

Code: https://github.com/gluthria/MT_Dynamics (commit 61acd79, master). "Data" in brief (GSE118828): WRONG accession — that GEO series is single-cell RNA-seq of ovarian cancer (PMID 30383866), unrelated to this imaging paper. This paper has no sequencing data; its inputs are intravital/confocal microscopy movies of EB3-comet (microtubule plus-end) dynamics. The mis-enrichment is noted and ignored.

What the computational pipeline does (in scope)

extract_features_from_track_matrices.ipynb + track_analysis_functions_v1.py take EB3-comet track matrices (CSV, exported from U-track/plusTipTracker via a MATLAB step) plus per-cell segmentation masks, and compute a MT-track × features DataFrame per movie. Feature families (match the paper's Methods):

Pipeline feature (code) Paper concept
calc_motion_features → speed, net_displacement, path length, persistance, curvature speed (µm/s), displacement, persistence, curvature
calc_mass_features → mass mean/std comet intensity
compute_cosine_distancesimilarity_20/200/5000 coherence = cosine similarity of a track to nearby tracks within a radius (the paper's defining metric)
getOrientationFromAxismaj_orient/min_orient orientation vs cell major/minor axis
identifyCellFromTrack, distance-to-axis/centre spatial localisation

Pipeline = Python 3.6 (numpy/pandas/scipy/scikit-image/trackpy/pims). The repo ships example_data/: 3 "Control EB3" movies' track matrices + masks + resolution/framerate table, and the authors' own output pickles (pickle_objs/t2_features.p, t3_features.p, t4_features.p, t6_features.p, plus *_tracks.p, track_distance_dicts.p, file_names.p).

In-scope reproduction target (what we attempt)

RU-1 (pipeline reproducibility, 1:1): Re-run the shipped notebook on the shipped example_data under a Python-3.6 env pinned to python_dependencies.txt, and check it regenerates the authors' shipped feature pickles (t2/t3/t4/t6_features.p) numerically. This is the cleanest honest reproduction: their code, their data, their reference outputs. Graded by per-column numeric agreement after aligning images by file_names.p.

RU-1b (feature-value sanity): report the reproduced speed and coherence (similarity_20) distributions actually computed on the example control data, in the paper's units, to show the metric machinery yields sensible values.

Out of scope (not attempted — and why)

  • Headline biological numbers — in-vivo HT1080 speed 0.38±0.18 µm/s vs in-vitro 0.35±0.15; cellular coherence 0.10±0.006 vs 0.03±0.002 (3.2-fold); orientation

    45° deviation 16.7% vs 32.6% (paper Figs 1–3). These require the full set of intravital + in-vitro EB3 movies, which are NOT publicly released (only 3 control example movies ship). Not reproducible from available data → not attempted. Flagged: these values are not derivable from the shipped data/code alone.

  • Wet-lab / imaging: intravital two-photon microscopy, macrophage polarization staining/flow cytometry, drug treatments, co-culture. Out of scope by definition.
  • U-track tracking step + MATLAB create_track_matrices.mlx: upstream of the shipped CSVs; the example CSVs are the documented pipeline input, so we start there.
  • Downstream stats (PCA, permutation tests, K-L divergence): depend on the full dataset's feature tables; not attempted.

80/20 statement

The low-hanging, clearly-specified output is RU-1: does the shipped pipeli

PIPE-1
Reported
shipped t2_features.p (motion+mass feature table, 17 numeric cols) = authors' reference output of the MT_Dynamics pipeline
Reproduced
2/3 example movies bit-identical (worst abs diff ~1e-9..1e-10); 1 movie (03 .oib Series 1) mismatch due to shipped-CSV/shipped-pickle inconsistency
partial
PIPE-3
Reported
shipped t4_features.p coherence/similarity table (similarity_20/200/5000 = mean cosine similarity to nearby tracks) = the paper's 'coherent MT dynamics' metric
Reproduced
2/3 movies exact (all 23 cols, worst ~2e-10); 1 movie mismatch (same provenance cause)
partial
PIPE-4
Reported
shipped t6_features.p final feature table (30 cols incl. orientation + cell IDs)
Reproduced
2/3 movies exact (all 30 cols); 1 movie mismatch
partial
VAL-speed
Reported
in-vitro 2D control MT speed 0.35 +/- 0.15 um/s (Results/Fig 1)
Reproduced
new_speed = 0.53 +/- 0.41 um/s on example control data (n=1172)
partial
VAL-coherence
Reported
in-vitro cellular coherence 0.03 +/- 0.002 (low; vs in-vivo 0.10, 3.2-fold)
Reproduced
cell-wide coherence similarity_5000 = 0.004 +/- 0.063 (near zero, qualitatively consistent with low in-vitro value)
partial
OUT-1
Reported
headline in-vivo vs in-vitro numbers (speed, 3.2-fold coherence, 16.7% vs 32.6% orientation>45deg), Figs 1-3
Reproduced
NOT ATTEMPTED - raw movie dataset not publicly released
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

Re-running the authors' MT_Dynamics pipeline on their own shipped example_data regenerated the feature/coherence tables to floating-point identity for 2 of 3 control movies (all 17->30 cols, worst ~1e-9), so the computational core — including the coherence similarity_* metric central to the paper — is genuinely reproducible. The single mismatched movie is an authors'-side provenance slip (shipped CSV inconsistent with its shipped reference pickle via a frame-threshold branch), not a divergence from our process. The paper's headline biological numbers cannot be reproduced because the underlying intravital/in-vitro EB3 movie set is unreleased — a data-availability limit, not demonstrated fabrication — and the brief's GEO accession (GSE118828) is wrongly attached to this imaging paper. Net: a solid pipeline-level 1:1 with explainable deviations and an unavoidable ceiling on the central claim.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

163.7 k
tokens (I/O) · 12.9 M incl. cache
36 min
runtime · 0.31 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine