Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Transcriptomic-Based Quantification of the Epithelial-Hybrid-Mesenchymal Spectrum across Biological Contexts.

Biomolecules · 2021
L1 100/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> faithful 1:1 reproduction of the in-scope (76GS + KS) results on the brief's dataset GSE118612 (Py2T epithelial vs MTdEcad mesenchymal mouse breast-cancer cells, one of the nine Fig-2 contexts). Ran the AUTHORS' OWN scoring functions (EMT76GS, KSScore from sushimndl/EMT_Scoring_RNASeq @50aaef0) VERBATIM on the public GEO deposit; only adaptation was bridging the deposit's Entrez gene IDs to the repo's Ensembl-keyed annotation via NCBI gene_info. Result: Py2T has higher 76GS (+4.61 vs -6.92) and lower KS (0.264 vs 0.365) than MTdEcad -- exactly the paper's Fig-2 directional claim -- and 76GS anti-correlates with KS (r=-0.988, p=9e-8), matching the paper's overall R<-0.3/p<0.05 concordance trend. This is a CLEAN RE-RUN («our HPC» «job», 2026-06-23) on a freshly rebuilt «infra» workdir+conda env; the output is BIT-FOR-BIT identical (emt_result.json SHA256 f032807f...) to the earlier run 2177258, confirming determinism under the fixed seed. NOT ATTEMPTED (the hard ~20%): (a) the MLR metric and its two correlations -- requires MATLAB (unlicensed) + a proprietary trained RelevantData.mat model; (b) the full 77-dataset survey and Table S1 exact per-dataset r-values -- supplement behind MDPI 403, and the main text reports only the aggregate 44/77=57.14% concordance, no single GSE118612 r to pin; (c) Fig-3 single-cell datasets (different data, out of this RU's scope). Grades are directional/sign confirmations (paper prints no single number for this dataset), provisional pending human audit. No fabrication flags.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-15 ⛓ a7872b9482b4
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether three distinct transcriptomic EMT scoring metrics (76GS, KS, MLR), each built on different gene lists and algorithms, can consistently quantify the epithelial-hybrid-mesenchymal spectrum across bulk and single-cell RNA-seq data, and whether these cancer-derived metrics generalize to non-cancer biological contexts.

Core claims
  • The 76GS, KS, and MLR EMT scoring metrics show concordant trends in quantifying EMP across bulk RNA-seq datasets spanning multiple cancer types finding
  • The microarray-trained MLR method was adapted for RNA-seq data via linear mapping of predictor/normalizer gene expression between platforms method
  • EMT scores from all three metrics recapitulate expected phenotypic shifts upon experimental perturbations (TGF-β treatment, Runx1/GRHL2/ZEB1 knockdown, HNF-1β loss) across multiple cell lines and models finding
  • Single-cell RNA-seq analysis reveals heterogeneity along the EMP spectrum, including bimodal KS score distributions indicating at least two EMT subpopulations finding
  • EMT metrics trained on cancer expression data can quantify EMP in non-cancer contexts such as lung fibrosis/COPD and iPSC reprogramming finding
  • Association between EMP scores and patient survival in TCGA pan-cancer analysis is context-specific finding
  • 76GS score is a weighted sum of 76 gene expression levels weighted by correlation with CDH1; higher score indicates a more epithelial sample resource
  • KS score (range -1 to +1) and MLR score (range 0 to 2) use distinct epithelial/mesenchymal gene signatures, with higher scores indicating more mesenchymal phenotype resource
Experimental setups
Assay System Perturbation Readout Platform
Transcriptomic EMT scoring (76GS, KS, MLR) 77 bulk high-throughput transcriptomic datasets (cell lines and primary tumors across cancer types) none/varied (dataset survey) EMT score concordance/correlation among three metrics
RNA-seq EMT scoring Py2T murine mammary tumor cells vs MTΔEcad cells (GSE118612) TGF-β treatment (reversible EMT) vs E-cadherin ablation (irreversible EMT) 76GS, KS, MLR EMT scores
RNA-seq EMT scoring MCF10A human mammary epithelial cells (GSE85857) Runx1 depletion 76GS, KS, MLR EMT scores
RNA-seq EMT scoring Primary airway epithelial cells (GSE72419) TGF-β treatment 76GS, KS, MLR EMT scores
RNA-seq EMT scoring HeLa cells (GSE61220) TGF-β and EGF treatment 76GS, KS, MLR EMT scores
RNA-seq EMT scoring OVCA429 ovarian cancer cells (GSE118407) GRHL2 knockdown 76GS, KS, MLR EMT scores
Single-cell RNA-seq EMT scoring 17 scRNA-seq datasets including 5902 cells from 18 oral cavity tumor/HNSCC patients (GSE103322) none (patient tumor samples) EMT score heterogeneity and correlation between metrics
Survival analysis (Kaplan-Meier, log-rank test, Cox regression) TCGA pan-cancer patient cohort RNA-seq expression data none (stratified by EMT score high/low groups) overall survival, hazard ratio, 95% CI UCSC Xena browser
Key results
  • 76GS scores negatively correlated with MLR and KS scores, while MLR and KS scores positively correlated with each other across most of 77 bulk datasets r<-0.3 and r>0.3, p<0.05
  • 44 of 77 (57.14%) datasets showed all three pairwise trends (KS vs MLR, MLR vs 76GS, 76GS vs KS) significantly 57.14%
  • 72 of 77 (93.5%) datasets showed concordance in at least two of the three EMT metrics 93.5%
  • Py2T cells (reversible EMT) had higher 76GS but lower KS and MLR scores than MTΔEcad cells (irreversible EMT), indicating relatively more epithelial status
  • Runx1-depleted MCF10A cells had higher KS and MLR scores but lower 76GS scores than controls, consistent with EMT induction
  • GRHL2 knockdown in OVCA429 cells increased MLR and KS scores but decreased 76GS scores, reflecting a more mesenchymal status
  • Across single-cell datasets, correlation trends between metrics were largely concordant with bulk data: 76GS vs KS negative in 11/17 datasets, MLR vs 76GS negative in 10/17, MLR vs KS positive in 9/17 65%, 59%, 53%
  • KS scores across multiple single-cell datasets showed two distinct histogram peaks, suggesting at least two major EMT subpopulations
Key statistics
  • correlation r < -0.3, p < 0.05 (76GS vs MLR and KS scores across bulk RNA-seq datasets)
  • correlation r > 0.3, p < 0.05 (MLR vs KS scores across bulk RNA-seq datasets)
  • count 44/77 (57.14%) (datasets with all three pairwise trends significant)
  • count 72/77 (93.5%) (datasets concordant in at least two of three metrics)
  • count 52/77 (67.53%) (datasets showing expected 76GS vs KS trend)
  • count 56/77 (72.72%) (datasets showing expected MLR vs 76GS trend)
  • count 57/77 (74.02%) (datasets showing expected MLR vs KS trend)
  • count 11/17 (65%) (single-cell datasets with negative 76GS vs KS correlation)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper applies three transcriptomic EMT scoring metrics (76GS, KS, MLR) to over 80 bulk and single-cell RNA-seq datasets, assessing concordance among metrics via pairwise correlation analysis and evaluating their ability to recapitulate expected EMT trends across cancer and non-cancer contexts. Group differences in EMT scores were tested with two-tailed Welch's t-tests, while patient survival associations used Kaplan–Meier analysis with log-rank tests and Cox proportional hazards regression. The MLR metric, originally trained on microarray data, was adapted to RNA-seq via ordinary least-squares linear regression on 24 paired samples, and ssGSEA enrichment scores served as an orthogonal validation of EMT status.

Replicationmixed Sample sizeNumber of datasets described (77 bulk, 17 single-cell RNA-seq); per-dataset sample sizes cited via GEO accessions where available; no formal power analysis stated GroupsEMT-induced vs. control cells across cancer cell lines and mouse models; high vs. low EMT score patient groups in TCGA pan-cancer survival cohorts; cancer vs. non-cancer biological contexts (lung fibrosis, iPSC reprogramming) Pairingmixed Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Two-tailed Student's t-test with unequal variance (Welch's t-test) Pairwise comparisons of EMT scores between experimental groups shown in bar plots (e.g., TGF-β-treated vs. control, Runx1-depleted vs. control MCF10A, GRHL2-knockdown vs. control) not stated
Pearson or Spearman correlation (type not specified; r values reported) Concordance assessment between all pairs of 76GS, KS, and MLR scores across 77 bulk RNA-seq datasets and 17 single-cell RNA-seq datasets 77 bulk datasets; 17 single-cell datasets not stated
Log-rank test Kaplan–Meier survival analysis comparing EMT-score-high vs. EMT-score-low groups in TCGA pan-cancer cohorts not stated
Cox proportional hazards regression Estimation of hazard ratios and 95% confidence intervals for EMT score group contrasts in TCGA survival analysis not stated
Ordinary least-squares linear regression Cross-platform mapping of log2 RNA-seq (FPKM and TPM) values to microarray-equivalent space for adaptation of the MLR scoring method 24 paired microarray and RNA-seq samples (6 time points × 2 biological replicates each, from a previously published dataset) not stated
Single-sample GSEA (ssGSEA) Per-sample enrichment scoring of the HALLMARK_EMT gene set from MSigDB as an orthogonal validation of EMT status across datasets na
Approaches that could also have been used
  • Concordance between the three scoring metrics was assessed across 77 bulk and 17 single-cell datasets using a uniform p < 0.05 threshold with no stated correction for multiple comparisons
    Could also: A false discovery rate (FDR) correction such as Benjamini–Hochberg could also be applied across the full family of correlation tests conducted simultaneously — When many hypothesis tests are performed in parallel, FDR control is a standard approach that limits the expected proportion of false discoveries among significant results; it would complement the threshold-based approach used here and is routinely reported in multi-dataset transcriptomic studies
  • Multiple independent two-tailed t-tests were used to compare EMT scores between experimental groups across many individual datasets and conditions
    Could also: A one-way or two-way ANOVA followed by a post-hoc correction (e.g., Tukey HSD or Dunnett's test) could also be applied when three or more groups are compared within a single experiment — ANOVA with post-hoc testing addresses the family-wise error rate for multi-group comparisons within a single experiment, which is an alternative to running separate pairwise t-tests; this can be particularly relevant for time-course or multi-condition designs such as the TGF-β treatment series
  • Dispersion in bar plots was reported as standard deviation (SD)
    Could also: A 95% confidence interval (CI) around the group mean could also be used to convey spread and estimation precision — CIs directly communicate the precision of the estimated group mean and facilitate inference about population-level differences; reporting guidelines from journals such as Nature Methods and many style guides recommend CIs alongside or instead of SD, especially when within-group n is small
  • Samples were dichotomized into high and low EMT score groups using the mean or median cutpoint for Kaplan–Meier survival analysis
    Could also: Maximally selected rank statistics (e.g., R package 'maxstat') or pre-specified tertile/quartile cutpoints could also be applied for group definition — Mean/median splits are simple and reproducible; data-adaptive optimal cutpoints can increase sensitivity to detect survival associations, while pre-specified quantile cuts reduce dependence on any single threshold value — both are used in the pan-cancer survival literature as alternatives
  • The type of correlation coefficient (Pearson vs. Spearman) used to assess concordance between EMT metrics is not explicitly stated in the methods
    Could also: Spearman rank correlation could also be applied, particularly given that EMT scores span bounded intervals ([0,2] for MLR, [−1,+1] for KS) and may not follow a bivariate normal distribution — Spearman correlation is a distribution-free alternative that is robust to outliers and monotone nonlinear relationships; explicitly stating and justifying the choice between Pearson and Spearman is standard practice when the distributional form of scores is uncertain
  • Cross-platform normalization from RNA-seq to microarray expression space for the MLR method was performed using ordinary least-squares linear regression on 24 paired samples
    Could also: Empirical Bayes batch-correction methods such as ComBat, or quantile normalization, could also be used for cross-platform harmonization of transcriptomic data — ComBat and related approaches are widely used in multi-platform transcriptomic studies to account for platform-specific technical variation across all genes simultaneously, offering an alternative framework to gene-by-gene linear scaling of the predictor set
Software: R 4.0.3 · ggplot2 (ggplot function) · GEOquery (R/Bioconductor) · R/survival · R/ggfortify · R/ssgsea · STAR-aligner · FASTQC · Samtools · htseq-count · UCSC Xena browser

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
16
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35053177

Paper: Mandal et al. 2021, Biomolecules 12(1):29. "Transcriptomic-Based Quantification of the Epithelial-Hybrid-Mesenchymal Spectrum across Biological Contexts." (Jolly lab, IISc). Repo: https://github.com/sushimndl/EMT_Scoring_RNASeq @ 50aaef0 (HEAD, 2021-09-09). Data (this RU): GEO GSE118612 — Py2T long-term cells (epithelial) and MTΔECad mesenchymal breast-cancer cells, mouse (mm10). 10 samples (GSM3334294–GSM3334303); raw counts in GSE118612_counts.txt.gz (Entrez GeneID-keyed, 18 103 genes × 10 samples).

What the paper/repo does

The repo scores every sample on three established EMT metrics and correlates them:

  • 76GS (Byers 2013): weighted sum of 76 EMT genes, weight = correlation with CDH1; output is mean-centred. Higher = more epithelial.
  • KS (Tan 2014): Kolmogorov–Smirnov statistic of Mes vs Epi signature ECDFs. Higher = more mesenchymal.
  • MLR (George 2017 / Chakraborty 2020): multinomial-logistic-regression predictor. Higher = more mesenchymal. Pipeline: htseq raw counts → countToTpm (TPM, log2) → rnaToMA (MA = 0.57 + 0.37·log2TPM, RNA-seq→microarray-scale) → three scoring functions → pairwise Pearson correlations (all_scoreCor).

In scope (pipeline-derived, attempted)

id result pipeline source
C1 Py2T cells have higher 76GS than MTΔECad countToTpm+EMT76GS (repo R, verbatim) Fig 2 / Results text (PMC)
C2 Py2T cells have lower KS than MTΔECad countToTpm+KSScore (repo R, verbatim) Fig 2 / Results text (PMC)
C3 76GS and KS are negatively correlated (R<−0.3, p<0.05) all_scoreCor Fig 1 / Table S1 overall trend

Reported anchor (PMC, Results): "Py2T murine epithelial tumor cells … had lower KS and MLR scores but higher 76GS scores as compared to the MTΔEcad cells."

Out of scope (the hard ~20%, not attempted) — with reason

  • MLR metric: requires MATLAB (MLR3_Code/MLR3_automated.m) + a proprietary trained model RelevantData.mat. MATLAB is not installed/licensed; the model blob is not independently re-derivable. → MLR score and the two MLR-involving correlations (76GS–MLR, KS–MLR) are dropped. We still verify the KS and 76GS halves of the Fig-2 claim (C1,C2) and the 76GS–KS correlation (C3).
  • Full 77-dataset survey / Table S1 exact per-dataset r-values: the paper's per-dataset numbers live in a supplementary xlsx behind an MDPI 403; the main text reports only the aggregate (44/77 = 57.14 % datasets concordant) and the R≷±0.3 threshold, not a single-dataset r we can pin. We reproduce ONE dataset (GSE118612, the one named in the brief) and grade its directional claims; we do not claim to reproduce the survey.
  • Single-cell analyses (Fig 3): different datasets, out of this RU's data scope.

Adaptation note (faithful, documented)

GEO's GSE118612_counts.txt is keyed by Entrez GeneID; the repo's Annotation/mm10_gene_length_kb.txt is keyed by Ensembl ID → symbol → length. The authors evidently ran on Ensembl-keyed counts. We bridge Entrez→Ensembl using NCBI Mus_musculus.gene_info (dbXrefs Ensembl field), reformat the matrix into the repo's expected layout (col1=Ensembl, col2=symbol, cols3+=counts), then run the authors' countToTpm / EMT76GS / KSScore unchanged. Only input wrangling is ours; the scoring code is verbatim.

Determinism caveat

EMT76GS injects rnorm(n, sd=0.01) with no set.seed → 76GS scores are mildly non-deterministic run-to-run (noise is tiny vs signal; group ordering is stable). We fix a seed for our run and also report run-to-run spread.

Figures / tables: Fig 2Fig 1Table
C1
Reported
Py2T cells have higher 76GS than MTdEcad (Fig 2, directional)
Reproduced
76GS mean Py2T=+4.613 vs MTdEcad=-6.919 (Py2T higher; every Py2T sample > every MTdEcad sample)
exact
C2
Reported
Py2T cells have lower KS than MTdEcad (Fig 2, directional)
Reproduced
KS mean Py2T=0.264 vs MTdEcad=0.365 (Py2T lower; strict per-sample)
exact
C3
Reported
76GS negatively correlated with KS (R<-0.3, p<0.05; Fig 1 overall trend)
Reproduced
Pearson r(76GS,KS)=-0.987965, p=9.046e-08
exact
MLR
Reported
MLR EMT metric + 76GS-MLR & KS-MLR correlations (Fig 1/2, Table S1)
Reproduced
nicht durchgefuehrt (kein Grund vermerkt)
m.public.grade.not-attempted

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

The in-scope claims (C1 76GS Py2T>MTdEcad, C2 KS Py2T<MTdEcad, C3 76GS↔KS negative at R<-0.3/p<0.05) were reproduced 1:1 directionally by running the authors' own scoring code verbatim on the public GEO deposit GSE118612 (reproduced r=-0.988, p=9e-8; separation >11 score units, seed-independent). Nothing sits on the authors' side — every reported trend is derivable from the shipped code plus public data, no fabrication signal. The only limitations are that the paper prints no single numeric value for this dataset (so the endpoint match is directional, q2 yellow) and that the MLR metric and exact Table S1 r-values were out of reach (MATLAB+proprietary model; MDPI 403) — availability/tooling gaps, not discrepancies. Overall a solid, faithful reproduction of the central conclusion.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

159.1 k
tokens (I/O) · 10.7 M incl. cache
37 min
runtime · 0.01 CPU-h
2.6 GB
peak RAM
1
HPC jobs
hummel
machine