Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

TGF-β-dependent reprogramming of amino acid metabolism induces epithelial-mesenchymal transition in non-small cell lung cancers.

Commun Biol · 2021
L1 85/100 PQI 95
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Reported values are derivable from the shared data
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the survival result 1:1 in DIRECTION + SIGNIFICANCE. Reproduced Fig 6a — the prognostic association of P4HA3 (probe 228703_at) with poor overall survival in all three independent NSCLC cohorts the paper names (GSE3141 [the brief's accession], GSE30219, GSE31210), pulled directly from GEO series matrices on «our HPC» (GEOquery -> survival/survminer; P16 — corrplot is only the paper's plotting pkg for the unrelated correlation figures). Using the paper's stated split method (KM-Plotter 'auto select best cutoff' ~ survminer::surv_cutpoint), all three give HR>1 and log-rank P<0.05 (0.021 / 0.0006 / 0.0026), matching the paper's qualitative claim; KM curve saved for GSE3141. Honest caveat / robustness flag: the paper prints no per-dataset HR or P, and with a pre-specified MEDIAN split all three cohorts are non-significant (P=0.12-0.14) though still directionally consistent (HR>1) — so the reported significance is method-dependent (relies on threshold optimisation, which inflates significance and is what KM-Plotter does). Reproducible and legitimate under the stated method, not a fabrication, but the prognostic signal is borderline under a fixed cutoff. NOT attempted (hard ~20%): the corrplot correlation matrices (Fig 4b/5d; need CCLE metabolome+transcriptome + 76-gene EMT score, plus wet-lab cell-line metabolomics), TCGA Fig 6b, and all wet-lab figures — out of the clean GSE3141 pipeline scope.

💻 Code ↗ 🗄 Data: GSE3141

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 85
    assessed: 2026-06-15 ⛓ 1a9c210b3e6d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper tests whether and how TGF-β-induced reprogramming of intracellular amino acid metabolism is integrated with and required for epithelial-mesenchymal transition (EMT) in non-small cell lung cancer cells.

Core claims
  • TGF-β induces reprogramming of intracellular amino acid metabolism that is necessary to promote EMT in NSCLC cells finding
  • Prolyl 4-hydroxylase α3 (P4HA3) is upregulated during TGF-β stimulation and is required for TGF-β-dependent changes in amino acids, EMT, and tumor metastasis mechanism
  • Manipulation (depletion) of extracellular amino acids induces EMT-like responses without TGF-β stimulation finding
  • Integrated metabolomic and transcriptomic two-layer omics screening identifies P4HA3 as a key amino acid metabolism enzyme in EMT method
  • P4HA3 inhibitor is proposed as a potential therapeutic agent for cancer resource
  • P4HA3 overexpression alone is insufficient to induce EMT finding
  • P4HA3 expression correlates with EMT markers and EMT score across NSCLC cell lines in CCLE finding
Experimental setups
Assay System Perturbation Readout Platform
CE-TOFMS metabolomics (polar metabolites/amino acids) A549, HCC827, H358 NSCLC cell lines TGF-β treatment (2-5 ng/mL, 3 days to several weeks) polar metabolite/amino acid levels capillary electrophoresis time-of-flight mass spectrometry (CE-TOFMS)
Time-course metabolome analysis A549 cells TGF-β 5 ng/mL for 24, 48, 72 h amino acid levels (fold-change) CE-TOFMS
Real-time PCR (mRNA EMT markers) A549, SW1573 cells TGF-β stimulation/withdrawal; amino acid depletion media mRNA of CDH1, CDH2, FN1, ZEB1, MMP2, MMP9
Western blotting A549, SW1573 cells amino acid depletion; TGF-β; P4HA3 knockdown CDH1, CDH2, ZEB1, P4HA3 protein levels
Amino acid depletion media culture A549, SW1573 cells depletion of combined or single amino acids for 72 h EMT marker expression, cell growth, cell morphology/circularity
siRNA knockdown A549, HCC827, H358, SW1573 cells P4HA3 siRNA +/- TGF-β EMT marker mRNA, amino acid metabolome
P4HA3 overexpression A549, HCC827 cells P4HA3 overexpression EMT marker expression
Bioinformatic correlation analysis NSCLC cell lines (CCLE; 187/147 lines) and clinical datasets GSE3141, GSE30219, GSE31210 none P4HA3 expression vs EMT score, EMT markers, amino acid levels
Key results
  • 21 metabolic pathways altered by TGF-β across cell lines, with amino acid metabolism commonly altered
  • TGF-β (72 h) increased Asp, Glu, Lys and decreased Ala, Asn, citrulline, Gln, Gly, His, hydroxyproline, Ile, Leu, Phe, Pro, Thr, Tyr in A549
  • Amino acid changes were reversed after TGF-β withdrawal (MET), except for Asp
  • Amino acid depletion media induced EMT-like responses (down CDH1, up CDH2/ZEB1 mRNA) and elongated morphology
  • P4HA3 knockdown increased CDH1 and decreased CDH2, FN1, MMP9, MMP2, abrogating TGF-β EMT
  • P4HA3 knockdown abrogated TGF-β-mediated amino acid metabolic changes
  • P4HA3 expression and EMT score significantly correlated with citrulline, Arg, ornithine, His, and Lys levels in CCLE
  • P4HA3 overexpression produced insignificant changes in EMT markers
Key statistics
  • count 21 pathways altered (metabolic pathways changed by TGF-β across cell lines)
  • count 187 lung cancer cell lines (CCLE dataset for P4HA3-EMT correlation)
  • count 147 lung cancer cell lines (CCLE metabolome/transcriptome dataset for P4HA3-amino acid correlation)
  • count 76 genes (genes used to calculate EMT score in NSCLC)
  • other P < 0.05 (FDR <0.05) (pathway significance threshold in MPEA)
  • pvalue P < 0.01 (cell circularity change with amino acid depletion in A549)
  • other 5-year overall survival <20% (prognosis of NSCLC)
  • count n = 4 (replicates for metabolome heat map / PCA)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a multi-omics experimental design (metabolomics via CE-TOFMS combined with transcriptomics) in NSCLC cell lines treated with TGF-β, with comparisons made between stimulated and unstimulated (often paired) conditions and across knockdown/overexpression and amino-acid-depletion conditions. Dimensionality reduction (PCA) and metabolic pathway enrichment analysis (MPEA) were applied to metabolite data, with pathway significance judged by P values and false discovery rates (<0.05), and correlation analyses were used for CCLE datasets. Results were largely reported as heat maps of fold-changes relative to paired controls and as mean ± SD from replicate samples, with significance indicated by thresholds (e.g., **P < 0.01).

Replicationunclear Sample sizen = 4 for metabolome time-course/PCA experiments; triplicate samples for real-time PCR and circularity; CCLE analyses based on 187 and 147 cell lines; no formal power/sample-size justification stated GroupsTGF-β-stimulated vs unstimulated cells; withdrawal (MET); amino-acid-depleted vs full medium; P4HA3 knockdown/overexpression vs control siRNA Pairingmixed Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionFalse discovery rate (FDR < 0.05)
Statistical tests used
Test Applied to n Assumptions
Metabolic pathway enrichment analysis (MPEA) with P values and false discovery rate Identification of altered metabolic pathways in TGF-β-stimulated A549, HCC827, H358 cells (Fig. 1a) not stated
Principal component analysis (PCA) Metabolomic profiles of TGF-β-stimulated vs unstimulated and P4HA3-knockdown cells (Supplementary Fig. 1g–i; Fig. 5b) n = 4 na
Correlation analysis (significance by P value, P > 0.05 marked non-significant) P4HA3 mRNA vs EMT markers/EMT score and amino acid levels in CCLE NSCLC cell lines (Fig. 4b, Fig. 5d, Supplementary Figs. 4e, 5g) 187 and 147 lung cancer cell lines (as stated) not stated
Unspecified significance test reported as a P-value threshold Cell circularity after amino acid depletion in A549 cells (Fig. 3e, **P < 0.01) triplicate samples not stated
Approaches that could also have been used
  • Replicate data (e.g., real-time PCR, circularity) were summarized as mean ± SD from triplicate samples.
    Could also: Reporting could additionally include a 95% confidence interval or showing individual data points alongside the mean. — For small replicate numbers, plotting individual points and a CI can convey both the spread and the precision of the estimate, complementing the SD.
  • Significance was conveyed using P-value thresholds such as **P < 0.01 and P < 0.05.
    Could also: Exact P values together with effect-size estimates (e.g., fold-change with confidence intervals) could also be reported. — Exact P values and effect sizes give readers a continuous sense of evidence strength and practical magnitude beyond a pass/fail threshold.
  • Multiple group comparisons (e.g., several amino-acid-depletion conditions, multiple cell lines) were assessed individually.
    Could also: A single ANOVA model with a post-hoc multiple-comparison correction (e.g., Tukey HSD or Dunnett's against control) could also be applied. — An omnibus model with post-hoc correction controls the family-wise error rate when many conditions are compared simultaneously.
  • Pathway significance was assessed with P values and FDR (< 0.05) via MPEA.
    Could also: The specific multiple-testing method (e.g., Benjamini-Hochberg) and the family of tests it covered could be stated explicitly. — Naming the correction method and its scope helps readers reproduce the enrichment analysis and interpret the FDR threshold.
  • Correlations in the CCLE datasets were reported with a significance cutoff (P > 0.05 marked non-significant).
    Could also: The correlation type (Pearson vs Spearman), the coefficient values, and confidence intervals could also be reported. — Specifying the method and providing coefficients with CIs clarifies the strength and direction of association, especially when relationships may be non-linear.
  • The specific hypothesis test underlying threshold annotations (e.g., **P < 0.01 for circularity) is not named in the text.
    Could also: Stating the exact test used (e.g., two-tailed t-test or Mann-Whitney U) and whether assumptions were checked could also be included. — Naming the test and its assumptions aids reproducibility and helps readers judge fit between the data distribution and the chosen method.
Software: ImageJ (used for quantification of protein bands and cell circularity)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
52
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE136780 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34168290

Paper: Nakasuka et al. 2021, Commun Biol 4:782. "TGF-β-dependent reprogramming of amino acid metabolism induces epithelial–mesenchymal transition in non-small cell lung cancers." DOI 10.1038/s42003-021-02323-7 · PMCID PMC8225889.

Brief-listed artifacts: Code = github.com/taiyun/corrplot (generic R plotting package → P16: a third-party tool, not the authors' code; reproduce by running the described pipeline on the paper's own / public data). Data = GEO:GSE3141.

Which reported results are pipeline-derived (in scope) vs not

Result Figure Pipeline Data In scope?
P4HA3 expression vs NSCLC prognosis (high P4HA3 → poor survival, P<0.05) Fig 6a KM survival (KM-Plotter); probe 228703_at; auto-best-cutoff; univariate Cox HR+P GSE3141, GSE30219, GSE31210 (all GPL570, public) YES — primary
Correlation of P4HA3 / amino acids / EMT score / EMT genes Fig 4b, 5d Pearson correlation → corrplot viz CCLE metabolome (suppl) + transcriptome; cell-line metabolome (wet-lab) Partial — needs CCLE download + 76-gene EMT score; secondary
P4HA3 in tumor vs non-tumor / TNM Fig 6b TCGA expression query TCGA lung Out (different data source, not the brief's GSE3141)
TGF-β metabolome / EMT wet-lab assays (CE-MS, qPCR, WB, migration) Figs 1-3,5 wet-lab OUT (not pipeline)

What we attempt (80/20)

Primary (clear, low-hanging, uses the brief's GSE3141): reproduce Fig 6a — the Kaplan-Meier association between P4HA3 (probe 228703_at) expression and overall survival in NSCLC, on GSE3141 and (as the other two "independent studies" named in the same sentence) GSE30219 and GSE31210. Method per paper: KM-Plotter with auto-best-cutoff + univariate Cox. We reproduce the equivalent pipeline directly from the GEO series matrices (GEOquery → probe 228703_at → median split AND best-cutoff via survminer::surv_cutpoint, restricted to 25–75% like KM-Plotter → Cox HR + log-rank P). Expected: high expression → worse survival, P<0.05.

Not attempted (the hard ~20%): the corrplot Fig 4b/5d correlation matrices require CCLE metabolome+transcriptome assembly and the 76-gene EMT score, and the cell-line metabolome is wet-lab (CE-TOFMS) — out of the clean GSE3141 scope. TCGA Fig 6b and all wet-lab figures are out of scope. The paper reports only "P<0.05" qualitatively for Fig 6a (no per-dataset HR/P printed), so the comparison is on direction (high→poor) + significance, not an exact numeric match.

Possible-fabrication watch

The reported Fig 6a claim ("significant positive correlation, P<0.05") is checkable against the public GEO survival data; we record the actual HR + log-rank P per cohort so a human can see whether the direction/significance holds in each named dataset.

Figures / tables: Fig 6a
C1
Reported
Fig 6a: high P4HA3 (228703_at) -> poor NSCLC prognosis, P<0.05 (GSE3141; qualitative, no number printed)
Reproduced
best-cutoff HR=1.87 [1.09-3.20], log-rank P=0.021, n=111 (29 high/82 low, 58 events); median split HR=1.48, P=0.137 (n.s.)
within tolerance
C2
Reported
Fig 6a: high P4HA3 -> poor prognosis, P<0.05 (GSE30219)
Reproduced
best-cutoff HR=1.84 [1.29-2.63], log-rank P=0.0006, n=293; median split HR=1.24, P=0.122 (n.s.)
within tolerance
C3
Reported
Fig 6a: high P4HA3 -> poor prognosis, P<0.05 (GSE31210)
Reproduced
best-cutoff HR=2.67 [1.37-5.21], log-rank P=0.0026, n=226; median split HR=1.66, P=0.137 (n.s.)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

Fig 6a's qualitative claim — high P4HA3 (228703_at) associates with poor NSCLC overall survival, P<0.05 — reproduces in all three named GEO cohorts under the paper's stated best-cutoff method (HR=1.87/1.84/2.67, P=0.021/0.0006/0.0026), directly from public series matrices, and is not a fabrication. The key caveat is on the methodology side, shared between us and the authors: significance hinges on the threshold-optimising 'auto best cutoff' (a known significance-inflating procedure); under a pre-specified median split all three cohorts are non-significant (P=0.122–0.137) while direction (HR>1) is preserved. Severity is moderate — the central conclusion holds in direction and under the stated method, but the strength of the prognostic signal is borderline and method-sensitive, which the reproduction honestly flags.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

126.8 k
tokens (I/O) · 7.5 M incl. cache
14 min
runtime · 0.02 CPU-h
2.8 GB
peak RAM
2
HPC jobs
hummel
machine