Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Involvement of N4BP2L1, PLEKHA4, and BEGAIN genes in breast cancer and muscle cell development.

Front Cell Dev Biol · 2024
L1 73/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
73/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 41% of all assessed papers rank 664 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (DE strong, survival directional). Paper's named tool GDCRNATools 1.22.0 run end-to-end on «our HPC» against GDC TCGA-BRCA (1095 PrimaryTumor + 113 SolidTissueNormal, 60616 genes merged, 15048 tested). PRIMARY result (Table 3 limma DE) reproduces faithfully for ALL 3 genes: logFC within ~4-5% with identical sign and the same extreme-significance regime (N4BP2L1 -1.462->-1.538; PLEKHA4 -1.225->-1.273; BEGAIN -1.406->-1.481). SECONDARY result (Table 2 coxph OS) reproduces in direction (all protective, HR<1) and roughly magnitude, with drift: N4BP2L1 HR 0.725->0.776 (sig), PLEKHA4 0.612->0.851 (sig), BEGAIN 0.691->0.890 (p 0.0225->0.0578, borderline non-sig) — consistent with the GDC cohort having grown since 2024 + median-split sensitivity. Gene-mapping fix: prior prep used ENSG00000064601 for BEGAIN, which is CTSA; corrected by extracting genes by symbol (BEGAIN=ENSG00000183092). Engineering blockers solved along the way: gdc-client AF_UNIX-path-too-long (short node-local TMPDIR), node /tmp is tmpfs, and account-wide «infra»+home quota exhaustion (job made fully self-contained: submit via sbatch stdin, all bulk on node-local tmpfs, KB results retry-copied to «infra»). No fabrication signal: our independently-run values land in the paper's regime. NOT attempted (out of scope): web-portal (UALCAN/TACCO/GTEx/HPA/CPTAC) + wet-lab (RT-qPCR/IHC/in-vitro muscle) + GSE182338 GREIN/edgeR Table 5.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-14 ⛓ ef3e422d86f1
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether differential mRNA expression analysis of TCGA-BRCA data can identify previously uncharacterized, downregulated genes (N4BP2L1, PLEKHA4, BEGAIN) that function as tumor suppressors in breast cancer and are linked to myoepithelial/smooth muscle cell identity and muscle development.

Core claims
  • N4BP2L1, PLEKHA4, and BEGAIN, normally highly expressed in breast myoepithelial and smooth muscle cells, are significantly downregulated in breast tumor tissue of a 50-patient cohort finding
  • Low PLEKHA4 expression in patients with menopause below 50 years correlates with higher breast cancer risk finding
  • N4BP2L1 and BEGAIN are potential biomarkers of HER2-positive breast cancer finding
  • Low BEGAIN expression in patients with blood fat, heart problems, and diabetes correlates with higher breast cancer risk finding
  • N4BP2L1 protein and RNA expression analysis of TCGA-BRCA identifies it as a promising diagnostic protein biomarker in breast cancer finding
  • scRNA-seq in silico data show high expression of N4BP2L1, PLEKHA4, and BEGAIN in several normal breast tissue cell types, including myoepithelial and smooth muscle cells finding
  • GDCRNATools pipeline (gdcParseMetadata, gdcFilterDuplicate, gdcFilterSampleType, gdcVoomNormalization, gdcDEAnalysis, gdcDEReport, gdcSurvivalAnalysis) used to identify differentially expressed genes with prognostic significance in TCGA-BRCA method
  • H9 hESCs were differentiated into skeletal muscle cells over 46 days to validate expression of the three genes during muscle development method
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq differential expression (TCGA-BRCA) human breast tumor vs. matched adjacent non-tumor tissue none (tumor vs normal comparison) differential mRNA expression and overall survival association GDCRNATools (R/Bioconductor)
RT-qPCR 50 paired primary breast tumor and adjacent normal tissues, Iranian patient cohort none (patient tissue comparison) relative mRNA expression (ΔΔCt) of N4BP2L1, PLEKHA4, BEGAIN normalized to β2M RealQ Plus 2x Master Mix Green with HighROX
RT-qPCR H9 human embryonic stem cell-derived skeletal muscle cells (D0, D2, D7, D30, D46) directed differentiation (hESC to skeletal muscle) relative expression of PLEKHA4, N4BP2L1, BEGAIN, TBXT, PAX3, MYOG (PUM1 reference gene)
immunofluorescence staining H9 hESC-derived muscle cells, day 46 terminal differentiation directed differentiation MF20 (myosin heavy chain) protein expression DSHB antibody AB_2147781 / Alexa Fluor 488
bulk and single-cell RNA-seq (public database) normal human tissues (GTEx) none N4BP2L1, PLEKHA4, BEGAIN expression across tissues/cell types Human Protein Atlas
single-cell RNA-seq human skeletal muscle, embryonic/fetal/postnatal developmental stages none (developmental time course) gene expression across muscle developmental stages Human Skeletal Muscle Atlas
bulk RNA-seq (GREIN reanalysis) mammary myoepithelial (MEP) vs luminal epithelial (LEP) cells from primary breast organoids none (cell-type comparison) differential expression of N4BP2L1, PLEKHA4, BEGAIN, TP63, SERPINB5 GREIN (GSE182338)
immunohistochemistry human breast tissue none protein expression levels Human Protein Atlas
Key results
  • N4BP2L1, BEGAIN, and PLEKHA4 downregulated in TCGA-BRCA tumor vs normal samples
  • Several genes upregulated in TCGA-BRCA tumors (SLC35A2, DONSON, BRI3BP, NDUFAF6, C2CD4D, SEZ6L2)
  • Low PLEKHA4 expression associated with higher breast cancer risk in patients with menopause before age 50
  • N4BP2L1 and BEGAIN expression differ in HER2-positive breast cancer, supporting use as biomarkers
  • Low BEGAIN expression correlates with higher cancer risk in patients with blood fat, heart problems, and diabetes
  • N4BP2L1 RNA and protein expression analyses support its use as a diagnostic biomarker in breast cancer
  • scRNA-seq shows high expression of the three genes in myoepithelial and smooth muscle cells of normal breast tissue
Key statistics
  • count 1,222 RNA sequencing data (TCGA-BRCA dataset used for differential expression analysis)
  • count 1,097 clinical data records (TCGA-BRCA dataset used for survival analysis)
  • count 50 primary tumors and paired normal tissues (Iranian breast cancer patient cohort for RT-qPCR validation)
  • pvalue p < 0.05 (Threshold defined as statistically significant across all statistical tests)
  • other RNA integrity number 8–10 (RNA quality assessed by Agilent Bioanalyzer 2100 for muscle differentiation samples)
  • other 45 cycles: 95°C for 20 s, 60°C for 30 s (RT-qPCR cycling conditions)
  • other 46-day differentiation protocol (D0, D2, D7, D30, D46) (H9 hESC to skeletal muscle differentiation timeline)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper combines TCGA-BRCA bioinformatic analysis with RT-qPCR validation in 50 paired Iranian patient samples to characterize three candidate tumor suppressor genes. Differential expression in TCGA was assessed via a limma-voom pipeline (GDCRNATools), followed by Kaplan–Meier univariate survival analysis. In the patient cohort, paired and non-parametric tests compared tumor vs. adjacent-normal tissue expression, and associations with >30 clinicopathologic and demographic variables were evaluated with Mann–Whitney, Kruskal–Wallis, chi-square, and independent t-tests. IBM SPSS 26 was used for the cohort analyses, with p < 0.05 as the significance threshold.

Replicationbiological Sample size50 patient pairs stated for RT-qPCR cohort; 1,222 RNA-seq and 1,097 clinical TCGA-BRCA records stated; no formal power calculation or sample size justification mentioned GroupsTumor vs. adjacent non-tumor tissue; high vs. low gene expression (median split); numerous clinicopathologic and demographic subgroups (>30 variables); TCGA cancer types via TACCO; MEP vs. LEP cells; muscle differentiation time points (D0, D2, D7, D30, D46) Pairingmixed Randomization/blindingnot stated Dispersionmixed Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Paired sample t-test ΔΔCt comparison of RT-qPCR expression between tumor and paired adjacent non-tumor tissue 50 paired samples not stated
Mann–Whitney U test Log fold-change comparison between tumor and adjacent non-tumor tissue; also association between gene expression and clinicopathologic/demographic factors (two-group variables) 50 paired samples not stated
Kruskal–Wallis test Association between gene expression and clinicopathologic/demographic factors with more than two groups 50 samples (subgroup sizes not stated) not stated
Chi-square test Comparison of high vs. low expression groups (dichotomized by median) across categorical clinicopathologic variables 50 samples not stated
Independent samples t-test Comparison of continuous variables between high and low expression groups (dichotomized by median log fold-change) 50 samples not stated
Kaplan–Meier univariate survival analysis (gdcSurvivalAnalysis, log-rank implied) Overall survival by high vs. low expression groups across differentially expressed genes in TCGA-BRCA 1,097 clinical records from TCGA-BRCA not stated
Limma-voom differential expression analysis (gdcVoomNormalization + gdcDEAnalysis via GDCRNATools) Differential mRNA expression: TCGA-BRCA tumor vs. matched adjacent non-tumor tissue 1,222 RNA-seq samples from TCGA-BRCA not stated
Differential expression analysis via GREIN pipeline (DESeq2 or limma backend, not explicitly specified) Mammary myoepithelial (MEP) vs. luminal epithelial (LEP) cells in GSE182338 primary breast organoids not stated not stated
Approaches that could also have been used
  • More than 30 clinicopathologic and demographic variables were each independently tested for association with gene expression using alpha = 0.05, without a stated multiplicity correction.
    Could also: Apply a false discovery rate (FDR) correction (e.g., Benjamini–Hochberg) or Bonferroni adjustment across the family of association tests. — When many simultaneous comparisons are made, the expected number of false positives at alpha = 0.05 grows with the number of tests; FDR or family-wise error rate control quantifies and limits this inflation, making the reported associations easier to interpret in terms of expected error rates.
  • Overall survival was analyzed by dichotomizing patients into high- and low-expression groups based on the median, followed by Kaplan–Meier log-rank testing.
    Could also: Model expression as a continuous predictor in a Cox proportional hazards regression, optionally adjusting for clinical covariates. — Median-based dichotomization discards information about the continuous relationship between expression level and survival and can inflate Type I error; Cox regression retains the full continuous scale, yields a hazard ratio with confidence interval as an effect size, and allows covariate adjustment for factors such as tumor stage or receptor status.
  • The paired tumor vs. adjacent-normal comparison used both a paired t-test (on ΔΔCt) and a Mann–Whitney test (on log fold-change) applied to the same 50 pairs.
    Could also: Apply a single Wilcoxon signed-rank test on ΔΔCt values as a non-parametric paired alternative, or confirm normality first (e.g., Shapiro–Wilk) and then select one primary test. — Running two tests on essentially the same data and outcome can obscure which result is primary; a pre-specified single test with a stated justification (parametric vs. non-parametric) is easier to interpret and avoids ambiguity about which p-value governs the conclusion.
  • Dispersion for ΔΔCt is reported as SD while central tendency for fold-change is reported as median (without accompanying spread measure for fold-change).
    Could also: Report median with IQR for both ΔΔCt and fold-change (if data are skewed), or mean ± SD for both (if approximately normal), and include a 95% confidence interval for the difference. — Consistent dispersion reporting across all outcomes aids comparability; pairing a median with IQR (or mean with SD) consistently, and adding a confidence interval for the group difference, conveys both spread and uncertainty in the effect estimate.
  • Gene expression associations with clinicopathologic subgroups were evaluated with separate univariate tests for each variable.
    Could also: Fit a multivariable linear or logistic regression model including the clinicopathologic variables of primary interest simultaneously. — Univariate analyses cannot distinguish independent associations from confounding; a multivariable model identifies which variables have an independent relationship with expression after accounting for others, which is particularly relevant when variables such as age, menopausal status, and BMI are correlated.
  • Survival analysis was performed as a univariate Kaplan–Meier comparison without stated covariate adjustment.
    Could also: Perform multivariate Cox regression adjusting for established prognostic factors (e.g., tumor stage, nodal status, receptor subtype). — Univariate survival associations may reflect confounding by established prognostic variables; a multivariate Cox model tests whether the gene expression association persists independently and is standard practice for reporting potential prognostic biomarkers.
Software: R/Bioconductor GDCRNATools · IBM SPSS 26 · geNorm · NormFinder · Enrichr (web tool) · GREIN (web tool, GEO RNA-seq Experiments Interactive Navigator) · UALCAN (web resource) · TACCO (web server)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 56/100
partly built on non-reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38859961

Paper: Dastsooz et al. (2024) Involvement of N4BP2L1, PLEKHA4, and BEGAIN genes in breast cancer and muscle cell development. Front Cell Dev Biol. PMID 38859961 · PMCID PMC11163233 · DOI 10.3389/fcell.2024.1295403

Named code artifact: GDCRNATools (https://github.com/Jialab-UCR/GDCRNATools), a third-party Bioconductor tool. Per BRIEF rule P16, applying this third-party tool to the paper's own data is a fully valid reproduction.

Genes of interest (Ensembl)

  • N4BP2L1 = ENSG00000139597
  • PLEKHA4 = ENSG00000105559
  • BEGAIN = ENSG00000064601

Result inventory & scope classification

# Reported result Source Pipeline In scope?
R1 TCGA-BRCA differential expression (logFC, AvgExpr, t, p, FDR) for the 3 genes — Table 3 TCGA-BRCA RNA-seq (public, GDC) GDCRNATools gdcVoomNormalization + gdcDEAnalysis(method='limma') — the paper's named tool YES — PRIMARY
R2 TCGA-BRCA overall-survival HR + 95% CI + p for the 3 genes — Table 2 TCGA-BRCA clinical+expr GDCRNATools gdcSurvivalAnalysis (coxph) YES — SECONDARY
R3 Myoepithelial vs luminal DE (logFC/logCPM/p/FDR), Table 5 GSE182338 organoid RNA-seq GREIN web app (edgeR) Maybe (stretch) — different tool (web), 80/20 deprioritized
R4 Subtype/stage/nodal stratification p-values (Suppl S1) TCGA-BRCA UALCAN web portal OUT — external web portal, not a runnable pipeline
R5 Cross-cancer downregulation lists all-TCGA TACCO web portal OUT — external web portal
R6 GTEx nTPM, Human Protein Atlas single-cell nTPM GTEx / HPA portals portal lookup OUT — external portal values, not a pipeline we run
R7 CPTAC proteomics p-value CPTAC / UALCAN portal lookup OUT — external portal
R8 RT-qPCR in 50-patient cohort (Fig 3) wet-lab OUT — wet-lab
R9 IHC protein expression, in-vitro muscle differentiation (Fig 7) wet-lab OUT — wet-lab

Reproduction plan (80/20)

Reproduce R1 (primary) and R2 (secondary) with GDCRNATools on TCGA-BRCA, following the documented standard workflow verbatim:

  1. gdcRNADownload(project.id='TCGA-BRCA', data.type='RNAseq')
  2. gdcParseMetadatagdcFilterDuplicategdcFilterSampleType
  3. gdcVoomNormalization(counts, filter=FALSE)
  4. gdcDEAnalysis(counts, group=sample_type, comparison='PrimaryTumor-SolidTissueNormal', method='limma')
  5. extract Table-3 rows for the 3 genes + total DE-gene count
  6. gdcSurvivalAnalysis(method='coxph', ...) → Table-2 HR/CI/p for the 3 genes

The hard last ~20% (R3 web-tool edgeR on GSE182338; all R4–R9 portal/wet-lab results) is not attempted — those are external web portals or wet-lab and are not faithfully reproducible as a runnable pipeline on our side.

All compute runs on «our HPC»; all data stays on «infra».

Figures / tables: Table
DE_N4BP2L1_logFC
Reported
-1.462 (p=7.72E-91, FDR=2.84E-89)
Reproduced
-1.5377 (p=1.50E-95, FDR=5.72E-94)
within tolerance
DE_PLEKHA4_logFC
Reported
-1.225 (p=1.40E-27, FDR=6.87E-27)
Reproduced
-1.2734 (p=3.63E-29, FDR=1.91E-28)
within tolerance
DE_BEGAIN_logFC
Reported
-1.406 (p=8.85E-28, FDR=4.39E-27)
Reproduced
-1.4814 (p=1.83E-31, FDR=1.06E-30)
within tolerance
SURV_N4BP2L1_HR
Reported
HR=0.725 (0.526-0.999), p=0.0472
Reproduced
HR=0.7758 (0.6444-0.9339), p=0.00731
within tolerance
SURV_PLEKHA4_HR
Reported
HR=0.612 (0.445-0.842), p=0.00282
Reproduced
HR=0.8514 (0.7497-0.9670), p=0.01323
partial
SURV_BEGAIN_HR
Reported
HR=0.691 (0.501-0.951), p=0.0225
Reproduced
HR=0.8898 (0.7888-1.0039), p=0.05779
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 73/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is an incomplete run, not a discrepancy: data identity (public TCGA-BRCA via GDC) and endpoint comparability (limma logFC/p/FDR, coxph HR/CI/p) are both clean and 1:1-feasible using the authors' own GDCRNATools pipeline. However, «our HPC» «job» was still downloading the ~1,200 count files when the operator forced finalize, so no reproduced value was computed or graded for any of the 6 target claims. The deviation cannot be located, sized, or judged because nothing exists to compare; the cause is our operational truncation, with no sign of an authors'-side defect or fabrication. Overall: solid, well-scoped attempt that simply did not finish — yellow on all unverified axes pending a human pull of the «infra» outputs.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

433.5 k
tokens (I/O) · 36 M incl. cache
241 min
runtime · 192.08 CPU-h
peak RAM
2 (2 failed)
HPC jobs
hummel
machine