Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

An engineered tumor organoid model reveals cellular identity and signaling trajectories underlying SFPQ-TFE3 driven translocation RCC.

iScience · 2025
L1 88/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
88/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 74% of all assessed papers rank 276 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> 1:1 EXACT on the single-cell QC unit. The BRIEF mis-paired artifacts: it gave the BULK repo (teresouza/ganpat2023) with the SINGLE-CELL GEO accession (GSE260633); the paper actually has 3 units (bulk mRNA E-MTAB-13901, CUT&RUN E-MTAB-13900, single-cell GSE260633+Zenodo 10732016). Using Scanpy (third-party, rule P16) on the PUBLIC processed CellRanger matrices and applying the paper's stated QC (2000-100000 transcripts, mito<=50%), I reproduce exactly 863 e16 (wk16) cells passing QC, matching the authors' deposited pipeline (03-celltyping.R:190 '1661 -> 863 cells'). Also EXACT: the deposited e16 library = 1000 droplets = 613 ctrl + 387 fus (public demux map), and the QC thresholds match config-e16.R verbatim. The mito filter is a provable no-op (matrices ship 0 MT- genes), so the count filter alone yields 863. Re-run cleanly on a fresh «our HPC» job (2219174, n093) with env+data re-fetched into «infra». EXTENDED this pass: found config-e17.R uses a stricter min_txpts=3000 for wk17 (not stated in the paper text) and quantified that the wk17 (e17) analyzed set is NOT reconstructable from public data (requires a non-public luciferase-cell selection list + a deconfounded prep RDS, neither deposited) — a concrete reproducibility gap; the raw public-matrix e17 QC counts (lx419=600, lx420=467 at min3000) are reported as observations only. NOT attempted (80/20, the hard tail): the bulk DEG claim (75 DEGs / 8 CLEAR motif) — ganpat2023 hard-codes the author's laptop paths for every input and the CLEAR-motif step is absent from the repo; CUT&RUN; the stochastic pseudotime/SCENIC/trajectory analyses. Audit flag (not fabrication): the preliminary Zenodo deposit's Fig4c type-per-sample TSV sums to 1054 typed cells, inconsistent with the 863 figure -- 863 is nonetheless fully derivable from the shipped public data.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-14 ⛓ 90a422cb216f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether expression of the MiT/TFE fusion oncogene SFPQ-TFE3 is sufficient to transform normal human kidney tubular epithelial cells (tubuloids) into translocation renal cell carcinoma (tRCC), and seeks to define the cellular origin and signaling trajectories underlying this transformation.

Core claims
  • SFPQ-TFE3 expression is sufficient to transform normal kidney epithelial tubuloids into tRCC finding
  • SFPQ-TFE3-expressing tubuloids grow as clear cell RCC upon orthotopic xenotransplantation in mice finding
  • SFPQ-TFE3 reprograms gene expression via widespread, aberrant genome-wide DNA binding mechanism
  • Single-cell RNA-seq reveals that proximal tubular epithelial cells (or their progeny) are at the root of the tRCC transformation trajectory finding
  • Engineered SFPQ-TFE3-expressing tubuloid organoids serve as a representative human model of tRCC resource
  • Tub Fus-specific CUT&RUN peaks are enriched for the MITF motif, suggesting gain-of-function DNA binding beyond typical TFE3 sites mechanism
  • Disbalanced WNT signaling (decreased TCF7L2 activity, decreased RNF43 expression) is implicated in transformation toward tRCC mechanism
  • FOXP1 and HMGA2 transcription factor activity is decreased in fusion-expressing tubuloids, implicating them in malignant transformation finding
Experimental setups
Assay System Perturbation Readout Platform
RT-qPCR human kidney tubuloids (wk16, wk17) lentiviral expression of SFPQ-TFE3, TFE3, or control (luciferase) SFPQ-TFE3 and TFE3 mRNA expression
Western blot human kidney tubuloids and patient-derived tRCC organoids lentiviral expression of SFPQ-TFE3, TFE3, or control SFPQ-TFE3 fusion protein expression (β-actin loading control)
Histology (H&E) and immunostaining human tubuloids and patient-derived organoids SFPQ-TFE3/TFE3/control expression morphology, TFE3 nuclear localization, Ki67 proliferation
Orthotopic xenotransplantation immunodeficient mice (kidney) transplantation of Tub Ctrl or Tub Fus organoids tumor formation, histology (H&E, TFE3 IHC, human-specific keratin IHC)
Bulk RNA-seq Tub Ctrl, Tub TFE3, Tub Fus tubuloids, Tub Fus xenografts, and patient-derived kidney cancer organoids SFPQ-TFE3/TFE3/control expression gene expression profiles, hierarchical clustering
CUT&RUN Tub Ctrl and Tub Fus tubuloids SFPQ-TFE3 vs control expression genome-wide TFE3/fusion DNA binding peaks, motif enrichment, GSEA CUT&RUN
Single-cell RNA-seq (scRNA-seq) wk16 and wk17 Tub Ctrl and Tub Fus tubuloids SFPQ-TFE3 vs control expression cell type composition, EMT score, pseudotime trajectories, transcription factor activity (SCENIC) Chromium 10x Genomics
Key results
  • All mice transplanted with Tub Fus organoids developed ccRCC tumors, while none transplanted with Tub Ctrl did 5/5 vs 0/4
  • Tub Fus organoids and Tub Fus-derived xenografts clustered with patient-derived tRCC organoids by RNA-seq, unlike Tub Ctrl/Tub TFE3 which clustered with normal tubuloids
  • CUT&RUN identified peaks shared between Tub Ctrl and Tub Fus and additional peaks specific to Tub Fus, with no Tub Ctrl-specific peaks detected 2,774 shared vs 2,653 Tub Fus-specific peaks
  • Tub Fus-specific peaks were highly enriched for the MITF motif, unlike shared peaks enriched for the TFE3 CLEAR element
  • A subset of fusion-bound genes were differentially expressed in Tub Fus vs Tub Ctrl, suggesting direct transcriptional targets 75 genes (24.3% of all DEGs)
  • Cells with proximal tubule/CNT/principal cell signatures were underrepresented in Tub Fus scRNA-seq clusters compared to Tub Ctrl
  • TCF7L2 and RNF43 expression decreased along the trajectory from tubular epithelium to the presumed tRCC population
  • FOXP1 and HMGA2 transcription factor activity was decreased in both wk16 and wk17 Tub Fus models by SCENIC analysis
Key statistics
  • count 5/5 mice with tumors (Tub Fus) vs 0/4 (Tub Ctrl) (orthotopic xenotransplantation tumor incidence)
  • count 2,774 shared peaks; 2,653 Tub Fus-specific peaks (CUT&RUN genome-wide TFE3/fusion binding peaks)
  • fold_change 75 genes (24.3% of all differentially expressed genes) (fusion-bound genes that are also differentially expressed)
  • count 8 genes with CLEAR motif: COL14A1, RGS6, GRN, DTNBP1, GPNMB, EPHB1, RAB7A, SLC15A4 (differentially expressed fusion target genes containing CLEAR motif)
  • count Tub Ctrl n=3, Tub TFE3 n=2, Tub Fus n=3, PDO tRCC n=2, normal kidney PDO n=3, Wilms tumor n=3 (sample sizes for bulk RNA-seq hierarchical clustering)
  • count n=3 independent experiments (RT-qPCR quantification of SFPQ-TFE3 and TFE3 expression)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study employs a multi-omics approach combining bulk RNA-seq with unsupervised hierarchical clustering, CUT&RUN genome-wide binding analysis with de novo motif enrichment and GSEA, and 10x Genomics scRNA-seq with t-SNE visualization, diffusion-map pseudotime trajectory modeling, EMT module scoring, and SCENIC transcription factor activity inference. Comparisons were made between engineered tubuloid lines (Tub Ctrl, Tub TFE3, Tub Fus) and patient-derived organoids across two independent tubuloid models (wk16, wk17). Xenotransplantation outcomes were reported as proportions (0/4 vs 5/5) without a formal statistical test, and quantitative bench assays (RT-qPCR) were summarized as mean ± SD.

Replicationbiological Sample sizeIndependent experiments cited per group for bench assays (n=2–3); two independent tubuloid donor lines (wk16, wk17) used as cross-model validation; xenotransplantation used n=4 control and n=5 fusion mice GroupsTub Ctrl vs Tub Fus vs Tub TFE3 (in vitro); Tub Fus xenografts vs patient-derived tRCC (in vivo/histology); Tub Ctrl vs Tub Fus in wk16 and wk17 (scRNA-seq) Pairingmixed Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
Unsupervised hierarchical clustering (on differentially expressed genes) Bulk RNA-seq comparison of Tub Ctrl, Tub TFE3, Tub Fus, xenografts, Wilms tumor, normal kidney PDO, and PDO tRCC samples (Figure 3A, Figure S2A) Tub Ctrl n=3, Tub TFE3 n=2, Tub Fus n=3, Wilms tumor n=3, normal kidney PDO n=3, PDO tRCC n=2 (independent experiments each) not stated
Differential gene expression analysis (specific method not stated in provided text) Bulk RNA-seq: Tub Fus vs Tub Ctrl, yielding 75 directly bound differentially expressed genes (Figures S2D, S2E; Table S3) Tub Ctrl n=3, Tub Fus n=3 (independent experiments) not stated
Gene Set Enrichment Analysis (GSEA) Genes associated with Tub Fus-specific CUT&RUN peaks, testing for WNT pathway enrichment (Figure 3D) 2,653 Tub Fus-specific peaks; 75 differentially expressed direct target genes not stated
De novo motif enrichment analysis CUT&RUN peaks: shared (Tub Ctrl + Tub Fus) vs Tub Fus-specific peaks (Figure 3C) 2,774 shared peaks; 2,653 Tub Fus-specific peaks not stated
Diffusion map pseudotime trajectory modeling scRNA-seq data modeling differentiation trajectories from tubular epithelium toward tRCC-like cells (Figures 4E–4G, S3E–S3G) not stated
Module score (Tirosh et al. method) EMT gene signature scoring across individual scRNA-seq cells (Figures 4D, S3D) not stated
SCENIC (single-cell regulatory network inference and clustering) Transcription factor activity inference in Tub Ctrl vs Tub Fus scRNA-seq data (Figures 4H, 4I, S3H–S3M) not stated
Gene ontology / biological process enrichment analysis Malignant cell population vs normal nephron population in wk16 and wk17 scRNA-seq (Figures S4E, S4F; Table S5) not stated
Approaches that could also have been used
  • Xenotransplantation tumor take rates were reported as counts (0/4 vs 5/5) without a formal statistical test
    Could also: Fisher's exact test could also be applied to compare tumor formation proportions between the two groups — Fisher's exact test is the standard approach for two-group proportion comparisons with small sample sizes and yields an exact p-value, which would allow readers to assess the probability of observing such a complete separation by chance
  • Bulk RNA-seq differential expression was performed but the specific statistical method is not named in the provided text
    Could also: DESeq2 (negative binomial Wald test) or edgeR (likelihood ratio test) are widely used alternatives for count-based RNA-seq DEG analysis from small-n experiments — Both tools are designed for the overdispersion inherent in RNA-seq count data and provide built-in Benjamini-Hochberg FDR correction, making the DEG calls reproducible and the false discovery rate explicitly controlled
  • Cell type abundance changes between Tub Ctrl and Tub Fus were described qualitatively from bar graphs (Figure 4C) and t-SNE visualizations
    Could also: A compositional analysis method such as scCODA, propeller (limma-based), or DA-seq could also be used to statistically test changes in cell type proportions between conditions — Cell type proportions are compositional data; dedicated compositional or quasi-likelihood methods account for this constraint and provide uncertainty estimates for shifts in abundance, complementing the visual description
  • Transcription factor activity differences identified by SCENIC were presented as t-SNE expression maps without quantitative between-group tests
    Could also: A Mann-Whitney U test or permutation test comparing regulon activity scores between Tub Ctrl and Tub Fus cells could also be reported per transcription factor — Quantitative test statistics would allow readers to assess the magnitude and confidence of TF activity differences beyond visual inspection of embedding plots, and facilitate prioritization among many candidate regulators
  • RT-qPCR data were summarized as mean ± SD with n=3 independent experiments
    Could also: A 95% confidence interval could also accompany or replace SD as the measure of uncertainty around the mean — At n=3 the SD and CI carry similar information, but a CI explicitly conveys the precision of the estimate and scales interpretably with sample size, which can be more directly meaningful in small-n biological experiments
  • Motif enrichment in CUT&RUN peaks was analyzed by de novo discovery; the enrichment significance metric is not described in the provided text
    Could also: Reporting enrichment p-values and fold-enrichment over a matched background (as output by tools such as HOMER or MEME-AME) could also accompany the top-motif visualization — Explicit enrichment statistics and background model choices allow readers to calibrate confidence in the identified motifs and to compare results with other datasets using a common quantitative scale
Software: 10x Genomics Chromium (scRNA-seq platform) · SCENIC · Diffusion map / pseudotime (tool not named in provided text)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
5
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40463960

Paper: Ganpat et al. 2025, iScience. "An engineered tumor organoid model reveals cellular identity and signaling trajectories underlying SFPQ-TFE3 driven translocation RCC." PMID 40463960 · PMCID PMC12131257 · DOI 10.1016/j.isci.2025.112122

Artifact map (from the paper's Data & Code Availability — corrects the BRIEF)

The BRIEF paired code=github.com/teresouza/ganpat2023 with data=GEO:GSE260633. These are two different assays — the brief mis-paired them. The paper actually has THREE separate data/code units:

Assay Data accession Public? Code Notes
Bulk mRNA-seq ArrayExpress E-MTAB-13901 exists github.com/teresouza/ganpat2023 repo = bulk DESeq2 R script
CUT&RUN ArrayExpress E-MTAB-13900 exists github.com/nhungpham1707/CUTnRUN genome-wide binding
Single-cell RNA-seq GEO GSE260633 + Zenodo 10729944/10732016 processed MTX public (36 MB) Zenodo ganpat_et_al.zip (616 MB) cellranger 7.1 + Seurat 4.3
  • GSE260633 raw fastq: withheld ("Raw data are unavailable due to patient privacy concerns"). Processed CellRanger MTX sparse matrices ARE public (GSE260633_RAW.tar, 36 MB) + GSE260633_lx221-samples.txt.gz (multiplex barcode→sample map).

In scope (attempted) — single-cell QC, pipeline-derived, clearly specified

The single-cell unit is the cleanest reproducible target: public processed matrices + deposited code (Zenodo) + a deterministic, clearly-specified QC filter in Methods:

"Barcodes with fewer than 2000 or more than 100,000 transcripts were discarded, as well as barcodes with a mitochondrial percentage >50%." (cellranger 7.1.0, refdata-gex-GRCh38-2020-A; Seurat 4.3.0)

Claim to reproduce: number of cells passing this QC filter, computed independently with a third-party tool (Scanpy) on the paper's own public matrices, compared 1:1 against the authors' deposited processed object / scripts (Zenodo). Per rule P16, applying a standard third-party tool to the paper's data is equally valid.

Out of scope (not attempted) — with reasons

  • Bulk DESeq2 (ganpat2023 repo) — the repo script hard-codes the author's local laptop paths («path») for ALL inputs (metadata.csv, biomart.RDS, featureCounts *_features.txt, TitoRaw_annotated.txt); none shipped. The script is exploratory with many undefined variables (metaKidney, tfDf, norm3, gs). Raw bulk data on E-MTAB-13901 would need full STAR+featureCounts realignment. The DEG claim ("75 DEGs TubFus vs TubCtrl, 8 with CLEAR motif") is pinnable but the CLEAR-motif step is not in the repo. → docs_insufficient for a 1:1 of the repo; deferred (80/20).
  • CUT&RUN — separate repo/data, genome-wide binding; heavy, deferred.
  • Pseudotime / SCENIC / trajectory — the hard last 20%; stochastic (destiny, mclust, slingshot, SCENIC 20 runs). Not attempted.

Compute plan

«our HPC» SLURM (partition std), env via conda on «infra». Download GSE260633_RAW.tar + lx221 map + Zenodo zip to «infra»; inventory + read authors' scripts for ground-truth cell counts; Scanpy QC reproduction; compare.

sc-qc-thresholds
Reported
discard barcodes with <2000 or >100,000 transcripts; mito% >50
Reproduced
config-e16.R: min_txpts=2000, max_txpts=100000, max_pct_mito=50 (verbatim). NEW: config-e17.R uses min_txpts=3000 for wk17 (per-timepoint difference the paper text omits)
exact
sc-deposit-lx221
Reported
e16 (wk16) library = 1000 droplets = 613 e16ctrl + 387 e16fus
Reproduced
matrix has exactly 1000 barcodes; public demux map = 613 e16ctrl + 387 e16fus
exact
sc-qc-e16
Reported
863 e16 (wk16) organoid cells pass QC (the analyzed set)
Reproduced
863 cells pass QC (Scanpy on public GSE260633 lx221 matrix; mito filter a no-op, 0 MT- genes)
exact
sc-qc-e17
Reported
no single pinned paper value; deposited e17 pipeline reads non-public intermediates
Reproduced
lx419=600 / lx420=467 barcodes pass the deposited e17 threshold (min_txpts=3000) on the public matrices — observation only, not a 1:1 to the paper's analyzed e17 set
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 88/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

The single-cell QC unit reproduces 1:1 EXACTLY — 863 e16 cells, the 2000–100000/mito-50% thresholds, and the 1000-droplet 613/387 ctrl/fus split all match the authors' deposited pipeline, and 863 is fully derivable from the public GSE260633 matrix (no fabrication). However, this is only one of the paper's three data units; the central biological conclusions (DEGs/CLEAR motifs, SCENIC/pseudotime trajectories) were not reproduced, so the core claim is only narrowly supported (q7 yellow). Two non-fatal issues sit on the authors'/brief side: the brief mis-paired the bulk repo with the single-cell accession, and the 'preliminary' Zenodo deposit is internally inconsistent (Fig4c TSV = 1054 typed cells vs the 863 figure). Net: a solid, exact reproduction of what was attempted, but limited scope and deposit caveats keep the overall judgement at yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

214.5 k
tokens (I/O) · 14.2 M incl. cache
49 min
runtime · 0.03 CPU-h
1.9 GB
peak RAM
2
HPC jobs
hummel
machine