Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comprehensive analysis of metastatic gastric cancer tumour cells using single-cell RNA-seq.

Sci Rep · 2021
L1 47/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
47/100
Reproducibility score
1.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 6% of all assessed papers rank 1100 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Smart-seq2 gastric-cancer scRNA-seq paper; only code link is the generic scater package (P16 third-party tool, no authors' analysis repo). Reproduced by running scater(log)+Seurat 5.3.0 on the single deposited count matrix (GSE158631_count.csv.gz, 21,196 genes x 94 cells) on «our HPC». RESULT = PARTIAL. The dataset DESIGN reproduces 1:1: N=94 cells exactly, and the per-patient TT/LN split (19/4, 27/13, 19/12) matches the paper exactly from column names. The pipeline-derived gene/cluster numbers do NOT reproduce from the public deposit: (a) the reported 22,335-gene hg19 reference vs 21,196 genes actually in the matrix; (b) the headline '7601 genes passed avg-read>1 filtration' is unreproducible -- the literal criterion on the deposited 94-cell matrix yields 4174, and no natural threshold reproduces 7601 (the pre-QC 171-sample / full-annotation matrix on which 7601 was presumably computed was never deposited); (c) '12 significant PCs' has no stated significance criterion and the PCA scree shows no clean elbow at 12; (d) '4 main clusters in the tumour tissues' does not reproduce -- a standard Seurat pipeline gives 2-3 clusters (all cells) or 1-2 (TT-only) across resolutions 0.2-1.0, and the paper does not report the resolution used. Causes are under-specification (no resolution, no PC criterion) plus undocumented pre-deposit gene filtering; this is flagged for human review as 'not verifiable from the deposit' rather than confirmed fabrication. NOT attempted (out of scope): read mapping/featureCounts (no FASTQ/BAM deposited), TPM with true gene lengths, monocle2/TSCAN trajectory, Metascape GO, marker/TF gene lists, immunofluorescence (wet-lab).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 47
    assessed: 2026-06-18 ⛓ 15051193333f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What are the intratumoural single-cell transcriptomic differences driving gastric cancer lymph node metastasis, which bulk approaches mask? The study tests whether scRNA-seq of paired primary and metastatic lymph node gastric cancer tissue can reveal subpopulations and marker/driver genes of metastasis.

Core claims
  • CDK12, ERBB2, and CLDN11 are overexpressed in metastatic lymph node gastric cancer and serve as candidate lymph node metastasis marker genes. finding
  • NOTCH2, NOTCH2NL, KIF5B, and ERBB4 are highly expressed in primary gastric cancer. finding
  • A subgroup of cells bridges the metastatic and primary groups, implying a transition/transition state during the metastatic process. finding
  • Transcription factors FOS and JUN (and FOSB, JUNB, ZNF256) drive the regulatory networks and act as potential gastric cancer evolution-driving genes. mechanism
  • Single-cell RNA-seq of paired primary and metastatic gastric tumours reveals significant intratumoural heterogeneity and patient-specific cancer profiles while microenvironmental subsets are shared across patients. finding
  • Pseudotime trajectory analysis revealed a postulated evolution state from Cluster 0 > 2 > 1 among GC clusters. finding
  • Smart-seq2 scRNA-seq of manually picked single cells from paired primary and lymph node gastric tumours provides a single-cell resolution metastasis dataset (GSE158631). resource
  • Seurat marker analysis identified four main cell clusters in the overall single cells. finding
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-seq (Smart-seq2) primary tumour tissue (TT) and paired lymph node (LN) metastasis tissue from 3 gastric cancer patients none (primary vs metastatic comparison) single-cell whole-transcriptome gene expression (TPM) Smart-seq2 protocol; Illumina HiSeq 2500, 50 bp single-end
immunofluorescence GC tumour tissues (TT and LN), paraffin embedded none protein expression of ERBB4/ERBB2 and CLDN11 in TT vs LN Abcam Anti-ERBB2 and Anti-Oligodendrocyte Specific Protein (CLDN11) antibodies
single-cell trajectory / pseudotime analysis gastric cancer cell clusters (Seurat-identified) none evolutionary trajectory and driver genes TSCAN, diffusion map, monocle2, SLICER
single-cell clustering / dimensionality reduction 94 QC-passed single cells from 3 patients none t-SNE/PCA cluster separation of primary vs metastatic cells Seurat, scater (R)
immunohistochemistry / FISH (clinical characterization) 3 gastric cancer patient tumours none HER2 status, Ki67, P53, tumour markers
Key results
  • CDK12, ERBB2, and CLDN11 overexpressed in metastatic (LN) gastric cancer cells
  • NOTCH2, NOTCH2NL, KIF5B, and ERBB4 highly expressed in primary cancer cells
  • Correlations between individual tumour cells from different samples spanned a broad range of Pearson coefficients, implying prominent transcriptomic heterogeneity r = -0.1 ~ 0.98
  • 94 of 171 samples passed quality control; 7601 genes adopted for analysis 94/171
  • Significant tumour and stromal scoring differences found between primary and metastatic single cells
  • Pseudotime trajectory revealed postulated evolution state from Cluster 0 > 2 > 1
  • FOS, FOSB, JUN, JUNB, and ZNF256 identified as transcription factors driving regulatory networks in evolution
  • Each cell sequenced with 20,000~200,000 uniquely mapped reads, sufficient to detect subpopulation profiles 20,000~200,000 reads
Key statistics
  • correlation r = -0.1 ~ 0.98 (Pearson correlation range between individual tumour cells across samples)
  • count 94 out of 171 samples passed QC (samples passing per-gene average read >1 filtration)
  • count 7601 genes (genes passing filtration adopted in further analysis)
  • count 20,000 ~ 200,000 uniquely mapped reads per cell (sequencing depth per cell)
  • count 22,335 genes (total genes in hg19 reference used for mapping)
  • fold_change twofold cut-off, FDR-adjusted p < 0.05 (DEG criteria)
  • count TT/LN cells: PT1 19/4, PT2 27/13, PT3 19/12 (analysed cell numbers per patient after QC)
  • other 22 PCR amplification cycles (full-length cDNA amplification in Smart-seq2 library prep)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study applied Smart-seq2 scRNA-seq to 94 quality-controlled single cells from primary tumour and paired lymph node metastasis tissue across three gastric cancer patients. The primary analytical framework combined dimensionality reduction (PCA, t-SNE), unsupervised graph-based clustering (Seurat), and pseudotime trajectory analysis (TSCAN, Monocle2, diffusion map, SLICER) to characterise intratumoural heterogeneity and postulated metastatic evolution. Differentially expressed genes between tumour tissue (TT) and lymph node (LN) cells were identified using a ≥2-fold change threshold combined with FDR-adjusted p < 0.05 via R's stats package, and Student's t-test was applied to bulk stemness/immune/stromal/tumour scoring comparisons; results were reported principally as ranked gene lists and visualisation plots with no dispersion measures for group-level summaries.

Replicationbiological Sample sizeThree patients; 94 cells passed QC from 171 total; per-patient cell counts provided in Table 1; no formal power analysis described GroupsPrimary tumour tissue (TT) vs. paired lymph node metastasis (LN); additionally four Seurat-defined cell clusters Pairingpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFDR adjustment (specific algorithm not stated; Benjamini-Hochberg is the R stats default but is not named)
Statistical tests used
Test Applied to n Assumptions
Student's t-test (R stats::t.test) Bulk stemness, immune, stromal, and tumour scoring comparisons between TT and LN single cells 94 cells total (PT1: 19 TT/4 LN; PT2: 27 TT/13 LN; PT3: 19 TT/12 LN) not stated
Fold-change (≥2-fold) combined with FDR-adjusted p-value (threshold p < 0.05) Identification of differentially expressed genes between TT and LN single cells across 7,601 genes 94 cells across 3 patients not stated
Pearson correlation (R stats::cor) Pairwise inter-cell transcriptomic heterogeneity assessment; r range reported as −0.1 to 0.98 94 cells (pairwise) not stated
Unsupervised Seurat graph-based clustering Identification of four main cell clusters from all single cells using 12 principal components 94 cells na
Pseudotime trajectory analysis (TSCAN, Monocle2, diffusion map, SLICER) Postulated evolutionary trajectory of GC cell clusters (Cluster 0→2→1); TT vs. LN evolutionary trajectory Cells from Seurat cluster output; exact per-method n not re-stated na
Approaches that could also have been used
  • Student's t-test was applied to compare bulk scores between TT and LN with the single cell as the unit of analysis across n=3 patients
    Could also: A linear mixed-effects model with patient as a random effect could also have been used — Because cells are nested within patients, they are not fully independent observations; a mixed model would explicitly partition within-patient from between-patient variance, which is a standard consideration when biological replication consists of a small number of donors contributing multiple observations
  • Differentially expressed genes were identified using fold-change plus FDR p-value via R's base stats package
    Could also: Dedicated scRNA-seq DEG methods such as MAST, or a pseudobulk approach (DESeq2 or edgeR on per-patient aggregates) could also have been applied — MAST models the bimodal, zero-inflated distribution typical of scRNA-seq data; pseudobulk DESeq2/edgeR treats the patient as the unit of replication, which directly accounts for within-patient cell-level correlation — both are widely used alternatives in the scRNA-seq literature
  • The experimental design was paired (each patient contributed both TT and LN), but statistical comparisons were not described as paired
    Could also: Paired t-test, Wilcoxon signed-rank test, or a paired pseudobulk analysis could also have been used — Paired tests use within-patient TT–LN differences as the unit of analysis, removing inter-patient variability; this can increase statistical sensitivity and more directly addresses the matched design — a natural fit for n=3 matched-pair data
  • t-SNE was used as the sole single-cell visualisation method
    Could also: UMAP (Uniform Manifold Approximation and Projection) could also have been used alongside or instead of t-SNE — UMAP is a widely adopted complement to t-SNE in scRNA-seq analysis that tends to better preserve global inter-cluster distances while maintaining local structure; reporting both is common practice in contemporary scRNA-seq workflows
  • Group-level comparisons were reported without dispersion measures (no SD, SEM, or CI for scoring differences)
    Could also: Reporting 95% confidence intervals or SD alongside p-values and effect sizes could also have been included — Dispersion measures convey the spread and precision of group estimates, allowing readers to judge the practical magnitude of differences independently of sample size; their inclusion complements p-values and fold-change thresholds, and is recommended by MIQE and MIAME-style reporting guidelines
  • Four separate pseudotime trajectory algorithms (TSCAN, Monocle2, SLICER, diffusion map) were run and their outputs were presented separately
    Could also: A formal quantitative agreement measure across trajectory methods could also have been reported — Different trajectory algorithms can produce divergent cell orderings; computing pairwise agreement (e.g., Kendall's τ or Spearman ρ on inferred pseudotime ranks) is one approach to reporting trajectory robustness and is particularly informative when the biological ground truth is unknown
Software: R / stats package (t.test, cor, prcomp, cluster) · R / scater (newSCESet; TPM normalisation and log2 transformation) · R / ComplexHeatmap · R / ggplot2 · Seurat (unsupervised graph-based clustering and marker analysis) · TSCAN / Monocle2 / SLICER / diffusion map (pseudotime trajectory) · Metascape (Gene Ontology enrichment analysis) · HiSat2 / FeatureCounts / Trimmomatic (read alignment and preprocessing)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33441952

Paper: Wang et al. 2021, Comprehensive analysis of metastatic gastric cancer tumour cells using single-cell RNA-seq, Sci Rep. PMID 33441952 / PMC7806779.

Design: Smart-seq2 plate-based scRNA-seq, 3 gastric-cancer patients, paired primary tumour tissue (TT) + metastatic lymph node (LN). HiSat2 -> featureCounts -> scater (TPM/log2) -> PCA/t-SNE -> Seurat clustering -> TSCAN/monocle2 trajectory -> Metascape GO.

Code link in brief: github.com/davismcc/scater — this is the generic scater R package (a third-party tool), NOT the authors' analysis repo. Per P16 this is a valid third-party-tool-on-paper-data reproduction. There is no authors' own analysis repository.

Data: GEO GSE158631 — a single deposited file, GSE158631_count.csv.gz (per-cell raw count matrix, 21,196 genes x 94 cells).

In scope (pipeline-derived, reproducible from deposited data)

  • C1 N cells passing QC = 94 — directly from matrix columns. (171 sequenced is not in the deposit -> uncheckable.)
  • C2 Per-patient/tissue cell breakdown (19/4, 27/13, 19/12) — from column names.
  • C3 Reference gene count 22,335 vs deposited matrix gene count.
  • C4 Gene filter "per-gene average read > 1 across all samples" -> 7601 genes.
  • C5 12 significant principal components (scater log2-TPM + PCA).
  • C6 4 main clusters (Seurat).

Out of scope (wet-lab / external / not pipeline, not attempted)

  • Sequencing, library prep, read mapping (HiSat2/featureCounts) — raw FASTQ/BAM not deposited.
  • Immunofluorescence validation (wet-lab).
  • Metascape GO enrichment of top-1000 genes (external web platform, manual).
  • Specific marker/TF gene lists (NOTCH2, ERBB2, FOS/JUN, etc.) — descriptive, derived from the clustering + DE that is itself under-specified.

Reproducibility notes

  • C1, C2 are deterministic and cleanly reproducible.
  • C3, C4 fail against the deposit (the matrix dimensions match neither the 22,335 reference nor the 7601 filtered set; literal avg>1 filter on the 94-cell deposit = 4,174).
  • C5, C6 are under-specified in the paper (no PC-significance method, no Seurat resolution), so they are attempted but graded provisional — a standard pipeline result, not an exact match.
C1
Reported
94 of 171 samples passed QC
Reproduced
94 cells (matrix columns)
exact
C2
Reported
per-patient TT/LN: GC1 19/4, GC2 27/13, GC3 19/12
Reproduced
19/4, 27/13, 19/12 (column names)
exact
C3
Reported
22,335 reference genes (UCSC hg19)
Reproduced
21,196 genes in deposited matrix
did not match
C4
Reported
7601 genes passed per-gene-average-read>1 filter
Reproduced
4174 genes (literal criterion on 94-cell deposit)
did not match
C5
Reported
12 significant principal components
Reproduced
PCA stdev decays smoothly, no clean elbow at 12
partial
C6
Reported
4 main clusters in tumour tissues
Reproduced
2-3 clusters (94 cells) / 1-2 (65 TT cells) at res 0.2-1.0
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 47/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The dataset design reproduces exactly — 94 cells and the per-patient TT/LN split (19/4, 27/13, 19/12) match the paper 1:1 from the deposited column names. However, every pipeline-derived number fails to reproduce from the public deposit: the deposited matrix has 21,196 genes (not the 22,335 reference), the stated avg-read>1 filter yields 4174 not 7601, and a standard scater(log)+Seurat run gives 2-3 clusters (not 4) with no clean elbow at 12 PCs. The cause sits mainly on the authors'/data-availability side — undocumented pre-deposit gene filtering plus underspecified methods (no resolution, no PC criterion) and only a generic third-party code link — rather than confirmed fabrication. Deviations are substantive (≈1.8x on the gene filter, clustering not recovered) but explainable as not-verifiable-from-deposit, so overall partial/yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

107.8 k
tokens (I/O) · 9.3 M incl. cache
27 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.