Quality control method for RNA-seq using single nucleotide polymorphism allele frequency.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Partial reproduction. The upstream alignment pipeline (sratoolkit -> Bowtie2/TopHat2 against GRCm38, reused from a prior run: 84.9% overall mapping rate) was sound. The downstream snpexp SNP-allele-frequency counting step, which is the paper's actual novel contribution, was NOT reproducible out-of-the-box: the GitHub master branch crashes with glibc heap corruption near exit, and the last tagged release (0.4.2, which we had to patch just to compile under a modern C++ standard) silently produces ZERO output whenever its -G (exon-restriction) option is used against any modern Ensembl/GENCODE GTF, because its GTF parser hard-requires an obsolete 'tss_id' attribute that current annotation files never contain -- confirmed with 0 occurrences via grep and via the tool's own verbose diagnostic ('0 alleles loaded' with -G vs '78,772,544 alleles loaded' with -I). After a minimal, clearly-flagged source patch removing that requirement, the exon-restricted pipeline runs but crashes with a SECOND, independent heap-corruption bug partway through chromosome X -- yielding a genuine, disclosed PARTIAL result covering only autosomes 1-19. That partial result (26,411 discriminating B6/129 SNPs at >=20x coverage) is reasonably close to the paper's reported ~23,838 SNPs for the ESC sample (~11% higher, and would be higher still with X/Y/MT included), and a fully genome-wide unrestricted cross-check (30,368 SNPs) is in the same order of magnitude. The allele-frequency distribution shows real concentration in the 0.3-0.5 range expected for a 129B6F1 hybrid sample, but with heavier tails at both extremes than an idealized single peak at 50% -- plausibly reference-mapping bias, un-filtered low-confidence VCF calls, or genuine QC signal, not disambiguated here. NOT attempted: exact replication of the paper's full FILTER-column/quality criteria beyond minimum coverage; the paper's other RNA-seq samples beyond the RU-designated SRR1047502; digit-level comparison to the paper's Figure 1B/2A histograms (not available as machine-readable data, only as a plotted figure). Two genuine, well-evidenced third-party-tool defects (not environment or self-imposed-limit issues) are the main reason this is graded 'partial' rather than 'reproduced': (1) master-branch heap corruption on exit, and (2) release-0.4.2's silently-broken GTF exon filter due to the tss_id dependency, compounded by a second heap-corruption bug once that filter is patched to actually run.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-08-06
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-06no human curator yet
- Last updated
- 2026-08-06
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan allele frequencies of heterozygous SNPs observed in RNA-seq reads be used as a quality-control measure to detect contaminating cells with a different genomic background and chromosomal aberrations in transcriptome datasets, both prospectively and retrospectively?
- ★ SNP allele frequency distributions from RNA-seq reads can detect contaminating cells whose genomic background differs from the target cells; the mode of the distribution reflects the cellular composition while its variance reflects PCR bias. method
- ★ Chromosome-wise SNP allele frequency analysis detects aneuploidy and can determine the parental origin of duplicated chromosomes, without requiring control cells of normal karyotype. method
- ★ FI-SCs in the STAP dataset, annotated as 129B6F1, were not 129B6F1: their RNA-seq data derive from ~90% B6 cells with an ESC-like expression pattern and ~10% CD1-like cells with a TSC-like pattern, i.e. TSC contamination. finding
- ★ STAP cells in the original dataset carried trisomy of chromosome 8, with the duplicated chromosome originating from the 129 parent. finding
- ★ The reported contribution of FI-SCs to the placenta may have been an artifact of contamination by trophoblast stem cells, which are known organ stem cells. mechanism
- Because pure trisomy 8 is embryonic lethal in mice, the samples annotated as neonatal spleen cells must instead have been cultured cells with ESC-like expression characteristics. finding
- MEF feeder-cell contamination of FI-SCs is negligible because MEF marker cytokine/extracellular matrix genes were not expressed in FI-SCs. finding
- Sensitivity of the method depends on sequencing depth; distributions become too noisy to interpret when fewer than ~1000 SNPs are available. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| In silico simulation of SNP allele frequency (modified binomial distribution with PCR-bias term) | mathematical model, N = 50 fragments per locus | simulated PCR bias, Gaussian sd = 0 (no bias) or 1 (high bias); varied allele composition | probability distribution of reference-allele frequency | — |
| RNA-seq SNP allele frequency profiling of public datasets | mouse ESCs (SRR1047502, 129B6F1), fibroblast-derived iPSCs (SRR1047504, 129B6F1), MEFs (SRR104220, 129B6F1), normal fibroblasts (SRR1191170, B6 x BALB/c), cancer-associated fibroblasts (SRR1191171, B6 x BALB/c), HSCs (SRR892995, B6) | none | distribution of reference (B6) vs alternative allele frequency; number of applicable exonic SNPs per sample | dbSNP build 137 VCF + Sanger Mouse Genomes Project variants; iGenomes exon annotation; bowtie/bowtie2 build, tophat/tophat2 alignment (50-bp fragmented reads, 2 mismatches, --no-coverage-search -G genes.gtf); sratoolkit 2.3.4-2; Bowtie2 v2.1.0 |
| In silico contamination titration by random sampling/mixing of RNA-seq datasets | pure C57BL/6 hematopoietic stem cells mixed with 129B6F1 embryonic stem cells | artificial contamination at varying ESC percentages | shape and peak position of allele frequency curves versus mixing ratio | — |
| RNA-seq allele frequency re-analysis of the STAP dataset | mouse ESCs, STAP cells, STAP stem cells (STAP-SCs), FI-SCs, TSCs; seven replicate experiments | none (re-analysis); original perturbations included Fgf4-induced conversion of STAP cells to FI-SCs | per-sample allele frequency distributions; genotype (B6 vs non-B6) of SNPs in ESC-specific, TSC-specific and other genes | TruSeq reagent; SRA project SRP038104; mouse genome mm10 (GenBank) |
| Gene-wise / locus-wise SNP genotype analysis | ESCs, TSCs and FI-SCs from the STAP dataset | none | homozygous vs heterozygous SNP counts at ESC markers (Sall4, Klf4), TSC markers (Elf5, Sox21) and at Des, Grb2, Setd7, Fbxo21, Chd4; Fisher's exact test of genotype distribution between TSC- and ESC-specific genes | — |
| Chromosome-wise SNP allele frequency (virtual karyotyping) analysis | STAP cells and ESCs, derived from 129 and B6 | none | per-chromosome allele frequency peak position (whole chromosomes vs chromosome 8) | SMARTer reagent kit (distinct from the TruSeq data in Figs 1-2) |
| Chromosome-wise differential gene expression analysis | STAP cells vs ESCs | none | FPKM-based expression of genes grouped by chromosome; two-sided Student t-tests | tophat/tophat2 FPKM quantification |
| Marker gene expression profiling (heatmap / marker quantification) | ESCs, TSCs, FI-SCs, MEFs | none | normalized log ratios of FPKM against median of all samples for MEF cytokine/extracellular matrix genes; TSC marker gene expression relative to average TSC level | — |
- – Simulation showed the mode of the allele frequency distribution corresponds to the SNP allele composition while the variance depends on PCR bias; public RNA-seq datasets agreed with the simulation, with heterozygous SNPs averaging ~50% independent of cell type. ~50% mean allele frequency
- – FI-SCs annotated as 129B6F1 failed to show the balanced allele distribution, with the majority of SNPs B6-like, indicating a nearly pure B6 origin; ESC marker genes (Sall4, Klf4) carried only B6 SNPs in FI-SCs while ESCs carried both 129 and B6 alleles, and this B6 dominance was absent at TSC markers (Elf5, Sox21).
- – FI-SC allele frequency peaks were at ~95-96% while the contaminating population peak is expected at ~10%, consistent with a two-population sample of ~90% B6 ESC-like cells and ~10% CD1-like TSC-like cells. peak ~95-96%; ~10% contaminating population
- – FI-SC-specific SNPs, at positions where B6 and 129 share the same nucleotide, mostly matched the CD1 background of the TSCs used in the experiments and appeared heterozygous, confirming shared TSC-specific alleles.
- – Chromosome 8 of STAP cells showed an allele frequency peak at ~33% instead of ~50%, indicating three copies of chromosome 8 with the duplicate derived from the 129 parent. peak ~33% vs expected ~50%
- – Chromosome 8 gene expression was significantly higher in STAP cells than ESCs, concordant with trisomy 8, whereas chromosome 13 genes were significantly less expressed. 1.3-fold higher for chromosome 8
- ▼ MEF marker cytokine and extracellular matrix genes were not expressed in FI-SCs, excluding feeder-cell contamination as the source of the skewed allele frequencies.
- – Detection sensitivity degrades with low SNP counts: MEF data with 5948 SNPs produced markedly more spikes than ESC data with 23 838 SNPs, and distributions become very noisy below ~1000 SNPs. 5948 vs 23 838 SNPs; noise threshold ~1000 SNPs
- pvalue 2.89 × 10−25 (Two-sided Student t-test, chromosome 8 gene expression higher in STAP cells than ESCs)
- fold_change 1.3 times higher (Chromosome 8 gene expression in STAP cells vs ESCs)
- count 6859 and 7243 heterozygous SNPs; 24 and 14 non-B6 homozygous alleles (Duplicated FI-SC RNA-seq experiments, at alleles where B6 and 129 differ)
- other approximately 95-96% (Peak allele frequency in FI-SCs, versus expected population peak at ~10%)
- other approximately 33% (Peak allele frequency of chromosome 8 in STAP cells, versus ~50% expected for diploid)
- count 31 of 97 (Mouse ESC lines reported to carry trisomy 8 (Mayshar et al. 2010))
- count 1 016 227 SNPs (Exonic SNPs retained from dbSNP build 137 VCF after excluding non-exonic variants)
- count 5948 SNPs (MEF) vs 23 838 SNPs (ESC); ~1.3 × 10^9 nucleotides required for 20× average coverage (Sensitivity/depth requirements of the SNP allele frequency method)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper is a computational reanalysis of previously published RNA-seq datasets, using SNP allele-frequency patterns (modeled with a binomial-distribution-based simulation incorporating PCR bias) to assess sample composition, detect contaminating cell populations, and identify chromosomal aberrations. Genotype-distribution comparisons between marker-gene sets were assessed with Fisher's exact test, and chromosome-wise gene-expression differences between cell populations were assessed with two-sided Student's t-tests. Results were conveyed mainly through allele-frequency distribution plots, a stated fold-change, and p-values (including one exact value), without a described replicate-structure classification, formal power analysis, or multiplicity correction across the many gene/chromosome comparisons performed.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Fisher's exact test | genotype distribution comparison between TSC-specific genes and ESC-specific genes (Fig. 2D) | — | not stated |
| two-sided Student's t-test | chromosome-wise gene expression comparison (chromosome 8 and chromosome 13) between STAP cells and ESCs (Fig. 3B) | — | not stated |
-
Genotype-distribution comparisons between marker-gene sets were evaluated with Fisher's exact test.↳ Could also: a chi-square test of independence — For larger SNP/read counts, a chi-square test can also be used to compare categorical distributions and gives similar results to Fisher's exact test asymptotically, while Fisher's exact test is often preferred specifically when counts are small.
-
Chromosome-wise expression differences between STAP cells and ESCs were assessed with a two-sided Student's t-test.↳ Could also: a non-parametric test such as the Mann-Whitney U test, or a count-based RNA-seq framework such as DESeq2/edgeR (negative binomial) or limma-voom — RNA-seq read counts are often over-dispersed and not normally distributed; non-parametric methods avoid the normality assumption, and negative-binomial or moderated frameworks are widely used alternatives specifically designed to model RNA-seq count variance.
-
Multiple genes, marker-gene sets, and chromosomes were compared without a stated multiple-testing correction.↳ Could also: a Benjamini-Hochberg false discovery rate (FDR) correction or Bonferroni correction — Applying a correction across the family of comparisons performed would control the overall false-positive rate when many statistical tests are run in parallel.
-
Point estimates (e.g., allele frequencies, FPKM ratios, fold-changes) were reported without an accompanying dispersion or interval measure in the main text.↳ Could also: reporting standard deviation, standard error, or a 95% confidence interval alongside each estimate — Adding a dispersion or interval measure would let readers directly gauge the precision and variability of the reported estimates.
-
Sensitivity to sequencing depth and SNP count was discussed descriptively (e.g., noting that fewer than ~1000 SNPs produced noisy distributions) rather than through a formal statistical framework.↳ Could also: a formal power or sensitivity analysis, or explicit confidence intervals on estimated contamination percentages — A formal analysis could quantify, for a given SNP count and read depth, the minimum detectable contamination fraction or aneuploidy skew with a stated confidence level.
-
The replicate structure of the reanalyzed RNA-seq experiments (e.g., the seven replicate experiments, duplicated FI-SC experiments) was not explicitly classified as biological or technical, nor modeled as such.↳ Could also: a mixed-effects or hierarchical model that explicitly accounts for replicate/batch structure — Explicitly modeling replicate structure can separate biological variability between samples from technical variability between sequencing runs when combining results across replicate experiments.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The correct raw inputs were fully available and used 1:1 (SRR1047502, 35.9M read pairs; Sanger MGP v5 129S1_SvImJ), and the exonic discriminating-SNP count landed at 26,411 vs the paper's ~23,838 — same order of magnitude, +10.8%, and that from an autosome-only run truncated by a heap-corruption abort, so the complete figure and hence the true gap are larger. The decisive defect is on the authors' side: snpexp 0.4.2 gates every GTF record on a legacy Cufflinks tss_id attribute that occurs 0 times in modern Ensembl GTFs, so the paper's central exon-restriction step silently loads '0 alleles loaded' and writes a header-only file with exit code 0 — the headline number was recoverable only after patching the source, and the master-branch binary separately aborts before flushing -o. The residual numeric offset is most plausibly ours: no PASS-only VCF filter, and self-chosen -m 20 / GTF build, none of which the paper specifies. The qualitative QC claim survives in weakened form — real mass at 0.3–0.5 (6,814 SNPs) as predicted, but bimodal tails (7,817 at 0.0–0.1, 3,699 at 0.9–1.0) instead of a clean unimodal ~50% peak, and Fig 1B/2A exists only as a figure so no numeric check was possible.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.