Comparison of RNA-Seq by poly (A) capture, ribosomal RNA depletion, and DNA microarray for expression profiling.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Downstream comparative statistics (protocol-pair correlations, detected-gene counts, differential expression between RNA-seq library-prep protocols) were reproduced directly from the paper's own processed/normalized output (GSE51783_normalized_gene_counts.txt.gz), since raw sequencing reads needed for the upstream MapSplice->UBU->RSEM pipeline are dbGaP controlled-access and unavailable to this account. Correlation results reproduce closely (exact/within-tol: e.g. RiboZero-vs-DSN FF r=0.977 reproduced vs 0.961 reported; all RNA-seq protocol pairs confirmed >0.9 in FF as stated). Differential-expression gene counts (SAM) are directionally and order-of-magnitude consistent but not numerically exact (174-293 vs 410 reported; 27-37 vs 104 reported), because the exact Stanford samr R package was not available in this environment and was replaced with a documented, simplified permutation-based SAM-equivalent implemented in Python/numpy. Sample design and GEO sample counts (48 RNA-seq + 11 microarray = 59 total) match the paper exactly. NOT attempted: the raw-read alignment/quantification pipeline itself (data access blocker), RNA-seq-vs-microarray correlation (probe annotation only available as an impractically large 2.9GB combined file), and the TCGA external validation-cohort clustering claims (separate dataset, out of scope this pass).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-08-01
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-01no human curator yet
- Last updated
- 2026-08-01
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan RNA-Seq library preparation protocols that do not depend on intact poly(A) tails (Ribo-Zero rRNA depletion and DSN normalization) provide gene expression profiling from degraded FFPE-derived RNA that is concordant with fresh-frozen mRNA-Seq and DNA microarray data? The study evaluates whether these protocols are suitable alternatives to poly(A)-capture mRNA-Seq for clinically archived material.
- ★ Ribo-Zero-Seq removes rRNA with efficiency comparable to poly(A)-based mRNA-Seq in both FF and FFPE RNA, whereas DSN-Seq leaves significantly more rRNA and shows greater variation. finding
- ★ Transcript quantification is highly concordant across all protocols and between paired FF and FFPE RNAs, supporting the use of RNA-Seq on FFPE-derived RNA for gene expression profiling. finding
- ★ rRNA-depletion protocols (Ribo-Zero-Seq, DSN-Seq) map far fewer bases to the transcriptome than mRNA-Seq, with most reads mapping to intronic/intergenic regions, suggesting capture of pre-mRNA. finding
- ★ Ribo-Zero-Seq on FF RNA provides equivalent or less biased 5′-to-3′ coverage than mRNA-Seq because it does not rely on poly(A) selection. finding
- ★ Substantially greater sequencing depth is required for rRNA-depletion protocols (45–65 M reads) than for mRNA-Seq (~14 M reads) to detect the same number of genes as an Agilent DNA microarray. finding
- ★ Biologically relevant expression signatures (breast tumor intrinsic subtypes) are preserved across protocols, with samples from the same tumor co-clustering with their protocol partners. finding
- Genes differentially detected between mRNA-Seq and Ribo-Zero-Seq are enriched for non-polyadenylated species (snoRNAs, histone RNAs), which are also depleted by DSN's Cot-kinetics normalization. mechanism
- A benchmarking framework and paired FF/FFPE multi-protocol dataset (11 UNC breast tumors plus 10 TCGA tumors) for comparing expression profiling platforms. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| DNA microarray (gene expression) | Fresh-frozen human breast tumor RNA (UNC dataset, 11 tumors) | none | Expressed Entrez gene detection and expression levels; baseline gene count for depth comparison | Custom Agilent 244,000-feature microarray |
| mRNA-Seq (poly(A) capture RNA-Seq) | Fresh-frozen human breast tumor RNA (UNC, n = 11; TCGA, n = 10) | none | % rRNA, % aligned bases, exon/intron/intergenic base distribution, CV of coverage, 5′-to-3′ bias, transcript abundance | oligo(dT) primer poly(A) enrichment; alignment to UCSC hg19 |
| Ribo-Zero-Seq (rRNA depletion RNA-Seq) | Fresh-frozen human breast tumor RNA (UNC, n = 11; TCGA, n = 6) | none | rRNA depletion efficiency, alignment profile, coverage uniformity, 5′/3′ bias, transcript quantification | Ribo-Zero kit (hybridization capture of rRNA with magnetic bead subtraction) |
| DSN-Seq (Duplex-Specific Nuclease normalization RNA-Seq) | Fresh-frozen human breast tumor RNA (UNC, n = 10) | none | rRNA level, alignment profile, coverage uniformity, 5′/3′ bias, transcript quantification | Duplex-Specific Nuclease, Cot-kinetics-based normalization |
| Ribo-Zero-Seq on FFPE RNA | FFPE human breast/tumor RNA (UNC, n = 8; TCGA, n = 18, including 8 technical replicates) | none (modified protocol omitting selected steps) | rRNA level, % aligned bases, CV coverage, 5′/3′ bias, correlation to microarray and to FF libraries, technical reproducibility | Ribo-Zero kit; alignment to UCSC hg19 |
| DSN-Seq on FFPE RNA | FFPE human tumor RNA (UNC, n = 4; TCGA, n = 10) | none (modified protocol omitting selected steps) | rRNA level, % aligned bases, CV coverage, 5′/3′ bias, correlation to microarray, required read depth | Duplex-Specific Nuclease normalization |
| Hierarchical clustering with intrinsic gene list | 904 human breast tumor samples (88 UNC/TCGA samples profiled here plus 725 TCGA breast tumors and 91 normal breast tissues) | none | Co-clustering of protocol-partner libraries from the same tumor; subtype assignment | — |
| Significance Analysis of Microarrays (SAM) differential expression | UNC breast tumors: 3 Basal-like and 3 Luminal samples profiled by mRNA-Seq, Ribo-Zero-Seq and DSN-Seq; also protocol-vs-protocol comparisons | none (subtype and protocol comparison) | Number of differentially expressed genes at FDR 0; overlap of top 500 most variably expressed genes | SAM |
- – In FF RNA, Ribo-Zero-Seq rRNA content was near mRNA-Seq levels while DSN-Seq rRNA was far higher (UNC: 5.04% vs 116% relative to mRNA-Seq); FFPE showed 7.14% for Ribo-Zero vs 585% for DSN. 5.04% vs 116% (FF); 7.14% vs 585% (FFPE), relative to mRNA-Seq; p < 0.001
- ▼ Bases mapping to transcripts (coding + UTR) in FF were 62.3% for mRNA-Seq versus 31.5% for Ribo-Zero-Seq and 22.7% for DSN-Seq, while intronic/intergenic bases rose from 31.6% to 62.5%. 62.3% → 31.5% / 22.7% transcriptome; 31.6% → 62.5% intron+intergenic
- ▼ In FFPE, both DSN-Seq and Ribo-Zero-Seq mapped ~20% of bases to the transcriptome and >60% to intronic/intergenic regions. ~20% transcriptome; >60% intronic/intergenic
- ▲ All FF RNA-Seq protocols correlated highly with matched Agilent microarray data (Pearson > 0.8), while FFPE protocols correlated at ~0.7. Pearson 0.851 (mRNA-Seq), 0.832 (Ribo-Zero), 0.855 (DSN); 0.636 (Ribo-Zero-FFPE), 0.700 (DSN-FFPE)
- ▲ The two rRNA-depletion protocols were the most highly correlated pair of methods in both FF and FFPE samples; all pairwise FF comparisons exceeded 0.9 and FFPE-vs-FF mRNA-Seq exceeded 0.8. Pearson 0.961 (FF) and 0.934 (FFPE)
- ▲ Ribo-Zero-Seq-FFPE technical replicates were essentially indistinguishable. Pearson = 0.991 across 8 technical replicates
- ▲ 13.5 million reads from FF mRNA-Seq matched the microarray gene-detection baseline, whereas FF DSN-Seq/Ribo-Zero-Seq and Ribo-Zero-Seq-FFPE required 35–65 M reads and DSN-Seq-FFPE required 90 M reads. 13.5 M vs 35–65 M vs 90 M reads (baseline n = 16,975 genes)
- ▲ 41/44 UNC samples and 40/44 TCGA samples co-clustered tightly with their same-tumor protocol partner using the intrinsic gene list; the non-clustering UNC samples were all Ribo-Zero-Seq FFPE libraries. 41/44 and 40/44; non-clustered TCGA partners still correlated > 0.6
- correlation Pearson 0.961 (FF) and 0.934 (FFPE) between Ribo-Zero-Seq and DSN-Seq (Cross-protocol transcript quantification concordance, most correlated pair)
- correlation Pearson = 0.991 (Technical reproducibility of Ribo-Zero-Seq-FFPE replicates (TCGA))
- correlation 0.851 (mRNA-Seq), 0.832 (Ribo-Zero-Seq), 0.855 (DSN-Seq), 0.636 (Ribo-Zero-FFPE), 0.700 (DSN-FFPE) (Median Pearson correlation of each RNA-Seq protocol to matched Agilent microarray, UNC dataset)
- pvalue p < 0.001 (Higher rRNA fraction in DSN-Seq libraries vs mRNA-Seq/Ribo-Zero-Seq; also for Ribo-Zero-Seq's less biased 5′-to-3′ ratio vs mRNA-Seq on FF)
- pvalue mRNA-Seq p < 0.001; Ribo-Zero-Seq p = 0.002 (Lower median CV of transcript coverage vs DSN-Seq in FF libraries (top 1000 expressed transcripts))
- count 16,975 expressed Entrez genes on Agilent arrays; 15,206 detected by both microarray and RNA-Seq; 20,531 non-ribosomal genes in GAF 2.1 reference (Gene detection baseline for depth simulation and platform overlap)
- count 410 genes at FDR 0 differentially expressed between mRNA-Seq and Ribo-Zero-Seq; 104 genes at FDR 0 between Ribo-Zero-Seq and DSN-Seq (38 lowly quantified by DSN-Seq) (SAM protocol-vs-protocol differential expression, enriched for snoRNAs and histone RNAs)
- count ~350 shared differentially expressed genes per protocol pair; >300 genes consistently identified by all protocols (SAM of top 500 variable genes between 3 Basal-like and 3 Luminal UNC tumors)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study compares six gene-expression profiling protocols (DNA microarray, mRNA-Seq, Ribo-Zero-Seq, and DSN-Seq, applied to fresh-frozen and/or FFPE RNA from paired UNC and TCGA tumor samples) using descriptive metrics (percent rRNA, percent aligned bases, coverage uniformity, 5'-3' bias) summarized as central values with ranges in a table, pairwise Pearson correlations and Deming regression to assess concordance/sensitivity between protocols, hierarchical clustering to assess sample-pairing fidelity, and Significance Analysis of Microarrays (SAM) to identify differentially expressed genes between subtypes and between protocols. Some pairwise comparisons report threshold or exact p-values (e.g., p < 0.001, p = 0.002) without naming the specific statistical test used.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Unspecified statistical test (test name not stated in text) | Comparison of % rRNA between DSN-Seq vs mRNA-Seq/Ribo-Zero-Seq; comparison of median CV coverage across protocols; comparison of 5' to 3' bias between Ribo-Zero-Seq and mRNA-Seq | Per-protocol sample sizes as listed in Table 1 (e.g., 11 mRNA-Seq, 11 Ribo-Zero-Seq, 10 DSN-Seq in UNC dataset) | not stated |
| Pearson correlation | Concordance between RNA-Seq protocols and DNA microarray (Table 1); pairwise concordance among RNA-Seq protocols (Figure 3A, B, C); technical replicate reproducibility (Ribo-Zero-Seq-FFPE replicates) | Paired samples per protocol as listed in Table 1; 8 technical replicates for the Ribo-Zero-Seq-FFPE reproducibility check | not stated |
| Deming regression | Estimating relative sensitivity (slope) between pairs of RNA-Seq protocols (Figure 3D) | UNC dataset paired samples | not stated |
| Significance Analysis of Microarrays (SAM) | Differential expression between Basal-like and Luminal subtypes across protocols; differential expression between mRNA-Seq vs Ribo-Zero-Seq and Ribo-Zero-Seq vs DSN-Seq (FF samples) | Three Basal-like and three Luminal UNC samples for the subtype comparison; FF sample sets per protocol from Table 1 for the protocol comparisons | not stated |
| Hierarchical clustering analysis (unsupervised) | Assessment of whether same-tumor samples profiled by different protocols co-cluster, using an intrinsic gene list | 904 samples (88 UNC/TCGA paired samples plus 725 additional breast tumors and 91 normal breast tissues) | not stated |
-
Differences in % rRNA, coverage CV, and 5'-3' bias between protocols are reported with p-values (e.g., p < 0.001, p = 0.002) without naming the specific statistical test applied.↳ Could also: Explicitly naming and reporting the test statistic for a parametric test (e.g., Student's t-test) or nonparametric test (e.g., Wilcoxon rank-sum/Mann-Whitney U) used for each comparison. — Naming the test and reporting its statistic alongside the p-value lets readers independently evaluate the assumptions (e.g., normality, equal variance) and the magnitude of the underlying difference.
-
Numerous pairwise comparisons of protocols are made (rRNA %, aligned bases, CV coverage, 5'-3' bias, correlations) without a stated family-wise multiple-testing correction, aside from SAM's FDR use in the differential expression analyses.↳ Could also: Applying a family-wise correction such as Bonferroni or a Benjamini-Hochberg FDR procedure across the full set of pairwise metric comparisons. — This would control the overall false-positive rate when many related comparisons are examined together, complementing the FDR already used within the SAM analyses.
-
Differential expression between protocols and between subtypes was assessed using SAM, a method originally developed for microarray intensity data.↳ Could also: A count-based RNA-seq differential expression framework (e.g., DESeq2, edgeR, or limma-voom) that models the mean-variance relationship of sequencing read counts directly. — These tools are designed specifically for count-distributed RNA-seq data and could be used alongside or as an alternative lens on the same comparisons, particularly for the RNA-Seq-based protocol comparisons.
-
Variability in Table 1 metrics is summarized using minimum-maximum ranges rather than SD, SEM, or confidence intervals.↳ Could also: Reporting SD, SEM, or 95% confidence intervals alongside or instead of ranges. — SD/SEM/CI are less sensitive to single extreme values than a min-max range and allow more direct statistical comparison of variability and precision across protocols and sample sizes.
-
Concordance across protocols and against microarray data is assessed primarily using Pearson correlation.↳ Could also: Complementing Pearson correlation with Spearman rank correlation or a Bland-Altman-style agreement analysis. — Spearman correlation is less sensitive to non-linear relationships or outliers, and Bland-Altman-style plots can visualize systematic bias and its relationship to expression level, which can add to the Pearson correlation and Deming regression already used.
-
Sample sizes for each protocol/dataset combination are presented directly (Table 1) without a described power calculation.↳ Could also: Including an a priori or post hoc power/precision analysis for the key comparisons. — This would give readers additional context on the sensitivity of each comparison to detect a given effect size, particularly for protocols represented by smaller sample sizes (e.g., n = 4 for DSN-FFPE).
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
What deviates: Correlation-type claims reproduce cleanly and slightly better than reported (RiboZero-vs-DSN FF r=0.977 vs 0.961; FFPE 0.965 vs 0.934; all FF pairs >0.9), and the detected-gene surplus (657 vs ~550) is consistent. The real gaps are the SAM differential-expression counts (410→174/293; 104→27/37) and the technical-replicate claim, where the paper's 8 FFPE Ribo-Zero replicates are simply absent from GSE51783 and only a 2-sample mRNA-seq FF pair (r=0.996 vs 0.991) exists. Whose side: predominantly ours and the data-availability layer — samr was replaced by a documented simplified permutation SAM-equivalent, the gene universe and 'FDR=0' cut had to be guessed, raw reads are dbGaP-locked, and the S1..S48 column-to-sample mapping was assumed from GEO ordering. Severity: moderate on the DE numbers, negligible everywhere else; the central conclusion — rRNA-depletion RNA-seq is highly concordant with poly(A) RNA-seq and viable on FFPE — holds fully. No fabrication or non-derivability concern: nothing looks 'too perfect', and the direction of every reported effect is preserved.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.