Optimal scaling of digital transcriptomes.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Different-scope reproduction: the Normalizer.pm software itself was verified to run correctly end-to-end (19/19 attempted method variants, exit=0) and the paper's real dataset (GSE30611/ERP000546, 16/16 tissue FASTQ files) was downloaded and byte-verified intact, but the completed normalization runs used the repo's bundled SYNTHETIC test fixture, not the real data -- the Blat+Cufflinks alignment/quantification step needed to turn the real FASTQ into the paper's actual 29,665-gene count matrix was not executed this pass. Consequently every headline number in the paper (2,507/6,944 ubiquitous genes; per-method uniform-gene counts of 13-539; ES best of 539) remains unreproduced-not-yet-attempted, not mismatched: no invented numbers are reported. Housekeeping/geNorm methods were also skipped (missing guide-gene input file). This is a genuine, substantial remaining gap for a future pass to close by running the alignment/quantification stage on the already-downloaded, already-verified raw data.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusBecause digital transcript counts scale with sequencing depth and total RNA content, cross-sample comparison requires normalization; the paper asks which of many existing and novel scaling algorithms best renders expression levels comparable across samples, evaluated by application-independent, data-driven metrics rather than gold-standard RT-PCR or simulations.
- ★ Fifteen existing and novel transcript-count normalization algorithms can be compared with two novel, mutually independent metrics: the number of "uniform" genes (sufficiently low coefficient of variation after normalization) and low average Spearman correlation between normalized expression profiles of gene pairs. method
- ★ The two most commonly used normalization methods — scaling to a fixed total value (CPM/RPKM) and equalizing expression of selected "housekeeping" genes — perform particularly poorly, worse even than normalization guided by randomly selected gene sets. finding
- ★ Seven of the fifteen algorithms approach what appears to be optimal normalization, and several independently converge on very similar solutions that differ markedly from standard methods. finding
- ★ Three of the best-performing algorithms rely on identifying "ubiquitous" genes — expressed in all samples but never at very high or very low levels — which by construction excludes most differentially expressed genes. mechanism
- ★ Ubiquitous genes include a "core" of genes expressed across many tissues in a mutually consistent (proportional) pattern, suitable as an internal normalization guide. finding
- ★ Four novel scaling algorithms are introduced — Random (mock control), Total Ubiquitous, Network Centrality Scaling (NCS), and Evolution Strategy — the last explicitly maximizing the number of uniform genes. method
- High Spearman correlation between gene-pair sample rankings is mostly an artifact of distorted (incorrectly scaled) expression values rather than shared biology, so decorrelation is a valid normalization-quality criterion. mechanism
- Robustly normalized expression values are a prerequisite for reliable identification of differentially expressed and tissue-specific genes as candidate biomarkers. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (75 bp single-end reads), Illumina BodyMap 2.0 | 16 human tissues: adipose, adrenal, brain, breast, colon, kidney, heart, liver, lung, lymph node, prostate, skeletal muscle, white blood cells, ovary, testes, thyroid | none (tissue comparison) | average fold coverage per transcribed base for 133,810 transcripts summarized to 29,665 genes | Illumina, Inc.; GEO series GSE30611 |
| read alignment / mapping benchmark | BodyMap 2.0 RNA-seq reads vs human reference genome hg19/GRCh37 | aligner comparison (Blat vs TopHat/Bowtie) | percent of total reads mapped and percent uniquely mapped | Blat v34 (default parameters), TopHat/Bowtie, custom Perl scripts, SAM output |
| transcript assembly and expression quantification | mapped BodyMap 2.0 reads, Ensembl Gene Models GRCh37.62 | none | coverage column (average fold coverage per transcribed base), exonic uniquely mapped reads only, GC- and multi-mapping-corrected, excluding novel isoforms, pseudogenes, rRNA, mitochondrial sequences | Cufflinks v1.0.3 (-b -G -u -M) |
| computational normalization benchmark of 15 algorithms | 16-sample BodyMap 2.0 gene-level expression matrix | algorithm choice: CPM/RPKM, Total Counts, Upper Quartile, Upper Decile, Housekeeping (single gene), geNorm, Stability, All Ubiquitous, Random Ubiquitous, TMM, Quantile Normalization, Random, Total Ubiquitous, NCS, Evolution Strategy | sample-specific scaling factors, number of uniform genes (low coefficient of variation), average/median Spearman correlation of gene-pair sample rankings | — |
| ubiquitous-gene definition / fitness terrain analysis | 16-sample BodyMap 2.0 expression matrix | systematic variation of lower and upper expression-percentile trimming cutoffs at 5% resolution | number of uniform genes resulting from each cutoff combination | — |
| decorrelation analysis (pairwise Spearman correlation of gene-derived sample rankings) | ubiquitous genes across the 16 BodyMap 2.0 tissue samples | normalization method applied (e.g. Total Counts vs NCS) | distribution and average Spearman correlation over 100,000 randomly selected gene pairs | — |
| pairwise sample expression comparison (scatter analysis) | human liver vs testes (BodyMap 2.0) | normalization method (none, Total Counts, NCS) | average coverage per base for 15,861 genes with nonzero expression in both; NCS guide-gene weights | — |
| housekeeping-gene expression assessment | 16 BodyMap 2.0 tissue samples | none | detection/nonzero expression of the 10 geNorm control genes (ACTB, B2M, GAPDH, HMBS, HPRT1, RPL13A, SDHA, TBP, UBC, YWHAZ) | — |
- ▲ Blat mapped more RNA-seq reads than TopHat/Bowtie, raising the sample-average mapped fraction from 89.2% to 97.2% +7.94% +/− 2.3% of total reads
- ▲ Uniquely mapped reads increased with Blat, from a sample average of 82.3% to 91.1% +8.85% +/− 2.05%
- – Nearly one third of quantified genes were observed in all 16 tissue samples 9,588 of 29,665 genes (32.2%)
- – Using 30th–85th percentile trimming, 2,507 ubiquitous genes were identified in the 16-sample BodyMap 2.0 data set; with the more permissive 5% trimming of Robinson and Oshlack, 6,944 genes were retained 2,507 and 6,944 genes
- – Fitness terrain analysis showed the choice of trimming cutoffs is robust: the upper cutoff in the range 70%–95% and the lower cutoff from 0% to 60%; the 30th–85th percentile cutoffs maximize the number of uniform genes upper 70–95%, lower 0–60%
- ▼ Under Total Counts scaling, most genes appear higher in testes than liver so most gene pairs rank testes over liver, producing high Spearman correlations; under NCS scaling roughly half the gene pairs rank each way, giving much lower average correlation approximately half of gene pairs each direction under NCS
- – In the liver-vs-testes NCS comparison, 39 genes carried weight >0.5 and thus dominated the scaling among positively weighted guide genes 39 genes with weight >0.5 of 15,861 compared
- – YWHAZ, one of the ten geNorm housekeeping genes, was detected in only 9 of 16 tissues and had to be excluded, leaving nine housekeeping genes for the study 9 of 16 samples
- other +7.94% +/−2.3% additional total reads mapped (89.2% → 97.2%) (Blat v34 vs TopHat/Bowtie mapping rate per sample)
- other +8.85% +/−2.05% additional uniquely mapped reads (82.3% → 91.1%) (Blat vs TopHat/Bowtie unique mapping per sample)
- count 133,810 observed transcripts; 29,665 genes (Cufflinks quantification of BodyMap 2.0)
- count 9,588 (32.2%) (genes observed in all 16 samples)
- count 2,507 ubiquitous genes (30th–85th percentile trimming); 6,944 (5% trimming) (ubiquitous gene sets used as normalization guides)
- count 100,000 gene pairs (random sample used to characterize the gene-pair Spearman correlation distribution)
- count 15,861 genes (genes with nonzero expression in both liver and testes (Fig. 2))
- other random scaling factors 2^r, r uniform in [−0.5,0.5), i.e. 0.707 to 1.414 (range of the mock 'Random' normalization method)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational/methods-comparison paper rather than a classical hypothesis-testing study: the authors compare fifteen normalization algorithms for RNA-seq transcript counts across 16 human tissue samples (BodyMap 2.0), evaluating success using two custom aggregate metrics — the count of 'uniform' genes (low coefficient of variation after normalization) and the average Spearman correlation between normalized expression rankings of gene pairs — rather than per-gene inferential tests. Results are reported mainly as descriptive counts, percentages, and sample averages with a stated ± dispersion value, without classical significance testing, p-values, or confidence intervals in the excerpted text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Spearman rank correlation (pairwise gene-to-gene, decorrelation analysis) | comparing normalized expression rankings between ubiquitous gene pairs across samples/methods (e.g., Fig. 2) | 100,000 randomly sampled gene pairs per comparison | not stated |
| Coefficient of variation (used to define 'uniform' genes) | evaluating normalization success across the 16 BodyMap 2.0 tissue samples | 16 samples; 2,507–6,944 ubiquitous genes depending on trimming cutoffs | not stated |
| Jongeneel's specificity measure (rule-based, not a formal significance test) | defining genes 'specific' to a given tissue sample | genes observed in at least half of the 16 samples | not stated |
-
Alignment/mapping improvements (e.g., percent reads mapped by Blat vs. TopHat/Bowtie) are reported as an average across samples with a ± value whose type (SD, SEM, or other) is not specified.↳ Could also: Explicitly labeling the dispersion measure (e.g., stating SD vs. SEM) or reporting a 95% confidence interval alongside the mean — Would let readers directly distinguish variability across the 16 samples from precision of the estimated average, and CIs are often preferred for conveying uncertainty in small sample averages.
-
Normalization algorithms are ranked and compared using aggregate metrics (count of 'uniform' genes, average Spearman correlation across 100,000 gene pairs) without an accompanying formal statistical test of whether differences between methods are significant.↳ Could also: A repeated-measures/blocked comparison across methods (e.g., Friedman test) with post-hoc pairwise correction, or bootstrap resampling of the summary metrics to obtain confidence intervals — Would provide an explicit inferential statistic and interval estimate for method-to-method differences, complementing the descriptive ranking already presented.
-
The distribution of pairwise gene-rank Spearman correlations is summarized mainly by its average (with a note that the median gives similar results).↳ Could also: Reporting the full distribution shape (e.g., median with IQR, or a histogram/density plot) alongside the mean — Correlation distributions across large numbers of gene pairs can be skewed, and pairing the mean with IQR or a visual distribution can convey shape and spread more completely than a single summary value.
-
Genes are classified as tissue-'specific' using a fixed rule-based threshold (Jongeneel's specificity measure: expression in one sample exceeding the sum of all others), rather than a model-based statistical test with an associated p-value.↳ Could also: A dedicated differential-expression testing framework such as edgeR or DESeq2 (negative binomial generalized linear model) producing per-gene p-values and FDR-adjusted q-values — Would allow explicit control of the false discovery rate when flagging many genes as tissue-specific across thousands of comparisons simultaneously.
-
The main RNA-seq dataset consists of one sample per tissue (16 tissues total) rather than biological replicates within each tissue.↳ Could also: Incorporating biological replicates per tissue where feasible, paired with a variance-partitioning or mixed-effects approach — Replicates would let within-tissue (technical/biological) variability be separated from between-tissue and normalization-driven variability, supporting variance-based statistical comparison of normalization methods.
-
No correction for multiple comparisons is described for the large-scale pairwise gene-to-gene correlation analysis (100,000 sampled gene pairs).↳ Could also: Applying a multiple-testing correction such as Benjamini-Hochberg FDR if individual gene-pair correlations were to be tested and interpreted as statistically significant — Would guard against an inflated false-positive rate if, beyond the current aggregate summary, individual gene-pair correlations were later used as discrete significance calls.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Different-scope, not divergent-result. The authors' raw input is fully public and was obtained 1:1 (16/16 GSE30611/ERP000546 FASTQ files, byte-exact against ENA metadata) and Normalizer.pm was verified runnable end-to-end (19/19 method variants, exit=0, ES early-stopping reproduced after ~41 rounds). However, the Blat v34 + Cufflinks v1.0.3 stage that builds the paper's 29,665-gene matrix was never run, so all 19 normalization runs used the repo's synthetic sample.data.gz fixture — meaning none of the headline numbers (2,507/6,944 ubiquitous genes; uniform-gene counts 13–539; ES best 539) were checked. The gap is squarely on our side (scope truncation plus a missing --guide= housekeeping file that also blocked geNorm), not an authors' defect, and crucially no numbers were invented — the claims are graded 'not computed', which is exactly the right disclosure. Verdict yellow throughout: honest and technically clean, but the paper's substantive conclusion remains untested.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.