Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Optimal scaling of digital transcriptomes.

PLoS One · 2013
L1 67/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
67/100
Reproducibility score
0.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 29% of all assessed papers rank 795 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Different-scope reproduction: the Normalizer.pm software itself was verified to run correctly end-to-end (19/19 attempted method variants, exit=0) and the paper's real dataset (GSE30611/ERP000546, 16/16 tissue FASTQ files) was downloaded and byte-verified intact, but the completed normalization runs used the repo's bundled SYNTHETIC test fixture, not the real data -- the Blat+Cufflinks alignment/quantification step needed to turn the real FASTQ into the paper's actual 29,665-gene count matrix was not executed this pass. Consequently every headline number in the paper (2,507/6,944 ubiquitous genes; per-method uniform-gene counts of 13-539; ES best of 539) remains unreproduced-not-yet-attempted, not mismatched: no invented numbers are reported. Housekeeping/geNorm methods were also skipped (missing guide-gene input file). This is a genuine, substantial remaining gap for a future pass to close by running the alignment/quantification stage on the already-downloaded, already-verified raw data.

💻 Code ↗ 🗄 Data: GSE30611

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Because digital transcript counts scale with sequencing depth and total RNA content, cross-sample comparison requires normalization; the paper asks which of many existing and novel scaling algorithms best renders expression levels comparable across samples, evaluated by application-independent, data-driven metrics rather than gold-standard RT-PCR or simulations.

Core claims
  • Fifteen existing and novel transcript-count normalization algorithms can be compared with two novel, mutually independent metrics: the number of "uniform" genes (sufficiently low coefficient of variation after normalization) and low average Spearman correlation between normalized expression profiles of gene pairs. method
  • The two most commonly used normalization methods — scaling to a fixed total value (CPM/RPKM) and equalizing expression of selected "housekeeping" genes — perform particularly poorly, worse even than normalization guided by randomly selected gene sets. finding
  • Seven of the fifteen algorithms approach what appears to be optimal normalization, and several independently converge on very similar solutions that differ markedly from standard methods. finding
  • Three of the best-performing algorithms rely on identifying "ubiquitous" genes — expressed in all samples but never at very high or very low levels — which by construction excludes most differentially expressed genes. mechanism
  • Ubiquitous genes include a "core" of genes expressed across many tissues in a mutually consistent (proportional) pattern, suitable as an internal normalization guide. finding
  • Four novel scaling algorithms are introduced — Random (mock control), Total Ubiquitous, Network Centrality Scaling (NCS), and Evolution Strategy — the last explicitly maximizing the number of uniform genes. method
  • High Spearman correlation between gene-pair sample rankings is mostly an artifact of distorted (incorrectly scaled) expression values rather than shared biology, so decorrelation is a valid normalization-quality criterion. mechanism
  • Robustly normalized expression values are a prerequisite for reliable identification of differentially expressed and tissue-specific genes as candidate biomarkers. resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (75 bp single-end reads), Illumina BodyMap 2.0 16 human tissues: adipose, adrenal, brain, breast, colon, kidney, heart, liver, lung, lymph node, prostate, skeletal muscle, white blood cells, ovary, testes, thyroid none (tissue comparison) average fold coverage per transcribed base for 133,810 transcripts summarized to 29,665 genes Illumina, Inc.; GEO series GSE30611
read alignment / mapping benchmark BodyMap 2.0 RNA-seq reads vs human reference genome hg19/GRCh37 aligner comparison (Blat vs TopHat/Bowtie) percent of total reads mapped and percent uniquely mapped Blat v34 (default parameters), TopHat/Bowtie, custom Perl scripts, SAM output
transcript assembly and expression quantification mapped BodyMap 2.0 reads, Ensembl Gene Models GRCh37.62 none coverage column (average fold coverage per transcribed base), exonic uniquely mapped reads only, GC- and multi-mapping-corrected, excluding novel isoforms, pseudogenes, rRNA, mitochondrial sequences Cufflinks v1.0.3 (-b -G -u -M)
computational normalization benchmark of 15 algorithms 16-sample BodyMap 2.0 gene-level expression matrix algorithm choice: CPM/RPKM, Total Counts, Upper Quartile, Upper Decile, Housekeeping (single gene), geNorm, Stability, All Ubiquitous, Random Ubiquitous, TMM, Quantile Normalization, Random, Total Ubiquitous, NCS, Evolution Strategy sample-specific scaling factors, number of uniform genes (low coefficient of variation), average/median Spearman correlation of gene-pair sample rankings
ubiquitous-gene definition / fitness terrain analysis 16-sample BodyMap 2.0 expression matrix systematic variation of lower and upper expression-percentile trimming cutoffs at 5% resolution number of uniform genes resulting from each cutoff combination
decorrelation analysis (pairwise Spearman correlation of gene-derived sample rankings) ubiquitous genes across the 16 BodyMap 2.0 tissue samples normalization method applied (e.g. Total Counts vs NCS) distribution and average Spearman correlation over 100,000 randomly selected gene pairs
pairwise sample expression comparison (scatter analysis) human liver vs testes (BodyMap 2.0) normalization method (none, Total Counts, NCS) average coverage per base for 15,861 genes with nonzero expression in both; NCS guide-gene weights
housekeeping-gene expression assessment 16 BodyMap 2.0 tissue samples none detection/nonzero expression of the 10 geNorm control genes (ACTB, B2M, GAPDH, HMBS, HPRT1, RPL13A, SDHA, TBP, UBC, YWHAZ)
Key results
  • Blat mapped more RNA-seq reads than TopHat/Bowtie, raising the sample-average mapped fraction from 89.2% to 97.2% +7.94% +/− 2.3% of total reads
  • Uniquely mapped reads increased with Blat, from a sample average of 82.3% to 91.1% +8.85% +/− 2.05%
  • Nearly one third of quantified genes were observed in all 16 tissue samples 9,588 of 29,665 genes (32.2%)
  • Using 30th–85th percentile trimming, 2,507 ubiquitous genes were identified in the 16-sample BodyMap 2.0 data set; with the more permissive 5% trimming of Robinson and Oshlack, 6,944 genes were retained 2,507 and 6,944 genes
  • Fitness terrain analysis showed the choice of trimming cutoffs is robust: the upper cutoff in the range 70%–95% and the lower cutoff from 0% to 60%; the 30th–85th percentile cutoffs maximize the number of uniform genes upper 70–95%, lower 0–60%
  • Under Total Counts scaling, most genes appear higher in testes than liver so most gene pairs rank testes over liver, producing high Spearman correlations; under NCS scaling roughly half the gene pairs rank each way, giving much lower average correlation approximately half of gene pairs each direction under NCS
  • In the liver-vs-testes NCS comparison, 39 genes carried weight >0.5 and thus dominated the scaling among positively weighted guide genes 39 genes with weight >0.5 of 15,861 compared
  • YWHAZ, one of the ten geNorm housekeeping genes, was detected in only 9 of 16 tissues and had to be excluded, leaving nine housekeeping genes for the study 9 of 16 samples
Key statistics
  • other +7.94% +/−2.3% additional total reads mapped (89.2% → 97.2%) (Blat v34 vs TopHat/Bowtie mapping rate per sample)
  • other +8.85% +/−2.05% additional uniquely mapped reads (82.3% → 91.1%) (Blat vs TopHat/Bowtie unique mapping per sample)
  • count 133,810 observed transcripts; 29,665 genes (Cufflinks quantification of BodyMap 2.0)
  • count 9,588 (32.2%) (genes observed in all 16 samples)
  • count 2,507 ubiquitous genes (30th–85th percentile trimming); 6,944 (5% trimming) (ubiquitous gene sets used as normalization guides)
  • count 100,000 gene pairs (random sample used to characterize the gene-pair Spearman correlation distribution)
  • count 15,861 genes (genes with nonzero expression in both liver and testes (Fig. 2))
  • other random scaling factors 2^r, r uniform in [−0.5,0.5), i.e. 0.707 to 1.414 (range of the mock 'Random' normalization method)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational/methods-comparison paper rather than a classical hypothesis-testing study: the authors compare fifteen normalization algorithms for RNA-seq transcript counts across 16 human tissue samples (BodyMap 2.0), evaluating success using two custom aggregate metrics — the count of 'uniform' genes (low coefficient of variation after normalization) and the average Spearman correlation between normalized expression rankings of gene pairs — rather than per-gene inferential tests. Results are reported mainly as descriptive counts, percentages, and sample averages with a stated ± dispersion value, without classical significance testing, p-values, or confidence intervals in the excerpted text.

Replicationunclear Sample size16 human tissue RNA-seq samples (BodyMap 2.0, one sample per tissue type); 29,665 genes total, 9,588 (32.2%) observed in all 16 samples; gene-pair correlation analyses subsampled to 100,000 pairs Groups15 normalization algorithms applied to the same 16-tissue RNA-seq dataset Pairingna Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Spearman rank correlation (pairwise gene-to-gene, decorrelation analysis) comparing normalized expression rankings between ubiquitous gene pairs across samples/methods (e.g., Fig. 2) 100,000 randomly sampled gene pairs per comparison not stated
Coefficient of variation (used to define 'uniform' genes) evaluating normalization success across the 16 BodyMap 2.0 tissue samples 16 samples; 2,507–6,944 ubiquitous genes depending on trimming cutoffs not stated
Jongeneel's specificity measure (rule-based, not a formal significance test) defining genes 'specific' to a given tissue sample genes observed in at least half of the 16 samples not stated
Approaches that could also have been used
  • Alignment/mapping improvements (e.g., percent reads mapped by Blat vs. TopHat/Bowtie) are reported as an average across samples with a ± value whose type (SD, SEM, or other) is not specified.
    Could also: Explicitly labeling the dispersion measure (e.g., stating SD vs. SEM) or reporting a 95% confidence interval alongside the mean — Would let readers directly distinguish variability across the 16 samples from precision of the estimated average, and CIs are often preferred for conveying uncertainty in small sample averages.
  • Normalization algorithms are ranked and compared using aggregate metrics (count of 'uniform' genes, average Spearman correlation across 100,000 gene pairs) without an accompanying formal statistical test of whether differences between methods are significant.
    Could also: A repeated-measures/blocked comparison across methods (e.g., Friedman test) with post-hoc pairwise correction, or bootstrap resampling of the summary metrics to obtain confidence intervals — Would provide an explicit inferential statistic and interval estimate for method-to-method differences, complementing the descriptive ranking already presented.
  • The distribution of pairwise gene-rank Spearman correlations is summarized mainly by its average (with a note that the median gives similar results).
    Could also: Reporting the full distribution shape (e.g., median with IQR, or a histogram/density plot) alongside the mean — Correlation distributions across large numbers of gene pairs can be skewed, and pairing the mean with IQR or a visual distribution can convey shape and spread more completely than a single summary value.
  • Genes are classified as tissue-'specific' using a fixed rule-based threshold (Jongeneel's specificity measure: expression in one sample exceeding the sum of all others), rather than a model-based statistical test with an associated p-value.
    Could also: A dedicated differential-expression testing framework such as edgeR or DESeq2 (negative binomial generalized linear model) producing per-gene p-values and FDR-adjusted q-values — Would allow explicit control of the false discovery rate when flagging many genes as tissue-specific across thousands of comparisons simultaneously.
  • The main RNA-seq dataset consists of one sample per tissue (16 tissues total) rather than biological replicates within each tissue.
    Could also: Incorporating biological replicates per tissue where feasible, paired with a variance-partitioning or mixed-effects approach — Replicates would let within-tissue (technical/biological) variability be separated from between-tissue and normalization-driven variability, supporting variance-based statistical comparison of normalization methods.
  • No correction for multiple comparisons is described for the large-scale pairwise gene-to-gene correlation analysis (100,000 sampled gene pairs).
    Could also: Applying a multiple-testing correction such as Benjamini-Hochberg FDR if individual gene-pair correlations were to be tested and interpreted as statistically significant — Would guard against an inflated false-positive rate if, beyond the current aggregate summary, individual gene-pair correlations were later used as discrete significance calls.
Software: Blat v34 · TopHat/Bowtie · Cufflinks v1.0.3 · custom Perl scripts

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

code-executability
Reported
Normalizer.pm implements a cascade of normalization methods (CPM, Total, Upper Quartile/Decile, Housekeeping, geNorm, Stability, All/Random/Total Ubiquitous, TMM, Quantile, Random, NCS, ES) callable via normalize.pl; requires Graph + Graph::Centrality::Pagerank Perl modules.
Reproduced
All 19 method/variant runs (cpm,total,ubiquitous,upperquartile,upperdecile,tmm,quantile,random,net,stability,allubiquitous,10ubiquitous,100ubiquitous,alltrimmed,100trimmed,median,deseq,desequbiq,es) completed with exit=0 and empty .err on the repo's bundled sample.data.gz fixture (SLURM «job», 2423259), producing scaling factors + normalized matrices for each.
exact
real-dataset-acquisition
Reported
Paper uses 16 raw single-end 75bp RNA-seq datasets from Illumina BodyMap 2.0, GEO GSE30611 (ENA ERP000546), one per tissue.
Reproduced
16/16 single-end FASTQ files downloaded to «infra» scratch («job»); file count and byte size match ENA-reported fastq_bytes exactly for all 16 files.
exact
ubiquitous-gene-sets
Reported
2,507 ubiquitous genes (strict 30th-85th percentile cutoff) / 6,944 genes (loose 5-95% cutoff), computed on the real 29,665-gene GSE30611 expression matrix.
Reproduced
not computed — requires the real-data gene count matrix, which requires Blat v34 alignment to hg19 + Cufflinks v1.0.3 quantification against Ensembl GRCh37.62; this alignment/quantification stage was not run in this pass (only raw FASTQ download was completed).
partial
per-method-uniform-gene-counts
Reported
Housekeeping best 13 (TBP); geNorm 305 (7 combined genes); Upper Decile 373; Quantile Normalization 374; Stability(100 genes) 481; Total Ubiquitous 492; NCS 492; ES best 539 (liver factor ~1.698), all on real GSE30611 data.
Reproduced
not computed on real data (same upstream blocker as ubiquitous-gene-sets). The 19 completed method runs («job») instead ran on the repo's synthetic sample.data.gz test fixture (31 samples x 5000 synthetic genes), which is not comparable to these paper numbers. On that synthetic fixture the ES driver converged to uniform_genes=0.
partial
housekeeping-genorm-methods
Reported
Housekeeping (single control gene) and geNorm (combination of housekeeping genes) normalization methods, using candidate genes GAPDH/TBP/RPL13A etc.
Reproduced
not run — normalize.pl --method=guide|genorm requires a user-supplied guide-gene weight file (--guide=) that was not prepared/available in this pass.
partial
es-convergence-behavior
Reported
ES stops when no improvement in fitness (uniform-gene count) for 100 consecutive rounds, or a time limit; typically converges in a few hours on real data; best real-data result 539 uniform genes.
Reproduced
Custom run_es.pl driver (wraps Normalizer.pm's evolution_strategy()) on the synthetic fixture correctly exhibits the described early-stopping behavior ("ES seems to have converged" after ~41 rounds, 38s), but reaches uniform_genes=0 on synthetic data with no real expression structure — behavior sanity-checked, final value not comparable to the paper's real-data result.
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 67/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

Different-scope, not divergent-result. The authors' raw input is fully public and was obtained 1:1 (16/16 GSE30611/ERP000546 FASTQ files, byte-exact against ENA metadata) and Normalizer.pm was verified runnable end-to-end (19/19 method variants, exit=0, ES early-stopping reproduced after ~41 rounds). However, the Blat v34 + Cufflinks v1.0.3 stage that builds the paper's 29,665-gene matrix was never run, so all 19 normalization runs used the repo's synthetic sample.data.gz fixture — meaning none of the headline numbers (2,507/6,944 ubiquitous genes; uniform-gene counts 13–539; ES best 539) were checked. The gap is squarely on our side (scope truncation plus a missing --guide= housekeeping file that also blocked geNorm), not an authors' defect, and crucially no numbers were invented — the claims are graded 'not computed', which is exactly the right disclosure. Verdict yellow throughout: honest and technically clean, but the paper's substantive conclusion remains untested.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.