Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

An extensive evaluation of read trimming effects on Illumina NGS data analysis.

PLoS One · 2013
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Both Table-2 claims reproduce closely: untrimmed mappability 72.559% (paper: 72.189%, +0.37pp) and Sickle Q=20-trimmed mappability 96.041% (paper: 95.971%, +0.07pp), using Tophat 2.1.1/Bowtie2 2.5.4 against hg19 (paper used Tophat 2.0.5 and a 2013 GTF snapshot no longer separately retrievable) -- both graded within-tol given these unavoidable version/annotation differences. A supplementary Q=10/30/40 sweep (not a formal Table-2 claim) reproduces the paper's qualitative narrative of an optimal intermediate quality threshold: mappability peaks at Q20 (96.0%) and degrades at both lower (Q10: 93.4%) and higher (Q30: 91.7%, Q40: 70.4%) thresholds. The SRR002073 dataset (9,104,944 single-end 33bp human HepG2 RNA-Seq reads, confirmed via 3 independent concordant counts) is complete and delivers exactly what the paper describes, though it shows the high per-base quality variance typical of 2008-era Illumina GA data (69% discarded by Sickle at Q20). Out of scope for this RU (not attempted): the paper's other 3 datasets and ~8 other trimming tools, and downstream SNP/assembly experiments elsewhere in the paper.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper asks whether and how quality trimming of Illumina NGS reads affects downstream analyses, and which trimming algorithm/quality-threshold combinations optimize the tradeoff between information retained and reliability across RNA-Seq, SNP calling and de novo genome assembly.

Core claims
  • Read trimming increases the quality and reliability of downstream NGS analyses (RNA-Seq mapping, SNP identification, genome assembly) while reducing execution time and computational resources. finding
  • Intermediate quality thresholds (Q between 20 and 30) give the best overall trimming benefit; very stringent thresholds (Q>30) remove too much data and degrade assembly quality. finding
  • Different trimming tools have different optimal Q thresholds and different Q-versus-mappability trends, so no single tool/threshold is universally best; performance depends on the Q distribution of the input dataset. finding
  • Nine trimming algorithms belonging to two algorithm families (running-sum and window-based) were systematically compared on four datasets across three NGS applications — the first such comprehensive assessment. method
  • A novel optimized implementation of Mott's trimming algorithm is introduced (ERNE-FILTER). resource
  • Low-quality untrimmed reads generate false k-mers that increase misassembly chance and computational demand; trimming mitigates this, though error-modeling assemblers such as ABySS partially compensate. mechanism
  • Tools operating on both 5' and 3' read ends (e.g. ERNE-FILTER) or allowing low-quality islands within high-quality stretches (e.g. ConDeTri) benefit low-quality datasets more than 3'-only tools such as FASTX. finding
  • APOMAC (Average Percentage Of Minor Allele Calls) in dihaploid samples serves as a direct estimate of false-positive SNP calling. method
Experimental setups
Assay System Perturbation Readout Platform
RNA-Seq read alignment (spliced/gapped alignment to reference genome) Homo sapiens RNA-Seq dataset (low quality, publicly available) read trimming with 9 algorithms at varying Q thresholds vs. untrimmed number of surviving reads and percentage of reads/nucleotides aligning to the reference genome; percentage of reads mapping within UCSC gene models Illumina; FastQC for quality assessment
RNA-Seq read alignment (spliced/gapped alignment to reference genome) Arabidopsis thaliana RNA-Seq dataset (high quality) read trimming with 9 algorithms at varying Q thresholds vs. untrimmed number of surviving reads and percentage mapped reads over reference genome Illumina; FastQC
SNP identification / variant calling from DNA-Seq alignment Prunus persica (peach) Lovell dihaploid variety, DNA-Seq reads aligned to peach reference genome (227 Mb) read trimming with 9 tools at varying Q thresholds (default Q=20; ConDeTri HQ=25, LQ=10) vs. untrimmed APOMAC (Average Percentage of Minor Allele Calls), APONAC (Average Percentage of Non-reference Allele Calls), number of high-confidence SNPs, number of covered nucleotides above minimum coverage thresholds Illumina
SNP identification / variant calling from DNA-Seq alignment Saccharomyces cerevisiae YDJ25 dihaploid strain DNA-Seq (near-perfect quality dataset) read trimming with 9 tools at varying Q thresholds vs. untrimmed APOMAC, APONAC, number of called SNPs Illumina
de novo genome assembly Prunus persica (peach) DNA-Seq reads, assembly compared to peach reference genome read trimming with 9 tools at varying Q thresholds vs. untrimmed N50 (bp), average and longest scaffold length, accuracy (% assembled nucleotides mappable to reference), recall (% reference genome covered), RAM peak and execution time ABySS assembler
de novo genome assembly Saccharomyces cerevisiae DNA-Seq reads read trimming with 9 tools at default/varying Q thresholds vs. untrimmed N50 (bp), accuracy, recall ABySS assembler
Read quality control / quality metric profiling All four Illumina datasets (Homo sapiens RNA-Seq, Arabidopsis thaliana RNA-Seq, Saccharomyces cerevisiae DNA-Seq, Prunus persica Lovell DNA-Seq) none average PHRED error score, GC content bias, position-specific quality variation, Q distribution FastQC
Computational resource benchmarking Prunus persica genome assembly pipeline trimming tool × Q threshold combinations vs. untrimmed RAM peak and wall-clock execution time ABySS assembler
Key results
  • In the low-quality Homo sapiens RNA-Seq dataset, mapping rate rose from 72.2% untrimmed to above 90% after trimming, peaking at 97.0% with ConDeTri (HQ=15, LQ=10) and 96.7% with SolexaQA (Q=5). 72.189% → 96.973% mapped reads
  • In the high-quality Arabidopsis thaliana RNA-Seq dataset, mappability increased from an untrimmed baseline of 82.8% to above 98.5% for all tools at stringent thresholds (Q>30); maximum 99.422% for Cutadapt, Sickle and Trimmomatic at Q=40. 82.774% → 99.422%
  • Trimming drastically reduced false-positive SNP indicators: APOMAC fell from about 30% to 10% or less of aligned nucleotides in both Prunus persica and yeast, achievable with any trimmer at Q≥20. ~30% → ≤10%
  • For the yeast genotyping dataset, APOMAC at default threshold dropped from 0.2367% (untrimmed) to 0.0409% with SolexaQA-BWA, the best-performing tool; in peach it dropped from 0.2909% to 0.0645%. yeast 0.2367% → 0.0409%; peach 0.2909% → 0.0645%
  • SolexaQA achieved the highest RNA-Seq quality while retaining the highest number of reads, making it the optimal tool for the loss-versus-quality tradeoff in low-quality RNA-Seq datasets.
  • Untrimmed peach data gave the highest assembly N50 (18,093 bp) but lower accuracy (95.116%) and much higher computational demand; stringent trimmers gave more fragmented assemblies with higher accuracy (ConDeTri N50 14,525 bp, accuracy 96.389%; SolexaQA N50 13,571 bp, accuracy 96.223%). N50 18,093 → 13,571 bp; accuracy 95.116% → 96.389%
  • In yeast assembly, trimming reduced N50 from 9,095 bp (untrimmed) to as low as 3,209 bp (SolexaQA) while increasing accuracy from 99.196% to 99.692% (Cutadapt, FASTX, PRINSEQ, SolexaQA-BWA). N50 9,095 → 3,209 bp; accuracy 99.196% → 99.692%
  • The drop in called SNPs occurs at trimming thresholds matching each dataset's Q-distribution inflection point: around Q=35 (abrupt) for Prunus persica, whose inflection is ~Q=35, and above Q=36 (gradual) for the higher-quality S. cerevisiae dataset, whose inflection is Q=37. Prunus inflection Q=35; yeast inflection Q=37
Key statistics
  • other 72.189% (untrimmed) vs 96.973% (ConDeTri HQ=15,LQ=10) (Max % mapped reads, Homo sapiens RNA-Seq dataset)
  • other 82.774% (untrimmed) vs 99.422% (Cutadapt/Sickle/Trimmomatic, Q=40) (Max % mapped reads, Arabidopsis thaliana RNA-Seq dataset)
  • other 0.2367% untrimmed vs 0.0409% SolexaQA-BWA (APOMAC at default threshold, yeast genotyping dataset)
  • other 0.2909% untrimmed vs 0.0645% SolexaQA-BWA (APOMAC at default threshold, peach genotyping dataset)
  • other N50 9,095 bp untrimmed; range 3,209–6,357 bp after trimming (Yeast de novo genome assembly N50 at default threshold)
  • other N50 18,093 bp untrimmed; range 13,571–17,692 bp after trimming (Peach de novo genome assembly N50 at default threshold)
  • count 9 trimming algorithms, 4 datasets, 3 NGS applications (Study design scale)
  • other p = 10^(−Q/10); Q range 0–41, error rate 1 to 0.000079433 (PHRED quality score to error probability conversion for recent Illumina runs)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper is a comparative benchmarking study rather than a hypothesis-testing study: nine read-trimming tools were applied at multiple quality thresholds to four Illumina datasets, and their effects on RNA-Seq mapping percentages, SNP-calling false-positive indices (APOMAC/APONAC), and de novo genome-assembly metrics (N50, accuracy, recall) were compared as point-estimate percentages, counts, and summary metrics presented in tables and bar/scatter plots. Comparisons are made across tools and quality thresholds against an untrimmed baseline, with conclusions drawn from the magnitude and direction of these metrics rather than from reported significance tests, replicate-based variance, or multiplicity-adjusted comparisons.

Replicationunclear Groupsnine trimming tools across a range of quality (Q) thresholds, compared to each other and to an untrimmed baseline, across four datasets spanning RNA-Seq, SNP calling, and de novo genome assembly Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionno
Approaches that could also have been used
  • Performance metrics (e.g., % mapped reads, APOMAC, N50) are each reported as a single value per tool/threshold/dataset combination, without replicate runs.
    Could also: Generating replicate estimates, for example via bootstrap resampling of reads or repeated subsampling of the dataset, and reporting a mean with SD or a 95% CI. — Would add a measure of estimation uncertainty around each metric, letting readers gauge how much a given percentage or count might vary by chance.
  • Differences between trimming tools and thresholds are described narratively (e.g., tools that 'quickly reduce' APOMAC or achieve higher mapping rates) rather than assessed with a formal statistical test.
    Could also: A paired or repeated-measures approach (e.g., Wilcoxon signed-rank test or repeated-measures ANOVA) treating datasets or Q thresholds as matched units across tools. — Would provide a formal basis for stating whether an observed difference between tools exceeds what could arise from chance variation alone.
  • Nine tools are compared across many Q thresholds and four datasets without a stated correction for multiple comparisons.
    Could also: Applying a multiplicity correction such as Benjamini-Hochberg FDR or Bonferroni, if formal pairwise statistical tests among tools/thresholds were conducted. — Would control the familywise error rate or false discovery rate across the many implicit pairwise comparisons being made.
  • Mapping percentages and allele-call percentages are reported as proportions without an interval estimate.
    Could also: Reporting a proportion confidence interval (e.g., Wilson or Clopper-Pearson) alongside each percentage. — Would convey the precision of each percentage estimate given the underlying number of reads or nucleotides it is based on.
  • An 'optimal' Q threshold per tool is identified by visual inspection of trends in bar plots and scatter plots (e.g., Figure 2).
    Could also: A model-based optimization approach, such as fitting a curve to the threshold-performance relationship and selecting/validating the optimum via cross-validation. — Would provide an objective, reproducible criterion for the optimal threshold along with an estimate of confidence in that choice.
Software: FastQC · ABySS · BWA

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

untrimmed_mappability
Reported
72.189%
Reproduced
72.559% (6,606,456/9,104,944 mapped; Tophat 2.1.1/Bowtie2 2.5.4/hg19)
within tolerance
sickle_q20_mappability
Reported
95.971% (Q=20)
Reproduced
96.041% (2,707,338/2,818,935 mapped; sickle se -t sanger -q 20 -l 23 -g)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

Both in-scope Table 2 claims reproduce essentially 1:1 from the exact deposited accession SRR002073: untrimmed mappability 72.559% vs the reported 72.189% (+0.37 pp) and Sickle Q=20 mappability 96.041% vs the reported 95.971% (+0.07 pp). The residual sits entirely on the technical side — Tophat 2.1.1/Bowtie2 2.5.4 instead of Tophat 2.0.5 and a current hg19.refGene.gtf standing in for the paper's irretrievable 8-April-2013 snapshot — not on the authors' side and not in the core computation. The paper's qualitative claim of an optimal intermediate trimming threshold is independently confirmed by the Q10/20/30/40 sweep (93.4 / 96.0 / 91.7 / 70.4%), an inverted U peaking at Q20. Two caveats worth logging for statistics but not for downgrading: the Sickle invocation had to be reconstructed because the paper prints no command line, and the RU covers only 1 of 4 datasets and 1 of ~9 tools, so 'reproduced' here certifies two Table-2 cells rather than the paper as a whole.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.