Analysis and comprehensive comparison of PacBio and nanopore-based RNA sequencing of the Arabidopsis transcriptome.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL (described well enough; honest mixed 1:1). P16 third-party reproduction: ran the paper's named tool AlignQC v2.0.5 on its own deposited ONT PCR-cDNA data (ENA SRR10611193/94/95) via «our HPC» SLURM «job» (192 cores, all intermediates in /dev/shm to survive a chronically-full shared «infra» quota; only ~1KB of stats written to «infra»). Pipeline: seqtk -s11 100k reads reservoir-sampled by STREAMING each full replicate (curl|zcat|seqtk) -> pooled 300k -> GMAP -f samse -n 0 (chimera-aware) -> TAIR10 (Ensembl Plants r57) -> AlignQC analyze. RESULT: read-length distribution (Table 1) reproduces EXACTLY -- mean 1231.307 vs 1231.308, median 1097=1097, max 8236=8236 across 300k reads -- and is now CONFIRMED across two independent compute runs (file-based «job» and this streaming «job» gave the identical aligned-read set, 271554), strong evidence the deposit IS the analyzed data with no fabrication concern. IMPROVEMENT over the prior run: GMAP -n 0 RECOVERS chimera detection -> chimeric reads 1.009% (3028/300000) vs paper 0.51% (same order of magnitude); the prior -n 1 run's chimeric=0.00 was a pure aligner-flag artifact, now resolved to PARTIAL. Alignment fraction (90.5 vs 97.0%) and error partition (mism/del/ins) remain partial because the paper does NOT specify the GMAP version/parameters (we ran GMAP 2024-11-20 vs the paper's ~2019 build); overall error rate is the right ONT magnitude (14.59% vs 12.669%). NOT attempted: PacBio column (under-specified read set, heavy Iso-Seq), ONT Direct-cDNA + Illumina (not deposited in GSE141641).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 54assessed: 2026-06-20 ⛓ ad85ff79ca0d
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests how Oxford Nanopore Technologies (ONT) direct cDNA and PCR cDNA long-read sequencing compare to PacBio SMRT sequencing for full-length transcriptome characterization in Arabidopsis, asking whether ONT (particularly PCR-amplified cDNA, ONT Pc) can serve as a viable, cost-effective alternative to PacBio for plant transcriptome analysis.
- ★ ONT Pc produces higher raw data quality (higher alignment rate, lower error rate) than ONT Dc, while PacBio generates the longest reads finding
- ★ PacBio and ONT Pc perform similarly in transcript identification, SSR analysis, and lncRNA prediction finding
- ★ PacBio is superior to ONT Pc in detecting alternative splicing events finding
- ★ ONT Pc estimates transcript expression levels with much higher correlation to Illumina data than ONT Dc finding
- ★ ONT Pc is a cost-effective and worthwhile method for full-length single-molecule transcriptome analysis in plants finding
- ONT Dc shows a low alignment rate and elevated self-chimeric read rate compared to PacBio and ONT Pc finding
- Long-read sequencing (PacBio/ONT) captures full-length transcripts, overcoming the fragmentation limitations of short-read Illumina RNA-seq mechanism
- PCR/Sanger validation confirmed a higher proportion of ONT Pc-predicted lncRNAs than PacBio-predicted lncRNAs finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Illumina short-read RNA-seq | Arabidopsis (CTRL1-3) | none | gene expression (FPKM/CPM), mapping rate | Illumina NovaSeq |
| PacBio SMRT Iso-Seq (long-read cDNA sequencing) | Arabidopsis full-length cDNA library (1-6 kb) | none | read length, error rate, transcript/isoform identification, AS events, SSRs, lncRNAs | PacBio Sequel |
| Nanopore direct cDNA sequencing (ONT Dc) | Arabidopsis (CTRL1-3) | none | read length, error rate, mappability, transcript identification, expression correlation with Illumina | Nanopore GridION |
| Nanopore PCR cDNA sequencing (ONT Pc) | Arabidopsis (CTRL1-3) | none | read length, error rate, mappability, transcript identification, expression correlation with Illumina | Nanopore PromethION |
| SSR (simple sequence repeat) analysis | PacBio- and ONT-derived transcripts (>500 bp) | none | number and type of SSR motifs | MISA |
| CDS/ORF prediction and lncRNA prediction | PacBio, ONT Dc, ONT Pc transcripts | none | complete ORF counts, candidate lncRNAs | TransDecoder, CPC, CNCI, Pfam, CPAT |
| PCR amplification and Sanger sequencing | 16 randomly selected candidate lncRNAs (8 PacBio, 8 ONT Pc) | none | sequence identity/mismatches vs long-read sequences | — |
| Alternative splicing (AS) event detection | PacBio, ONT Dc, ONT Pc mapped reads | none | counts of IR, ES, Alt.5', Alt.3', mutually exclusive exon events | — |
- – Mean raw read length: PacBio 1410.19 bp, ONT Dc 902.06 bp, ONT Pc 1231.31 bp; PacBio had the longest maximum read (89,075 bp) 1410.19 vs 902.06 vs 1231.31 bp
- – Alignment rate to reference genome: PacBio 94.5%, ONT Dc 66.0%, ONT Pc 97.0% 94.5% / 66.0% / 97.0%
- ▼ Error rate: PacBio 13.217%, ONT Dc 13.934%, ONT Pc 12.669%, indicating ONT Pc had slightly higher base quality 12.669% vs 13.217% vs 13.934%
- – 13,967 known genes commonly identified by ONT Pc and PacBio; 2,542 specific to ONT Pc and 1,283 specific to PacBio 13,967 common genes
- ▼ PacBio detected 12,979 total AS events (IR most frequent, 62.99%) vs only 509 AS events common between PacBio and ONT Pc, showing ONT Pc's relative weakness in AS detection 509 common of 12,979 PacBio events
- ▼ SSR detection: 29,394 SSRs in PacBio vs 13,415 SSRs in ONT Pc, with 3,551 SSRs common to both 29,394 vs 13,415
- ▲ Expression correlation with Illumina was much higher for ONT Pc (r=0.932, 0.928, 0.923) than for ONT Dc (r=0.747, 0.719, 0.711) r~0.93 vs r~0.73
- ▲ Of 8 selected ONT Pc lncRNAs, 5 were validated by Sanger sequencing (fewer than 3 mismatches), compared to fewer confirmed matches among 8 PacBio lncRNAs 5/8 vs 2/8 exact match
- correlation r=0.932, 0.928, 0.923 (Illumina vs ONT Pc expression correlation, CTRL1-3)
- correlation r=0.747, 0.719, 0.711 (Illumina vs ONT Dc expression correlation, CTRL1-3)
- other 94.5% (PacBio), 66.0% (ONT Dc), 97.0% (ONT Pc) (read alignment rates to reference genome)
- other 13.217% (PacBio), 13.934% (ONT Dc), 12.669% (ONT Pc) (overall sequencing error rates)
- mean 1410.186 bp (PacBio), 902.0619 bp (ONT Dc), 1231.308 bp (ONT Pc) (mean raw read length)
- count 38,011 non-redundant mapped transcripts (PacBio), 47,601 (ONT Dc), 36,775 (ONT Pc) (identified transcripts per platform)
- count 29,394 SSRs (PacBio) vs 13,415 SSRs (ONT Pc); 3,551 common (SSR detection comparison)
- count 509 common AS events out of 12,979 PacBio AS events (overlap of alternative splicing events between PacBio and ONT Pc)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a descriptive benchmarking study comparing PacBio, ONT Direct cDNA (ONT Dc), and ONT PCR cDNA (ONT Pc) long-read sequencing platforms against Illumina short-read RNA-seq for Arabidopsis transcriptome analysis. The primary analytical approach is descriptive: read-length distributions, error-rate tabulations, and percentage-based comparisons of alignment, chimeric read, and AS-event metrics across platforms. Quantitative comparison of expression-level agreement between ONT and Illumina was performed using pairwise correlation coefficients for each of three biological replicates; no formal inferential hypothesis tests or p-values are reported.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pairwise correlation (type not specified; likely Pearson) | Transcript expression correlation between Illumina and ONT Dc / ONT Pc (Fig. 7a–f) | Three replicates per condition (CTRL1, CTRL2, CTRL3); number of transcripts per correlation not stated | not stated |
| Descriptive percentage-based comparison | Alignment rates, error rates, chimeric read rates, AS-event type fractions, SSR counts across platforms (Tables 1–3, Fig. 4) | Randomly sampled 10 Mb PacBio subreads (3,112,439) and 100,000 ONT 1D reads per sample (300,000 total) | na |
| Overlap enumeration via Venn diagram | Known genes shared between PacBio and ONT Pc (Fig. 4a); AS events (Fig. 4e); SSRs (Fig. 4h); lncRNAs (Additional Table S8) | — | na |
| Sanger sequencing validation (qualitative concordance check) | 16 randomly selected lncRNAs (8 PacBio, 8 ONT Pc) validated by PCR and Sanger sequencing (Fig. 6) | n=16 lncRNAs total | na |
-
Expression correlations between platforms were summarised as single Pearson-like r values per replicate (type of correlation not specified)↳ Could also: Spearman rank correlation could also be used, and the correlation type (Pearson vs Spearman) could be explicitly stated — RNA-seq expression data (FPKM/CPM) are typically right-skewed and heteroscedastic; Spearman correlation is rank-based and does not assume bivariate normality, which is often more appropriate for expression data. Stating the method allows readers to assess the assumption.
-
Variability across the three biological replicates was shown by reporting three separate r values (one per replicate) rather than a summary with uncertainty↳ Could also: A summary statistic (mean r ± SD across replicates, or a single model-based intraclass correlation) with a 95% bootstrap or Fisher-z confidence interval could also be reported — Reporting spread across replicates in a single summary with uncertainty communicates both the typical agreement and its consistency, making cross-platform comparisons easier to interpret quantitatively.
-
Platform comparisons of error rates, alignment rates, AS-event counts, and SSR counts were made descriptively by tabulating percentages↳ Could also: Formal statistical tests (e.g., chi-square or Fisher's exact test for proportions; permutation tests for count-based metrics across replicates) could also be applied where replicate-level data are available — Formal tests provide a principled way to quantify whether observed differences between platforms exceed what would be expected from sampling variability, and are particularly useful when differences are modest.
-
lncRNA prediction validation used a randomly selected subset of 16 lncRNAs (8 per platform) assessed by Sanger sequencing concordance↳ Could also: A larger or stratified random sample (e.g., stratified by expression level or length) with a reported confidence interval on the validation rate could also be used — With n=8 per group, the binomial confidence interval around the observed validation rate (e.g., 5/8) is wide (roughly 30–95%); a larger or stratified sample and a reported CI would better characterise the expected accuracy across the full set of predicted lncRNAs.
-
No sample-size or power justification was provided for the choice of three biological replicates per ONT condition↳ Could also: A brief power or simulation-based rationale for n=3 could also be included, referencing typical within-condition variability from prior long-read RNA-seq studies — Documenting the basis for replicate number helps readers understand the study's sensitivity for detecting platform-level differences in quantitative metrics such as expression correlation.
-
Transcriptome-level metrics (gene counts, transcript counts, SSR counts) were compared across platforms as absolute numbers or via Venn overlaps without uncertainty estimates↳ Could also: Bootstrap resampling of reads could also be used to estimate variability in these counts and produce confidence intervals around the overlap and platform-specific figures — Because each platform was run as a single library preparation with three sequencing replicates, bootstrapping provides a way to characterise how stable the reported counts are with respect to sequencing depth and stochastic sampling.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32536962
Paper: Cui J, Shen N, Lu Z, Xu G, Wang Y, Jin B. Analysis and comprehensive comparison of PacBio and nanopore-based RNA sequencing of the Arabidopsis transcriptome. Plant Methods 2020. PMID 32536962 / PMC7291481 / DOI 10.1186/s13007-020-00629-x.
Tool (P16, third-party): AlignQC (https://github.com/jason-weirather/AlignQC,
v2.0.5, commit 2b471c2). The paper explicitly used "AlignQC software" to QC the
long-read alignments. Per the brief, applying this existing third-party tool to
the paper's own data is a fully valid reproduction.
Data: GEO GSE141641 / BioProject PRJNA594286 (ENA). Resolved runs:
| Run | Experiment | Platform | Library | reads (ENA) | bases |
|---|---|---|---|---|---|
| SRR10611192 | SRX7290464 | PacBio Sequel | Iso-Seq (subreads) | 21,531,816 | 43.95 Gb |
| SRR10611193 | SRX7290465 | ONT PromethION | PCR cDNA (rep CTRL1) | 8,258,511 | 10.09 Gb |
| SRR10611194 | SRX7290466 | ONT PromethION | PCR cDNA (rep CTRL2) | 7,849,574 | 9.78 Gb |
| SRR10611195 | SRX7290467 | ONT PromethION | PCR cDNA (rep CTRL3) | 7,025,983 | 8.61 Gb |
IN SCOPE (pipeline-derived, attempted)
AlignQC alignment-QC metrics for ONT PCR cDNA (paper Tables 1–3, "ONT Pc"
column). These are direct AlignQC outputs (error_stats.txt, alignment_stats.txt):
- error rate (%) — reported 12.669
- mismatch rate (%) — reported 4.352
- deletion rate (%) — reported 5.085
- insertion rate (%) — reported 3.232
- aligned reads (%) — reported 97.0
- mean / median / max read length (bp) — reported 1231.31 / 1097 / 8236
Pipeline per result: ENA FASTQ → (subsample) → GMAP align to TAIR10 (paper's aligner) → sorted BAM → AlignQC analyze → stat files.
Robustness note (why these are the 80%)
Per the paper's Methods (verified from PMC7291481 full text): AlignQC could not
hold the full data in 500 GB RAM, so the authors "randomly selected ... 100,000
1D reads from each ONT sample" → 100k × 3 ONT PCR-cDNA replicates
(SRR10611193/94/95) = 300,000 reads fed to AlignQC. Our job mirrors this exactly
(seqtk -s11, 100k from each of the 3 runs, pooled). We cannot recover the authors'
exact random selection, but:
- Per-base rates (error/mismatch/deletion/insertion) are unbiased over any random subsample → directly comparable (the strong 1:1 targets).
- mean/median read length of a random 300k subsample ≈ full-run value → comparable within tolerance.
- max read length is subsample-size dependent → weak, reported with caveat.
- aligned % is subsample-robust but aligner-sensitive (GMAP vs our GMAP run with documented params) → reported, compared honestly.
SECONDARY / OPTIONAL (the hard 20% — may skip)
- PacBio AlignQC column (Tables 1–3): the paper's "3,112,439 reads" does not match ENA subreads (21.5M), ROIs (516,364) or FLNC (416,662) — the exact read set fed to AlignQC is under-specified, and full PacBio Iso-Seq processing (CCS→classify→ICE cluster) is heavy. Attempt only if the ONT path completes with budget to spare; otherwise documented as not-attempted.
OUT OF SCOPE (not attempted, with reason)
- ONT Direct cDNA (Tables: "ONT Dc"): no data deposited in GSE141641 (only
PacBio + ONT PCR cDNA runs present) →
data_unavailablefor this subset. - Illumina correlation (Fig 7): Illumina NovaSeq data not in this series.
- Downstream biology — alternative-splicing event counts, SSR counts, lncRNA prediction, gene/transcript detection counts (Figs 4–5): multi-step downstream analyses with under-specified parameters and external tools (cDNA_Cupcake, ICE, SSR/lncRNA predictors). These are the hard 20%; not attempted.
Reference
TAIR10 genome from Ensembl Plants:
Arabidopsis_thaliana.TAIR10.dna.toplevel.fa (chr 1–5 + Mt + Pt).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.