Viral Diagnostics in Plants Using Next Generation Sequencing: Computational Analysis in Practice.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡Reported values were only indirectly comparable
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PMID 29123534 is a review/tutorial; the only pipeline-derived results that are the authors' OWN are their hands-on benchmark of public virus-detection tools on 3 public SRA datasets (Table 2). Described well enough to attempt: yes, for the open tools. CLEAN 1:1 reproduced (S1, exact): the grapevine sRNA-seq dataset the paper benchmarked VirusDetect on is 713.4M bases — ENA base_count for SRR3680863 = 713,356,593, matching exactly (immutable SRA fact, compute-independent). PRIMARY tool reproduction (S2 = VirusDetect on grapevine SRR3680863, an open third-party tool on the paper's own data) was authored and submitted to «our HPC» («job») but FAILED ~40s into the conda env-build, mid package download, BEFORE the pipeline executed — a transient HPC infra failure, NOT a statement about the paper (the env solve resolved every pin cleanly; a resubmit with the now-warm «infra» pkgs cache would likely proceed). Net: partial — one data-scale claim reproduced exact, the tool-output comparison incomplete. NOT attempted (the ~20%): Taxonomer on pear/pepper (closed web service, classifier+DB version unpinnable), VirFind (web service), VSD/Yabi (github.com/muccg/yabi is a generic workflow web-portal engine, not the analysis), Metavisitor (Galaxy suite). No fabrication concerns; S1 matches SRA exactly.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-16 ⛓ 5b777fdb3987
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan RNA-sequencing (NGS) enable unbiased, hypothesis-free detection and identification of multiple known and emergent viruses in crop plants, and what bioinformatics methods are required to do so? This review surveys studies and computational workflows addressing the question 'how many different viruses are present in this crop plant?' without prior knowledge of the targets.
- ★ NGS/RNA-seq enables unbiased, hypothesis-free detection of multiple known and emergent plant viruses, unlike RT-PCR which only detects one or a few known viruses per test. finding
- ★ Virus detection from RNA-seq requires a bioinformatics workflow comprising quality control, de novo assembly into contigs, host sequence removal by alignment to host genome, and viral read identification by mapping to virus databases. method
- ★ Co-infection of individual plants with multiple viruses is a consistent theme across crop studies, justifying multiplexed detection methods. finding
- ★ Sequencing of total small RNAs (sRNAs/virus-derived siRNAs enriched via host RNA interference antiviral immunity) is an effective alternative method for plant virus detection. method
- ★ Different assembly tools and analysis pipelines (e.g., Trinity vs Velvet) yield different sets of detected viruses, affecting which viruses are identified. finding
- ★ RNA-seq re-analysis can detect novel viruses and novel virus isolates, and identify viruses in asymptomatic plants or in new geographic regions/host species. finding
- ★ Numerous specialized bioinformatics tools/workflows exist (VirFind, Taxonomer, VSD toolkit, Metavisitor, VIP, ViromeScan, VirusHunter) but most focus on human clinical samples; plant virus detection is harder due to incomplete crop genomes and poor representation of plant virus sequences in databases. resource
- Future directions include deploying virus-detection bioinformatics tools in analytical environments using cloud computing. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA-seq (total RNA, de novo assembly) | Garlic (Allium sativum, cultivated and imported) and wild garlic (A. vineale), leaves, Australia | none | viral contigs/isolates identified via Blastn/Blastx against GenBank; contig coverage | Illumina HiSeq 2000; Geneious Pro and CLC Genomics Workbench |
| RNA-seq (de novo assembly) | Pepper (Capsicum annuum), cultivars Pusa Jwala (susceptible) and Taiwan-2 (resistant) | none (cultivar comparison) | viral contigs matched via MEGABLAST to RefSeq viral database; Trinity vs Velvet assembly comparison | Illumina HiSeq 2000; Trinity, Velvet+Oases |
| mRNA-seq and sRNA-seq transcriptome re-analysis | Pear (Pyrus pyrifolia), public SRA transcriptome (SRX532394), different developmental stages | none | viral contigs with read counts >5 via MEGABLAST against reference viral genomes | Trinity (assembly); SRA-sourced data |
| RNA-seq (paired-end, de novo assembly) | Grapevine (Vitis vinifera), lignified cane from Merlot vineyard, South Africa | none | viral contigs via Blast against NCBI nr DNA/protein databases | Illumina Genome Analyzer; Velvet |
| Transcriptome re-analysis (paired-end RNA-seq, de novo assembly) | Grapevine (Vitis vinifera) cultivar Tannat, libraries from grain, skin and seed tissues | none (tissue comparison) | viral prevalence per tissue via Blast against virus reference genomes | Illumina HiSeq 1000; Trinity |
| sRNA-seq (small RNA sequencing, de novo assembly) | Sweet potato (Ipomoea batatas), leaf material, Honduras and Guatemala | none | virus contigs via Blast against NCBI nr; short-read alignment with MAQ; co-infection vs symptom severity | Illumina Genome Analyzer; Velvet, MAQ |
| RNA-seq (de novo assembly) | Orange (Citrus sinensis), CSD-symptomatic and -asymptomatic trees, Sao Paulo, Brazil | disease state (symptomatic vs asymptomatic) | virus species/isolates/genotypes via Blastx (nr protein) and BLASTn (nt); genotype-symptom association | Illumina HiSeq 2000; CLC Assembly Cell, Trinity |
| sRNA-seq (small RNA sequencing, host alignment + de novo assembly) | Tomato (Solanum lycopersicum), symptomatic plants, US and Mexico | none | candidate virus genomes via BWA alignment to host then GenBank virus collection, Velvet assembly, BLAST vs nt/nr | Illumina Genome Analyzer II; BWA, Velvet |
- – Between 1 and 8 virus isolates were present in each cultivated garlic plant, and a single virus isolate in one wild garlic plant; 41 virus isolates identified in total (potyviruses, allexiviruses, carlaviruses); first complete genomes of two GarVD isolates and first detection of Asparagus virus 3 in wild garlic in Australia. 1-8 viruses per plant; 41 isolates total
- – Pepper study identified eight viruses common to all datasets (incl. BPEV, PepLCBV, TVCV with highest contig counts) and a novel virus Pepper Virus A (PepVA); Trinity produced longer contigs while Velvet was better for low virus titre. 8 common viruses
- – Pear transcriptome revealed 5 viruses with read counts >5 (ASGV, PrVT, AGCAV, ASPV); reads initially matched to Potato leaf roll virus were re-identified as host sequences. 5 viruses (read counts >5)
- – Grapevine Merlot study identified GLRaV-3, GRSPaV, GVA and Grapevine virus E (first report in South African vineyards) and was first to isolate mycoviruses in grapevine phloem.
- – Tannat grapevine re-analysis: most prevalent viruses were GYSVd1, GPGV, HSVd, GLRaV2; 4 viruses found in all 3 tissues while OBDV and PVS only in seed; skin tissue had higher virus prevalence than grain. 4 of viruses in all 3 tissues
- – Sweet potato sRNA-seq simultaneously detected three RNA viruses (SPCSV-WA, SPFMV-RC, SPVC) and two DNA viruses (SPLCGV, SPPV-B); severe SPPV-B symptoms were always accompanied by SPCSV-WA. 3 RNA + 2 DNA viruses
- – Citrus study found mixed infections (CTV, CSDaV, CitPRV) plus two putative novel viruses (CJLV, CVLV), differentiated two genotypes each for CTV and CSDaV, and associated one CSDaV genotype with symptomatic plants. 4 million orange trees lost to CSD
- – Tomato sRNA-seq assembled complete genomes of six PepMV isolates and a PSTVd isolate, differentially assembled two co-infecting PepMV strains (EU and US1), and detected and fully assembled a novel potyvirus. 6 PepMV isolates
- count 41 virus isolates identified (Total virus isolates found across cultivated and wild garlic plants)
- count 1 to 8 viruses per cultivated garlic plant (Range of virus isolates per individual A. sativum plant)
- other contig coverage <10-fold removed (Threshold for filtering putative viral contigs in garlic study)
- count 8 viruses common to all datasets (Viruses detected across all pepper assemblies)
- count 5 viruses with read counts >5 (Viruses detected in pear transcriptome)
- count more than 30 viruses (Number of viruses known to infect sweet potato)
- count over 4 million orange trees lost (Citrus sudden death disease losses in Sao Paulo State, Brazil)
- count 6 PepMV isolates and 1 PSTVd isolate (Complete genomes assembled from tomato sRNA-seq)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a narrative review article surveying computational and bioinformatics workflows for RNA-seq-based viral detection in crop plants. It qualitatively summarizes the methods, tools, and findings of previously published studies and does not itself collect data, define experimental groups, or apply inferential statistical tests; results are reported descriptively (e.g., counts of viruses identified, contig coverage thresholds).
-
The review synthesizes the literature qualitatively, describing each study in turn.↳ Could also: A structured or systematic review with predefined inclusion criteria and a tabulated cross-study comparison (e.g., PRISMA-style flow) could also have been used. — A systematic approach would add reproducibility and transparency about how studies were selected, complementing the narrative synthesis.
-
Bioinformatics tools are compared narratively in a summary table (Table 1).↳ Could also: A common benchmarking framework applying each tool to a shared reference dataset with standardized performance metrics (sensitivity, precision, F1) could also have been used. — Quantitative benchmarking on common data would allow direct, like-for-like comparison of tool performance across the surveyed workflows.
-
Virus detection in cited studies relied on coverage/length thresholds (e.g., >10-fold coverage, contigs >1,000 nt, read counts >5).↳ Could also: Reporting threshold sensitivity analyses or statistical confidence measures for calls could also accompany such cutoffs. — Sensitivity analysis would convey how robust detection counts are to the chosen thresholds, which is informative when comparing across datasets.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
SPCSV-WA co-infection consistently accompanied severe SPPV-B symptoms in Ipomoea batatas; five viruses (SPCSV-WA, SPFMV-RC, SPVC, SPLCGV, SPPV-B) simultaneously detected by sRNA-seq.other ipomoea-batatas-leaf 2017×1papers★ This paper is the founder (earliest)
-
Complete genomes of six PepMV isolates and a novel potyvirus assembled from symptomatic Solanum lycopersicum plants by sRNA-seq; co-infecting PepMV EU and US1 strains differentially assembled.other solanum-lycopersicum 2017×1papers★ This paper is the founder (earliest)
-
First complete genome sequences of GarVD isolates obtained from cultivated garlic (Allium sativum) in Australia; 41 virus isolates (potyviruses, allexiviruses, carlaviruses) identified across plants; Asparagus virus 3 first detected in wild garlic (A. vineale).RNA-seq allium-sativum-leaf 2017×1papers★ This paper is the founder (earliest)
-
Pepper Virus A (PepVA) identified as a novel virus in Capsicum annuum by de novo RNA-seq assembly across susceptible and resistant cultivar datasets.RNA-seq capsicum-annuum 2017×1papers★ This paper is the founder (earliest)
-
One CSDaV genotype associated with CSD-symptomatic Citrus sinensis trees; two putative novel viruses (CJLV, CVLV) identified alongside mixed CTV and CitPRV infections.RNA-seq citrus-sinensis 2017×1papers★ This paper is the founder (earliest)
-
Apple stem grooving virus (ASGV), PrVT, AGCAV, and ASPV detected in a public pear (Pyrus pyrifolia) transcriptome by MEGABLAST re-analysis; an apparent Potato leaf roll virus match was re-attributed to host sequences.RNA-seq pyrus-pyrifolia 2017×1papers★ This paper is the founder (earliest)
-
Grapevine virus E (GVE) reported for the first time in South African Vitis vinifera vineyards; GLRaV-3, GRSPaV, and GVA also detected; mycoviruses first isolated from grapevine phloem.RNA-seq vitis-vinifera 2017×1papers★ This paper is the founder (earliest)
-
GYSVd1, GPGV, HSVd, and GLRaV-2 detected across all three Tannat grapevine tissues (grain, skin, seed); skin tissue showed higher virus prevalence than grain; OBDV and PVS restricted to seed.RNA-seq vitis-vinifera 2017×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 29123534
Paper: Jones S, Baizan-Edge A, MacFarlane S, Torrance L. (2017) Viral Diagnostics in Plants Using Next Generation Sequencing: Computational Analysis in Practice. Front Plant Sci 8:1770. DOI 10.3389/fpls.2017.01770.
Nature of the paper
This is a review / "in practice" tutorial, not a primary-data study. Most of the numbers it cites (garlic 41 virus isolates, pepper novel virus, tomato 6 PepMV genomes, etc.) are quoted from other groups' published studies — those are out of scope: they are not results this paper's authors computed, and the inputs/parameters are in the cited papers, not here.
The authors' own computational contribution is a hands-on benchmark: they took three public SRA datasets (their Table 2) and ran them through several publicly available virus-detection tools (Table 1: VirFind, VirusDetect, Taxonomer, VSD/Yabi, Metavisitor), reporting how each performed. Those benchmark runs are the pipeline-derived results in scope.
Table 2 — datasets the authors re-analysed (verified against ENA)
| Organism | Paper SRA | Run (ENA) | Type | reads | bases | orig-study viruses |
|---|---|---|---|---|---|---|
| Pear (Pyrus pyrifolia) | SRR1269627 (= exp SRX532394) | SRR1269627 | RNA-seq SE | 97,896,223 | 3,524,264,028 | ASGV, AGCAV, ASPV, PrVT |
| Pepper (Capsicum annuum) | SRR1123893 | SRR1123893 | RNA-seq PE | 53,921,012 | 10,784,202,400 | 13 viruses (ALPV, BPEV, … TVCV) |
| Grapevine (Vitis vinifera) | SRR3680863 | SRR3680863 | sRNA-seq SE | 32,497,945 | 713,356,593 | GRSPaV, GVB, GFkV, GLRaV-3, HSVd |
Note: the RU brief lists the data accession as SRX532394; that is the SRA experiment for the pear run SRR1269627 (Table 2). Both resolve to the same pear RNA-seq library.
IN SCOPE (pipeline-derived, this paper's own runs)
- S1 — dataset facts. The reads/bases of the three Table-2 datasets are immutable SRA facts and the paper quotes one directly: "VirusDetect was tested on sRNA-seq data from grapevine (713.4M bases)." → ENA base_count for SRR3680863 = 713,356,593 ≈ 713.4 M. Verifiable with zero ambiguity (control-plane lookup).
- S2 — VirusDetect on grapevine sRNA-seq (SRR3680863). Paper:
"VirusDetect identified 11 virus isolates in the grape sRNA-seq data, 5 of
which were identified in the original study" and "After file upload it gave
results in < 4 h." Pipeline: VirusDetect (kentnf/VirusDetect; sRNA clean →
Velvet de-novo + reference-guided assembly → megablast/blastx vs
vrl_plantcurated plant-virus DB). This is open-source and runnable on the paper's exact data → chosen primary reproduction target. Metric we compare: number of reported virus isolates, and recovery of the 5 original-study viruses (GRSPaV, GVB, GFkV, GLRaV-3, HSVd).
IN SCOPE BUT NOT ATTEMPTED (the hard ~20%, with reasons)
- Taxonomer on pear / pepper ("8% of pear reads classified", "5,707 virus
reads pear / 364,959 pepper", "8 of 13 pepper viruses"). Taxonomer is a
closed web service (taxonomer.com) requiring an account and a remote
classifier whose database/version cannot be pinned → not reproducibly runnable;
skipped per 80/20 (
env_unresolvablefor that sub-result). - VirFind / VSD-Yabi / Metavisitor benchmark mentions: Yabi (github.com/ muccg/yabi) is a generic workflow web-portal engine, not the analysis itself; VirFind is a web service; Metavisitor is a Galaxy toolshed suite. Out of 80/20 budget; VirusDetect (S2) already provides one clean third-party-tool 1:1.
OUT OF SCOPE (not this paper's computation)
- All per-crop result numbers quoted from the reviewed primary studies (garlic Wylie 2014, tomato Li 2012, sweet-potato Kashif 2012, orange Matsumura 2017, etc.). These are literature citations, reproducing them would mean reproducing those papers.
- Cost/£ statements, wet-lab/library-prep, and general methodological commentary.
Caveat on the comparison (DB drift)
V
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Only one in-scope claim could be checked, and it reproduced exactly: the grapevine sRNA-seq input is 713,356,593 bases (= the paper's 713.4M) for SRR3680863, an immutable SRA fact. The primary scientific reproduction — VirusDetect's 11 isolates / 5-of-11 recovery on the same open data — never ran: «our HPC» «job» failed during the conda env-build (transient infra on our side), so the core benchmark claim is untested, not contradicted. No authors'-side or fabrication concern arises; the isolate count is additionally DB-version sensitive and only indirectly comparable. Net: solid 1:1 on the data-fact but incomplete on the central claim → overall yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.