Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Viral Diagnostics in Plants Using Next Generation Sequencing: Computational Analysis in Practice.

Front Plant Sci · 2017
L1 100/100 PQI 94
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PMID 29123534 is a review/tutorial; the only pipeline-derived results that are the authors' OWN are their hands-on benchmark of public virus-detection tools on 3 public SRA datasets (Table 2). Described well enough to attempt: yes, for the open tools. CLEAN 1:1 reproduced (S1, exact): the grapevine sRNA-seq dataset the paper benchmarked VirusDetect on is 713.4M bases — ENA base_count for SRR3680863 = 713,356,593, matching exactly (immutable SRA fact, compute-independent). PRIMARY tool reproduction (S2 = VirusDetect on grapevine SRR3680863, an open third-party tool on the paper's own data) was authored and submitted to «our HPC» («job») but FAILED ~40s into the conda env-build, mid package download, BEFORE the pipeline executed — a transient HPC infra failure, NOT a statement about the paper (the env solve resolved every pin cleanly; a resubmit with the now-warm «infra» pkgs cache would likely proceed). Net: partial — one data-scale claim reproduced exact, the tool-output comparison incomplete. NOT attempted (the ~20%): Taxonomer on pear/pepper (closed web service, classifier+DB version unpinnable), VirFind (web service), VSD/Yabi (github.com/muccg/yabi is a generic workflow web-portal engine, not the analysis), Metavisitor (Galaxy suite). No fabrication concerns; S1 matches SRA exactly.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-16 ⛓ 5b777fdb3987
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can RNA-sequencing (NGS) enable unbiased, hypothesis-free detection and identification of multiple known and emergent viruses in crop plants, and what bioinformatics methods are required to do so? This review surveys studies and computational workflows addressing the question 'how many different viruses are present in this crop plant?' without prior knowledge of the targets.

Core claims
  • NGS/RNA-seq enables unbiased, hypothesis-free detection of multiple known and emergent plant viruses, unlike RT-PCR which only detects one or a few known viruses per test. finding
  • Virus detection from RNA-seq requires a bioinformatics workflow comprising quality control, de novo assembly into contigs, host sequence removal by alignment to host genome, and viral read identification by mapping to virus databases. method
  • Co-infection of individual plants with multiple viruses is a consistent theme across crop studies, justifying multiplexed detection methods. finding
  • Sequencing of total small RNAs (sRNAs/virus-derived siRNAs enriched via host RNA interference antiviral immunity) is an effective alternative method for plant virus detection. method
  • Different assembly tools and analysis pipelines (e.g., Trinity vs Velvet) yield different sets of detected viruses, affecting which viruses are identified. finding
  • RNA-seq re-analysis can detect novel viruses and novel virus isolates, and identify viruses in asymptomatic plants or in new geographic regions/host species. finding
  • Numerous specialized bioinformatics tools/workflows exist (VirFind, Taxonomer, VSD toolkit, Metavisitor, VIP, ViromeScan, VirusHunter) but most focus on human clinical samples; plant virus detection is harder due to incomplete crop genomes and poor representation of plant virus sequences in databases. resource
  • Future directions include deploying virus-detection bioinformatics tools in analytical environments using cloud computing. finding
Experimental setups
Assay System Perturbation Readout Platform
RNA-seq (total RNA, de novo assembly) Garlic (Allium sativum, cultivated and imported) and wild garlic (A. vineale), leaves, Australia none viral contigs/isolates identified via Blastn/Blastx against GenBank; contig coverage Illumina HiSeq 2000; Geneious Pro and CLC Genomics Workbench
RNA-seq (de novo assembly) Pepper (Capsicum annuum), cultivars Pusa Jwala (susceptible) and Taiwan-2 (resistant) none (cultivar comparison) viral contigs matched via MEGABLAST to RefSeq viral database; Trinity vs Velvet assembly comparison Illumina HiSeq 2000; Trinity, Velvet+Oases
mRNA-seq and sRNA-seq transcriptome re-analysis Pear (Pyrus pyrifolia), public SRA transcriptome (SRX532394), different developmental stages none viral contigs with read counts >5 via MEGABLAST against reference viral genomes Trinity (assembly); SRA-sourced data
RNA-seq (paired-end, de novo assembly) Grapevine (Vitis vinifera), lignified cane from Merlot vineyard, South Africa none viral contigs via Blast against NCBI nr DNA/protein databases Illumina Genome Analyzer; Velvet
Transcriptome re-analysis (paired-end RNA-seq, de novo assembly) Grapevine (Vitis vinifera) cultivar Tannat, libraries from grain, skin and seed tissues none (tissue comparison) viral prevalence per tissue via Blast against virus reference genomes Illumina HiSeq 1000; Trinity
sRNA-seq (small RNA sequencing, de novo assembly) Sweet potato (Ipomoea batatas), leaf material, Honduras and Guatemala none virus contigs via Blast against NCBI nr; short-read alignment with MAQ; co-infection vs symptom severity Illumina Genome Analyzer; Velvet, MAQ
RNA-seq (de novo assembly) Orange (Citrus sinensis), CSD-symptomatic and -asymptomatic trees, Sao Paulo, Brazil disease state (symptomatic vs asymptomatic) virus species/isolates/genotypes via Blastx (nr protein) and BLASTn (nt); genotype-symptom association Illumina HiSeq 2000; CLC Assembly Cell, Trinity
sRNA-seq (small RNA sequencing, host alignment + de novo assembly) Tomato (Solanum lycopersicum), symptomatic plants, US and Mexico none candidate virus genomes via BWA alignment to host then GenBank virus collection, Velvet assembly, BLAST vs nt/nr Illumina Genome Analyzer II; BWA, Velvet
Key results
  • Between 1 and 8 virus isolates were present in each cultivated garlic plant, and a single virus isolate in one wild garlic plant; 41 virus isolates identified in total (potyviruses, allexiviruses, carlaviruses); first complete genomes of two GarVD isolates and first detection of Asparagus virus 3 in wild garlic in Australia. 1-8 viruses per plant; 41 isolates total
  • Pepper study identified eight viruses common to all datasets (incl. BPEV, PepLCBV, TVCV with highest contig counts) and a novel virus Pepper Virus A (PepVA); Trinity produced longer contigs while Velvet was better for low virus titre. 8 common viruses
  • Pear transcriptome revealed 5 viruses with read counts >5 (ASGV, PrVT, AGCAV, ASPV); reads initially matched to Potato leaf roll virus were re-identified as host sequences. 5 viruses (read counts >5)
  • Grapevine Merlot study identified GLRaV-3, GRSPaV, GVA and Grapevine virus E (first report in South African vineyards) and was first to isolate mycoviruses in grapevine phloem.
  • Tannat grapevine re-analysis: most prevalent viruses were GYSVd1, GPGV, HSVd, GLRaV2; 4 viruses found in all 3 tissues while OBDV and PVS only in seed; skin tissue had higher virus prevalence than grain. 4 of viruses in all 3 tissues
  • Sweet potato sRNA-seq simultaneously detected three RNA viruses (SPCSV-WA, SPFMV-RC, SPVC) and two DNA viruses (SPLCGV, SPPV-B); severe SPPV-B symptoms were always accompanied by SPCSV-WA. 3 RNA + 2 DNA viruses
  • Citrus study found mixed infections (CTV, CSDaV, CitPRV) plus two putative novel viruses (CJLV, CVLV), differentiated two genotypes each for CTV and CSDaV, and associated one CSDaV genotype with symptomatic plants. 4 million orange trees lost to CSD
  • Tomato sRNA-seq assembled complete genomes of six PepMV isolates and a PSTVd isolate, differentially assembled two co-infecting PepMV strains (EU and US1), and detected and fully assembled a novel potyvirus. 6 PepMV isolates
Key statistics
  • count 41 virus isolates identified (Total virus isolates found across cultivated and wild garlic plants)
  • count 1 to 8 viruses per cultivated garlic plant (Range of virus isolates per individual A. sativum plant)
  • other contig coverage <10-fold removed (Threshold for filtering putative viral contigs in garlic study)
  • count 8 viruses common to all datasets (Viruses detected across all pepper assemblies)
  • count 5 viruses with read counts >5 (Viruses detected in pear transcriptome)
  • count more than 30 viruses (Number of viruses known to infect sweet potato)
  • count over 4 million orange trees lost (Citrus sudden death disease losses in Sao Paulo State, Brazil)
  • count 6 PepMV isolates and 1 PSTVd isolate (Complete genomes assembled from tomato sRNA-seq)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a narrative review article surveying computational and bioinformatics workflows for RNA-seq-based viral detection in crop plants. It qualitatively summarizes the methods, tools, and findings of previously published studies and does not itself collect data, define experimental groups, or apply inferential statistical tests; results are reported descriptively (e.g., counts of viruses identified, contig coverage thresholds).

Replicationna Sample sizeNot applicable as a primary study; the review enumerates the sample basis of cited studies (e.g., a benchmarking set of 38 plant samples from 19 species for VirFind, 21 plant genomes for the VSD toolkit), but no sample size or power analysis is described for the review itself. GroupsTools/workflows and crop case studies compared narratively Pairingna Randomization/blindingna Dispersionnone
Approaches that could also have been used
  • The review synthesizes the literature qualitatively, describing each study in turn.
    Could also: A structured or systematic review with predefined inclusion criteria and a tabulated cross-study comparison (e.g., PRISMA-style flow) could also have been used. — A systematic approach would add reproducibility and transparency about how studies were selected, complementing the narrative synthesis.
  • Bioinformatics tools are compared narratively in a summary table (Table 1).
    Could also: A common benchmarking framework applying each tool to a shared reference dataset with standardized performance metrics (sensitivity, precision, F1) could also have been used. — Quantitative benchmarking on common data would allow direct, like-for-like comparison of tool performance across the surveyed workflows.
  • Virus detection in cited studies relied on coverage/length thresholds (e.g., >10-fold coverage, contigs >1,000 nt, read counts >5).
    Could also: Reporting threshold sensitivity analyses or statistical confidence measures for calls could also accompany such cutoffs. — Sensitivity analysis would convey how robust detection counts are to the chosen thresholds, which is informative when comparing across datasets.
Software: Geneious Pro (de novo assembly, cited study) · CLC Genomics Workbench / CLC Assembly Cell (de novo assembly, cited studies) · Trinity (de novo assembly, cited studies) · Velvet (with Oases) (de novo assembly, cited studies) · SPAdes (de novo assembly, VSD toolkit) · BLAST / Blastn / Blastx / MEGABLAST (sequence comparison, cited studies) · Bowtie2, BWA, MAQ (read mapping, cited studies/tools)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
110
Impact: high
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 75/100
stands on reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

SRX532394 ENA in Introduction (http://purl.org/orb/Introduction)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 29123534

Paper: Jones S, Baizan-Edge A, MacFarlane S, Torrance L. (2017) Viral Diagnostics in Plants Using Next Generation Sequencing: Computational Analysis in Practice. Front Plant Sci 8:1770. DOI 10.3389/fpls.2017.01770.

Nature of the paper

This is a review / "in practice" tutorial, not a primary-data study. Most of the numbers it cites (garlic 41 virus isolates, pepper novel virus, tomato 6 PepMV genomes, etc.) are quoted from other groups' published studies — those are out of scope: they are not results this paper's authors computed, and the inputs/parameters are in the cited papers, not here.

The authors' own computational contribution is a hands-on benchmark: they took three public SRA datasets (their Table 2) and ran them through several publicly available virus-detection tools (Table 1: VirFind, VirusDetect, Taxonomer, VSD/Yabi, Metavisitor), reporting how each performed. Those benchmark runs are the pipeline-derived results in scope.

Table 2 — datasets the authors re-analysed (verified against ENA)

Organism Paper SRA Run (ENA) Type reads bases orig-study viruses
Pear (Pyrus pyrifolia) SRR1269627 (= exp SRX532394) SRR1269627 RNA-seq SE 97,896,223 3,524,264,028 ASGV, AGCAV, ASPV, PrVT
Pepper (Capsicum annuum) SRR1123893 SRR1123893 RNA-seq PE 53,921,012 10,784,202,400 13 viruses (ALPV, BPEV, … TVCV)
Grapevine (Vitis vinifera) SRR3680863 SRR3680863 sRNA-seq SE 32,497,945 713,356,593 GRSPaV, GVB, GFkV, GLRaV-3, HSVd

Note: the RU brief lists the data accession as SRX532394; that is the SRA experiment for the pear run SRR1269627 (Table 2). Both resolve to the same pear RNA-seq library.

IN SCOPE (pipeline-derived, this paper's own runs)

  • S1 — dataset facts. The reads/bases of the three Table-2 datasets are immutable SRA facts and the paper quotes one directly: "VirusDetect was tested on sRNA-seq data from grapevine (713.4M bases)." → ENA base_count for SRR3680863 = 713,356,593 ≈ 713.4 M. Verifiable with zero ambiguity (control-plane lookup).
  • S2 — VirusDetect on grapevine sRNA-seq (SRR3680863). Paper: "VirusDetect identified 11 virus isolates in the grape sRNA-seq data, 5 of which were identified in the original study" and "After file upload it gave results in < 4 h." Pipeline: VirusDetect (kentnf/VirusDetect; sRNA clean → Velvet de-novo + reference-guided assembly → megablast/blastx vs vrl_plant curated plant-virus DB). This is open-source and runnable on the paper's exact data → chosen primary reproduction target. Metric we compare: number of reported virus isolates, and recovery of the 5 original-study viruses (GRSPaV, GVB, GFkV, GLRaV-3, HSVd).

IN SCOPE BUT NOT ATTEMPTED (the hard ~20%, with reasons)

  • Taxonomer on pear / pepper ("8% of pear reads classified", "5,707 virus reads pear / 364,959 pepper", "8 of 13 pepper viruses"). Taxonomer is a closed web service (taxonomer.com) requiring an account and a remote classifier whose database/version cannot be pinned → not reproducibly runnable; skipped per 80/20 (env_unresolvable for that sub-result).
  • VirFind / VSD-Yabi / Metavisitor benchmark mentions: Yabi (github.com/ muccg/yabi) is a generic workflow web-portal engine, not the analysis itself; VirFind is a web service; Metavisitor is a Galaxy toolshed suite. Out of 80/20 budget; VirusDetect (S2) already provides one clean third-party-tool 1:1.

OUT OF SCOPE (not this paper's computation)

  • All per-crop result numbers quoted from the reviewed primary studies (garlic Wylie 2014, tomato Li 2012, sweet-potato Kashif 2012, orange Matsumura 2017, etc.). These are literature citations, reproducing them would mean reproducing those papers.
  • Cost/£ statements, wet-lab/library-prep, and general methodological commentary.

Caveat on the comparison (DB drift)

V

S1
Reported
713.4M bases (grapevine sRNA-seq, VirusDetect benchmark input)
Reproduced
713356593 bases for SRR3680863 (=713.4M)
exact
S2a
Reported
VirusDetect identified 11 virus isolates in the grape sRNA-seq data
Reproduced
not produced — «our HPC» «job» failed at ~40s during conda env-build, before pipeline ran (transient infra, not a paper finding)
m.public.grade.error
S2b
Reported
5 of the 11 isolates were also found in the original study (GRSPaV, GVB, GFkV, GLRaV-3, HSVd)
Reproduced
not produced (same env-build failure)
m.public.grade.error
S2c
Reported
VirusDetect returned results in < 4 h
Reproduced
not produced (same env-build failure)
m.public.grade.error

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

Only one in-scope claim could be checked, and it reproduced exactly: the grapevine sRNA-seq input is 713,356,593 bases (= the paper's 713.4M) for SRR3680863, an immutable SRA fact. The primary scientific reproduction — VirusDetect's 11 isolates / 5-of-11 recovery on the same open data — never ran: «our HPC» «job» failed during the conda env-build (transient infra on our side), so the core benchmark claim is untested, not contradicted. No authors'-side or fabrication concern arises; the isolate count is additionally DB-version sensitive and only indirectly comparable. Net: solid 1:1 on the data-fact but incomplete on the central claim → overall yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

135.6 k
tokens (I/O) · 7.9 M incl. cache
25 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
1 (1 failed)
HPC jobs
hummel
machine