Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Viral Diagnostics in Plants Using Next Generation Sequencing: Computational Analysis in Practice.

Front Plant Sci · 2017
L1 100/100 PQI 94
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PMID 29123534 is a review/tutorial; the only pipeline-derived results that are the authors' OWN are their hands-on benchmark of public virus-detection tools on 3 public SRA datasets (Table 2). Described well enough to attempt: yes, for the open tools. CLEAN 1:1 reproduced (S1, exact): the grapevine sRNA-seq dataset the paper benchmarked VirusDetect on is 713.4M bases — ENA base_count for SRR3680863 = 713,356,593, matching exactly (immutable SRA fact, compute-independent). PRIMARY tool reproduction (S2 = VirusDetect on grapevine SRR3680863, an open third-party tool on the paper's own data) was authored and submitted to «our HPC» («job») but FAILED ~40s into the conda env-build, mid package download, BEFORE the pipeline executed — a transient HPC infra failure, NOT a statement about the paper (the env solve resolved every pin cleanly; a resubmit with the now-warm «infra» pkgs cache would likely proceed). Net: partial — one data-scale claim reproduced exact, the tool-output comparison incomplete. NOT attempted (the ~20%): Taxonomer on pear/pepper (closed web service, classifier+DB version unpinnable), VirFind (web service), VSD/Yabi (github.com/muccg/yabi is a generic workflow web-portal engine, not the analysis), Metavisitor (Galaxy suite). No fabrication concerns; S1 matches SRA exactly.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-16 ⛓ 5b777fdb3987
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Can hypothesis-free RNA-sequencing (RNA-seq/NGS) of plant material answer the question of how many and which different viruses are present in a crop plant, overcoming the limitation of RT-PCR-based diagnostics that only detect known, targeted viruses?

Core claims
  • Molecular techniques such as RT-PCR only allow detection of known viruses, with each test specific to one or a small number of related viruses, so unknown viruses can be missed finding
  • NGS/RNA-seq enables unbiased, hypothesis-free detection of multiple known and emergent viruses in plant material without prior knowledge of what is being sought finding
  • Across published plant RNA-seq virus-detection studies, co-infection of individual plants with more than one virus is a consistent theme finding
  • Sequencing of total small RNAs (sRNA-seq) is an effective method for virus detection because antiviral RNA interference enriches virus-derived siRNAs in the host method
  • Common elements across virus-detection RNA-seq workflows are quality control of raw reads, assembly into contigs, removal of host sequences, and identification of viral reads via mapping to a virus database method
  • Different de novo assembly and analysis tools (e.g., Trinity vs. Velvet) applied to the same dataset can identify different combinations of viruses finding
  • Plant virus detection is computationally harder than human virus detection because many crop genomes are unknown/incomplete and plant virus sequences are poorly represented in databases, relative to human data finding
  • Multiple bioinformatics tools/pipelines (VirFind, Taxonomer, VSD toolkit, Metavisitor, VIP, ViromeScan, VirusHunter) have been developed for virus identification from RNA-seq data, but most are focused on human clinical samples rather than plants resource
Experimental setups
Assay System Perturbation Readout Platform
RNA-seq (de novo assembly, Blastn/Blastx classification) Garlic (Allium sativum, A. vineale) leaves, imported and native to Australia none (natural field infection) number and identity of virus isolates per plant Illumina HiSeq 2000
RNA-seq (de novo assembly with Trinity and Velvet/Oases, MEGABLAST) Pepper (Capsicum annuum), cultivars Pusa Jwala and Taiwan-2 none (natural infection, susceptible vs resistant cultivar) identity and number of viruses per assembler Illumina HiSeq 2000
RNA-seq/sRNA-seq reanalysis of public transcriptome data (Trinity assembly, MEGABLAST) Pear (Pyrus pyrifolia) transcriptome across developmental stages none virus read counts and identity SRA dataset SRX532394
RNA-seq (paired-end, Velvet assembly, Blast) Grapevine (Vitis vinifera), lignified cane material from a merlot vineyard, South Africa none identity of infecting viruses/mycoviruses Illumina Genome Analyzer
RNA-seq reanalysis of public transcriptome (Trinity assembly, Blast) Grapevine cultivar Tannat, grain/skin/seed tissue libraries none (tissue-type comparison) virus prevalence and distribution by tissue type Illumina HiSeq 1000
sRNA-seq (Velvet assembly, Blast, MAQ alignment) Sweet potato (Ipomoea batatas) leaf material, Honduras and Guatemala none identity of co-infecting RNA and DNA viruses; symptom severity correlation Illumina Genome Analyzer
RNA-seq (CLC Assembly Cell and Trinity de novo assembly, Blastx/Blastn) Orange (Citrus sinensis), CSD-symptomatic and -asymptomatic trees, Sao Paulo, Brazil none (symptomatic vs asymptomatic comparison) virus/genotype identity associated with disease symptoms Illumina HiSeq 2000
sRNA-seq (BWA alignment to host/virus, Velvet de novo assembly, Blast) Tomato (Solanum lycopersicum), symptomatic plants from US and Mexico none complete virus/viroid genome assembly and strain differentiation Genome Analyzer II
Key results
  • Cultivated garlic plants contained between 1 and 8 virus isolates each; 41 virus isolates identified in total including potyviruses, allexiviruses and carlaviruses 1-8 isolates/plant; 41 total
  • Trinity and Velvet assemblies of pepper RNA-seq data identified different combinations of viruses, but 8 viruses were common to all datasets; a novel virus (Pepper Virus A) was identified 8 common viruses
  • Pear transcriptome analysis revealed 5 viruses with read counts >5, including ASGV plus three additional viruses (PrVT, AGCAV, ASPV) 5 viruses, read counts >5
  • Grapevine cane material contained GLRaV-3, GRSPaV and GVA; Grapevine virus E was newly reported in South African vineyards, and mycoviruses were isolated from grapevine phloem for the first time
  • In grapevine cv Tannat, the most prevalent virus differed by library/tissue; 4 viruses were found in all three tissues while OBDV and PVS were seed-specific, and skin tissue showed higher overall virus prevalence than grain
  • Sweet potato samples showed co-infection with three RNA viruses and two DNA viruses; when SPPV-B and SPCSV-WA co-occurred, disease symptoms were consistently severe 5 viruses detected
  • Orange trees showed mixed infections of CTV, CSDaV, CitPRV plus two novel viruses (CJLV, CVLV); two genotypes each of CTV and CSDaV were distinguished, with one CSDaV genotype linked to symptomatic plants
  • Complete genomes of six Pepino mosaic virus isolates and one Potato spindle tuber viroid isolate were assembled from tomato sRNA-seq data, including two co-infecting PepMV strains and a novel potyvirus 6 PepMV isolates
Key statistics
  • count 41 virus isolates identified in total (garlic (Allium sativum) RNA-seq survey)
  • count 1 to 8 virus isolates per cultivated garlic plant (garlic virus load per plant)
  • count 8 viruses common to all pepper datasets (pepper Trinity vs Velvet assembly comparison)
  • count 5 viruses with read counts greater than 5 (pear transcriptome virus detection)
  • other less than 10-fold coverage threshold used to exclude contigs (garlic study contig filtering criterion)
  • count more than 30 viruses known to infect sweet potato (background statistic on sweet potato virus susceptibility)
  • count over 4 million orange trees lost (Citrus sudden death disease impact in Sao Paulo State, Brazil)
  • count 6 complete Pepino mosaic virus (PepMV) isolate genomes assembled (tomato sRNA-seq virus genome assembly)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a narrative review article surveying computational and bioinformatics workflows for RNA-seq-based viral detection in crop plants. It qualitatively summarizes the methods, tools, and findings of previously published studies and does not itself collect data, define experimental groups, or apply inferential statistical tests; results are reported descriptively (e.g., counts of viruses identified, contig coverage thresholds).

Replicationna Sample sizeNot applicable as a primary study; the review enumerates the sample basis of cited studies (e.g., a benchmarking set of 38 plant samples from 19 species for VirFind, 21 plant genomes for the VSD toolkit), but no sample size or power analysis is described for the review itself. GroupsTools/workflows and crop case studies compared narratively Pairingna Randomization/blindingna Dispersionnone
Approaches that could also have been used
  • The review synthesizes the literature qualitatively, describing each study in turn.
    Could also: A structured or systematic review with predefined inclusion criteria and a tabulated cross-study comparison (e.g., PRISMA-style flow) could also have been used. — A systematic approach would add reproducibility and transparency about how studies were selected, complementing the narrative synthesis.
  • Bioinformatics tools are compared narratively in a summary table (Table 1).
    Could also: A common benchmarking framework applying each tool to a shared reference dataset with standardized performance metrics (sensitivity, precision, F1) could also have been used. — Quantitative benchmarking on common data would allow direct, like-for-like comparison of tool performance across the surveyed workflows.
  • Virus detection in cited studies relied on coverage/length thresholds (e.g., >10-fold coverage, contigs >1,000 nt, read counts >5).
    Could also: Reporting threshold sensitivity analyses or statistical confidence measures for calls could also accompany such cutoffs. — Sensitivity analysis would convey how robust detection counts are to the chosen thresholds, which is informative when comparing across datasets.
Software: Geneious Pro (de novo assembly, cited study) · CLC Genomics Workbench / CLC Assembly Cell (de novo assembly, cited studies) · Trinity (de novo assembly, cited studies) · Velvet (with Oases) (de novo assembly, cited studies) · SPAdes (de novo assembly, VSD toolkit) · BLAST / Blastn / Blastx / MEGABLAST (sequence comparison, cited studies) · Bowtie2, BWA, MAQ (read mapping, cited studies/tools)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
110
Impact: high
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 75/100
stands on reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

SRX532394 ENA in Introduction (http://purl.org/orb/Introduction)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 29123534

Paper: Jones S, Baizan-Edge A, MacFarlane S, Torrance L. (2017) Viral Diagnostics in Plants Using Next Generation Sequencing: Computational Analysis in Practice. Front Plant Sci 8:1770. DOI 10.3389/fpls.2017.01770.

Nature of the paper

This is a review / "in practice" tutorial, not a primary-data study. Most of the numbers it cites (garlic 41 virus isolates, pepper novel virus, tomato 6 PepMV genomes, etc.) are quoted from other groups' published studies — those are out of scope: they are not results this paper's authors computed, and the inputs/parameters are in the cited papers, not here.

The authors' own computational contribution is a hands-on benchmark: they took three public SRA datasets (their Table 2) and ran them through several publicly available virus-detection tools (Table 1: VirFind, VirusDetect, Taxonomer, VSD/Yabi, Metavisitor), reporting how each performed. Those benchmark runs are the pipeline-derived results in scope.

Table 2 — datasets the authors re-analysed (verified against ENA)

Organism Paper SRA Run (ENA) Type reads bases orig-study viruses
Pear (Pyrus pyrifolia) SRR1269627 (= exp SRX532394) SRR1269627 RNA-seq SE 97,896,223 3,524,264,028 ASGV, AGCAV, ASPV, PrVT
Pepper (Capsicum annuum) SRR1123893 SRR1123893 RNA-seq PE 53,921,012 10,784,202,400 13 viruses (ALPV, BPEV, … TVCV)
Grapevine (Vitis vinifera) SRR3680863 SRR3680863 sRNA-seq SE 32,497,945 713,356,593 GRSPaV, GVB, GFkV, GLRaV-3, HSVd

Note: the RU brief lists the data accession as SRX532394; that is the SRA experiment for the pear run SRR1269627 (Table 2). Both resolve to the same pear RNA-seq library.

IN SCOPE (pipeline-derived, this paper's own runs)

  • S1 — dataset facts. The reads/bases of the three Table-2 datasets are immutable SRA facts and the paper quotes one directly: "VirusDetect was tested on sRNA-seq data from grapevine (713.4M bases)." → ENA base_count for SRR3680863 = 713,356,593 ≈ 713.4 M. Verifiable with zero ambiguity (control-plane lookup).
  • S2 — VirusDetect on grapevine sRNA-seq (SRR3680863). Paper: "VirusDetect identified 11 virus isolates in the grape sRNA-seq data, 5 of which were identified in the original study" and "After file upload it gave results in < 4 h." Pipeline: VirusDetect (kentnf/VirusDetect; sRNA clean → Velvet de-novo + reference-guided assembly → megablast/blastx vs vrl_plant curated plant-virus DB). This is open-source and runnable on the paper's exact data → chosen primary reproduction target. Metric we compare: number of reported virus isolates, and recovery of the 5 original-study viruses (GRSPaV, GVB, GFkV, GLRaV-3, HSVd).

IN SCOPE BUT NOT ATTEMPTED (the hard ~20%, with reasons)

  • Taxonomer on pear / pepper ("8% of pear reads classified", "5,707 virus reads pear / 364,959 pepper", "8 of 13 pepper viruses"). Taxonomer is a closed web service (taxonomer.com) requiring an account and a remote classifier whose database/version cannot be pinned → not reproducibly runnable; skipped per 80/20 (env_unresolvable for that sub-result).
  • VirFind / VSD-Yabi / Metavisitor benchmark mentions: Yabi (github.com/ muccg/yabi) is a generic workflow web-portal engine, not the analysis itself; VirFind is a web service; Metavisitor is a Galaxy toolshed suite. Out of 80/20 budget; VirusDetect (S2) already provides one clean third-party-tool 1:1.

OUT OF SCOPE (not this paper's computation)

  • All per-crop result numbers quoted from the reviewed primary studies (garlic Wylie 2014, tomato Li 2012, sweet-potato Kashif 2012, orange Matsumura 2017, etc.). These are literature citations, reproducing them would mean reproducing those papers.
  • Cost/£ statements, wet-lab/library-prep, and general methodological commentary.

Caveat on the comparison (DB drift)

V

S1
Reported
713.4M bases (grapevine sRNA-seq, VirusDetect benchmark input)
Reproduced
713356593 bases for SRR3680863 (=713.4M)
exact
S2a
Reported
VirusDetect identified 11 virus isolates in the grape sRNA-seq data
Reproduced
not produced — «our HPC» «job» failed at ~40s during conda env-build, before pipeline ran (transient infra, not a paper finding)
m.public.grade.error
S2b
Reported
5 of the 11 isolates were also found in the original study (GRSPaV, GVB, GFkV, GLRaV-3, HSVd)
Reproduced
not produced (same env-build failure)
m.public.grade.error
S2c
Reported
VirusDetect returned results in < 4 h
Reproduced
not produced (same env-build failure)
m.public.grade.error

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

Only one in-scope claim could be checked, and it reproduced exactly: the grapevine sRNA-seq input is 713,356,593 bases (= the paper's 713.4M) for SRR3680863, an immutable SRA fact. The primary scientific reproduction — VirusDetect's 11 isolates / 5-of-11 recovery on the same open data — never ran: «our HPC» «job» failed during the conda env-build (transient infra on our side), so the core benchmark claim is untested, not contradicted. No authors'-side or fabrication concern arises; the isolate count is additionally DB-version sensitive and only indirectly comparable. Net: solid 1:1 on the data-fact but incomplete on the central claim → overall yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

135.6 k
tokens (I/O) · 7.9 M incl. cache
25 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
1 (1 failed)
HPC jobs
hummel
machine