Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

BaRTv2: a highly resolved barley reference transcriptome for accurate transcript-specific RNA-seq quantification.

Plant J · 2022
L1 94/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
94/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 87% of all assessed papers rank 133 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

1:1 reproduction of the paper's CORE pipeline-derived claims. R1 (deposited-artifact recompute, the primary check): recomputing BaRTv2.18 statistics directly from the deposited transfix GTF + annotation table reproduces the headline numbers EXACTLY - genes 39,434, transcripts 148,260, transcripts/gene 3.76, and protein-coding split 81%/19% (gene-level, via the annotation table's 'Coding potentiality' column); multi-exonic 68.2% vs reported 70% is within-tol. No sign of fabrication in the headline statistics - they are exactly derivable from the shipped deposit, which therefore delivers-what-promised (grade A). R2: the published third-party tool RTDmaker (github.com/anonconda/RTDmaker, the paper's named filtering pipeline) runs end-to-end (EXIT 0) on its shipped potato test_dataset with the README-pinned salmon 1.4.0/gffread 0.12.1; output is aggregate-reproducible (final count stable to 0.01% across reruns) with ~3% transcript-set churn traceable to salmon's non-deterministic quantification on the deliberately tiny subsampled test reads. R3: BUSCO 5.4.7 (embryophyta_odb10) on the deposited transcript FASTA gives 1,540 complete / 32 fragmented vs the paper's 1,530 / 38 (BUSCO v4.0.6) - within-tol, the +0.65% delta consistent with the BUSCO/odb version difference. R4 (the ambitious upstream rebuild of BaRT2.0-Illumina to 142,174 tx) was set up (assembler env built: STAR/Trimmomatic 0.39/StringTie 2.0/Scallop 0.10.5/Cufflinks 2.2.1; full pipeline recipe reconstructed) but is acknowledged-infeasible to reproduce exactly: the cv. Barke genome is locked in interactive IPK repositories (no headless download), the 60 intermediate assemblies are undeposited, and a manual curation step sits between assemblies and the published counts - all caveats the paper itself implies. PRJNA755156 profiled: 20 Illumina libraries match the reported 20 tissues; the PacBio Iso-seq raw reads are NOT under this accession (delivers_promised=partial, grade B). Verdict: described well enough; the paper's central transcriptome and its statistics reproduce 1:1 from the public deposit; the one unreproduced item is an upstream intermediate blocked by data-deposition gaps, not a mismatch.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 94
    assessed: 2026-06-21 ⛓ 3a47a2df5c23
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper addresses whether combining PacBio Iso-seq full-length transcript sequencing with stringently filtered Illumina short-read assemblies, using novel methods to determine high-confidence splice junctions and transcript start/end sites, can produce a more comprehensive and accurate barley reference transcriptome than the earlier short-read-only BaRTv1.0.

Core claims
  • BaRTv2.18 is the most comprehensive and resolved reference transcriptome in barley to date, containing 39,434 genes and 148,260 transcripts resource
  • BaRTv2.18 shows increased transcript diversity and completeness compared with the earlier BaRTv1.0 finding
  • Novel computational methods using high-confidence splice junction (SJ), transcription start site (TSS) and transcription end site (TES) datasets accurately determine transcript boundaries from Iso-seq data method
  • Accuracy of transcript quantification, splice junctions, and transcript start/end sites was validated using high-resolution RT-PCR and 5'-RACE method
  • Sequencing errors near splice junctions (edge wander, local mismatches, template switching) cause false SJ calls that must be filtered using combined Iso-seq and Illumina evidence mechanism
  • Two complementary methods (binomial enrichment for high-coverage genes, fixed sliding window for low-coverage genes) are needed to define high-confidence TSS/TES across genes with highly variable expression levels method
  • The Illumina-assembled dataset (BaRTv2.0-Illumina) supplements the Iso-seq dataset (BaRT2.0-Iso) to overcome incomplete gene coverage in Iso-seq alone method
  • BaRTv2.18 provides a resource including coding regions, protein variants, unproductive transcripts, and functional annotations resource
Experimental setups
Assay System Perturbation Readout Platform
PacBio Iso-seq (single-molecule long-read sequencing) Barley cv. Barke, 21 samples across diverse tissues/growth stages/treatments none full-length non-chimeric (FLNC) reads, splice junctions, transcription start/end sites PacBio Sequel
Illumina short-read RNA-seq Barley cv. Barke, 20 samples across diverse tissues/growth stages/treatments none assembled transcripts (via Cufflinks, Stringtie, Scallop) Illumina
Read mapping and transcript model collapsing/merging (TAMA collapse/merge) Barley cv. Barke Iso-seq FLNC reads mapped to Barke (TRITEX) genome none unified/non-redundant transcript models Minimap2, TAMA
High-resolution reverse transcriptase-PCR (HR RT-PCR) Barley cv. Barke none validation of splice junctions and transcript isoforms
5'-RACE Barley cv. Barke none validation of transcription start sites
Transcript quantification Barley cv. Barke RNA-seq reads vs BaRTv2.18 reference none transcript abundance used to filter low-expressed mono-exonic transcripts Salmon
Key results
  • Final BaRTv2.18 transcriptome contains 39,434 genes and 148,260 transcripts
  • 8,113,088 total CCS reads generated across samples, mean 405,654 reads per sample
  • 7,395,557 FLNC reads generated after processing, of which 93.7% (6,930,934) mapped to the Barke genome 93.7%
  • 257,496 splice junctions identified from Iso-seq data, of which 164,860 (64%) were designated high-confidence 64%
  • 43,302 significantly enriched TSS and 59,944 TES identified in 14,589 genes via binomial probability method
  • Unfiltered Iso-seq transcript dataset of 1,134,325 transcripts reduced via TAMA merge to a final BaRTv2.0-Iso of 103,330 transcripts from 24,630 genes ~11-fold reduction
  • Top 10% of expressed genes contained 79% of FLNC reads, while over 8000 genes had only one or two FLNC reads 79%
Key statistics
  • count 39,434 genes; 148,260 transcripts (final BaRTv2.18 reference transcriptome size)
  • count 8,113,088 CCS reads (total PacBio Iso-seq CCS reads across samples)
  • fold_change 93.7% (6,930,934 of 7,395,557 FLNC reads mapped) (FLNC read mapping rate to Barke genome)
  • count 257,496 SJs identified; 164,860 (64%) high-confidence (Iso-seq splice junction filtering)
  • count 92,636 SJs filtered out (31,541 mismatches within 10bp, 2,687 template switching, 58,408 non-canonical) (reasons for SJ exclusion)
  • count 43,302 TSS and 59,944 TES in 14,589 genes (high-confidence ends from binomial enrichment method)
  • count 1,134,325 transcripts reduced to 103,330 transcripts from 24,630 genes (BaRTv2.0-Iso after TAMA merge redundancy removal)
  • count 33,550 genes and 2,004,544 transcripts (unfiltered Iso-seq transcript dataset before HC filtering)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a bioinformatics resource paper describing the construction of a barley reference transcriptome (BaRTv2.18) by combining PacBio Iso-seq and Illumina RNA-seq data across 20-21 tissue/condition samples, followed by computational filtering pipelines (TAMA, RTDmaker) to define high-confidence splice junctions and transcript start/end sites. The main quantitative inference step described is a binomial probability model used to flag statistically enriched transcription start/end sites from read-end count distributions, with results otherwise reported as descriptive counts and percentages (e.g., numbers of genes, transcripts, splice junctions) rather than through classical inferential hypothesis testing (e.g., t-tests, ANOVA) between experimental groups.

Replicationunclear Sample sizeSample sizes reported as sequencing read counts (e.g., 8,113,088 CCS reads, 7,395,557 FLNC reads across 20-21 samples) rather than as replicate counts or power calculations GroupsIso-seq-derived transcriptome vs Illumina-derived transcriptome; final BaRTv2.18 vs earlier BaRTv1.0 version Pairingna Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Binomial probability test (per Loader, 2000) Identification of significantly enriched transcription start sites (TSS) and transcription end sites (TES) from Iso-seq read-end coordinates per gene (Figure 3a) Based on read-end counts per candidate site relative to total FLNC reads for that gene; a >10 read threshold was used to characterize end-site distributions (Figure S1a,b) not stated
Approaches that could also have been used
  • Significantly enriched TSS/TES sites were called using a binomial probability model applied separately to many thousands of genes, without a stated correction for testing across this large family of sites.
    Could also: A false discovery rate procedure (e.g., Benjamini-Hochberg) applied across the full set of per-gene/per-site binomial tests could also be used. — When many parallel hypothesis tests are performed (here, thousands of genes/sites), an FDR adjustment is a standard way to characterize the expected proportion of false positives among sites called 'significantly enriched.'
  • For genes with lower Iso-seq coverage, fixed empirical windows (e.g., ±20 bp for TSS, ±60 bp for TES) were used to retain transcript ends, based on the observed distribution of end sites (Figure S1).
    Could also: A model-based clustering or kernel-density approach to end-site calling (as used in some CAGE-seq TSS-calling methods) could also be used. — A density- or mixture-model-based method can yield a probabilistic confidence measure for each cluster/peak call, complementing a fixed-window heuristic, particularly useful when read coverage is uneven across genes.
  • Comparisons between BaRTv2.18 and the earlier BaRTv1.0 transcriptome (e.g., counts of genes, transcripts, splice junction categories) were reported descriptively as counts and percentages.
    Could also: A formal statistical comparison of proportions (e.g., a chi-square or two-proportion test) between the two transcriptome versions could also be used for specific categorical comparisons (e.g., canonical vs non-canonical splice site rates). — This would provide an explicit inferential statement about whether the observed differences between versions exceed what might be expected from sampling variation alone, complementing the descriptive summary.
  • High-resolution RT-PCR and 5'-RACE were used for experimental validation of transcript structures (per the abstract), though the excerpt does not detail replicate numbers or a quantitative concordance statistic.
    Could also: Reporting a validation rate with an accompanying confidence interval (e.g., a binomial/Wilson CI on percent concordance) could also be used. — This would let readers gauge the precision of the validation estimate given the number of loci or events tested, beyond a single point estimate.
Software: TAMA (collapse/merge) · PacBio Isoseq3 · Minimap2 · STAR · RTDmaker · Salmon

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35704392 (BaRTv2)

Paper: Coulter et al. 2022, Plant J 112:1373–1391. "BaRTv2: a highly resolved barley reference transcriptome for accurate transcript-specific RNA-seq quantification." DOI 10.1111/tpj.15871 · PMID 35704392 · PMCID PMC9546494.

Repos

Raw data: SRA PRJNA755156 (Illumina RNA-seq + PacBio Iso-seq, cv. Barke, 20 tissues). Deposited final products (open, ics.hutton.ac.uk/barleyrtd):


Reported pipeline-derived numbers (targets to reproduce)

id reported value location
genes_v218 39,434 genes in BaRTv2.18 Results / Abstract
transcripts_v218 148,260 transcripts in BaRTv2.18 Results / Abstract
tx_per_gene 3.76 transcripts/gene (derived) Results
multiexonic_pct 70% multi-exonic genes; 30% single-exon Results / Fig 6
coding_pct 81% protein-coding; 19% non-coding Results
illumina_genes 54,017 genes (BaRT2.0-Illumina, RTDmaker) Table S8
illumina_tx 142,174 transcripts (BaRT2.0-Illumina) Table S8
illumina_rejected 3,460,323 transcripts rejected by RTDmaker Table S8
busco_complete 1,530 complete BUSCO (vs 1,501 in v1.0) Table S12
sj_hc 164,860 HC splice junctions (Iso-seq) Table S6

IN SCOPE (computational, reproducible by running tools on the paper's own data)

  • R1 — Artifact summary statistics (PRIMARY, certain). Recompute genes, transcripts, transcripts/gene, mono/multi-exonic split, intron/SJ count, coding/non-coding from the deposited BaRTv2.18 GTF. Compare to genes_v218, transcripts_v218, tx_per_gene, multiexonic_pct, coding_pct. Pipeline = GTF parsing / gffread. This is the cleanest "delivers-what-promised" + 1:1 numeric check.
  • R2 — RTDmaker tool validation (certain). Run RTDmaker.py ShortReads on the repo's shipped potato test_dataset; confirm it runs end-to-end and is deterministic (re-run = identical output). Establishes the published tool is functional & reproducible.
  • R3 — BUSCO completeness (clean, medium). Run BUSCO (embryophyta/poales) on the deposited BaRTv2.18 transcript FASTA; compare complete-BUSCO count to busco_complete.
  • R4 — RTDmaker on real barley (AMBITIOUS, the floor-is-not-a-ceiling target). Reproduce BaRT2.0-Illumina (illumina_genes / illumina_tx) by regenerating the 60 short-read assemblies (STAR two-pass + Cufflinks + StringTie + Scallop on the 20 Illumina libraries of PRJNA755156) and running RTDmaker ShortReads with the paper's parameters. Caveat up front: the 60 intermediate assemblies are NOT deposited (only raw reads + final RTD), and exact tool versions / a manual v2.10→v2.18 curation step exist, so an exact count match is unlikely; attempt as far as honest, report agreement band.

OUT OF SCOPE (wet-lab / manual / external — not attempted)

  • HR RT-PCR validation (Fig 7, r=0.833) — wet-lab.
  • PacBio Iso-seq read generation (8.1M CCS) — sequencing instrument output.
  • Manual curation BaRTv2.10 → BaRTv2.18 — human decisions, not a deterministic pipeline.
  • TSS/TES motif enrichment (43,302 TSS) — depends on Iso-seq 5'/3' data + thresholds, complex.
  • Mercator4 functional binning — external web portal.

Plan / ordering

  1. («our HPC» front1) download deposited BaRTv2
Figures / tables: Fig6Table
genes_v218
Reported
39,434 genes (BaRTv2.18)
Reproduced
39,434
exact
transcripts_v218
Reported
148,260 transcripts (BaRTv2.18)
Reproduced
148,260
exact
tx_per_gene
Reported
3.76 transcripts/gene
Reproduced
3.76
exact
coding_pct
Reported
81% protein-coding / 19% non-coding genes
Reproduced
81.0% (31,933) / 19.0% (7,501)
exact
multiexonic_pct
Reported
70% multi-exonic / 30% single-exon genes
Reproduced
68.2% / 31.8%
within tolerance
rtdmaker_functional
Reported
RTDmaker ShortReads filtering pipeline (published tool)
Reproduced
runs end-to-end EXIT 0 on shipped test_dataset (salmon 1.4.0/gffread 0.12.1); final RTD 47,880 tx; count reproducible to 0.01% across reruns (set churn ~3% from salmon-quant on subsampled reads)
within tolerance
busco_complete
Reported
1,530 complete / 38 fragmented BUSCO (v4.0.6)
Reproduced
1,540 complete / 32 fragmented (BUSCO 5.4.7, embryophyta_odb10, 1614)
within tolerance
illumina_tx
Reported
142,174 transcripts (BaRT2.0-Illumina, RTDmaker)
Reproduced
not attempted - acknowledged-infeasible (Barke genome only in interactive IPK e!DAL/Galaxy; 60 intermediate assemblies undeposited; manual BaRTv2.10->2.18 curation). Full recipe reconstructed.
m.public.grade.uncheckable
illumina_genes
Reported
54,017 genes (BaRT2.0-Illumina)
Reproduced
not attempted (same blocker)
m.public.grade.uncheckable

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 94/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

477.9 k
tokens (I/O) · 53 M incl. cache
95 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.