Polymorphism identification and improved genome annotation of Brassica rapa through Deep RNA sequencing.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓No authors-side cause for any deviation
- ✓The central claim held under reproduction
- 🔴Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the pipeline-derived SNP-identification result of PMID 25122667 (Brassica rapa R500 vs IMB211 RNA-seq polymorphism calling) using the paper's own third-party toolkit SNPtools, run end-to-end on real SRA data (SRR1227842 + SRR1238059) after fixing three third-party code compatibility bugs (samtools version-regex, two files with removed legacy Perl keys/values-on-scalar syntax, and a non-portable Rscript shebang). The final run (SLURM «job») produced 296,705 SNP/indel calls across 10 chromosomes, 89.7% of the paper's reported 330,995 genome-wide SNPs -- graded partial given materially reduced input depth (1 SRA run per genotype vs. the paper's 33-lane, 8-tissue aggregate), a substituted reference assembly (Ensembl Brapa_1.0 vs. paper's BRAD v1.2, which is not available for unattended download), and no preceding splice-aware alignment pass. Both source datasets were verified 100% complete against ENA metadata (read counts and checksums matched exactly), with good BWA mapping rates (75.19%/71.03% and 76.37%/72.38% mapped/properly-paired for R500/IMB211 respectively). Scratch intermediates were cleaned up after curated results were confirmed (20G to 302M).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-08-03
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-03no human curator yet
- Last updated
- 2026-08-03
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan deep RNA sequencing of two Brassica rapa genotypes (R500 and IMB211, the parents of a mapping population) be used both to identify physically positioned, gene-based SNP markers and to reassemble the transcriptome so as to improve the current B. rapa genome annotation? The premise tested is that transcriptome-based marker discovery and re-annotation will support higher-resolution genetic mapping and more accurate mRNA abundance/eQTL analysis.
- ★ 330,995 SNPs were identified in transcribed regions between B. rapa genotypes R500 and IMB211, at an average frequency of one SNP per 200 bases. finding
- ★ A deep RNA-Seq reassembled B. rapa transcriptome identifies 44,239 protein-coding genes. resource
- ★ Relative to current B. rapa gene models, the reassembly detects 3537 novel transcripts, structurally modifies 23,754 gene models, and changes 3655 annotated proteins. finding
- ★ 780 transcripts could not be mapped to the current B. rapa genome assembly, highlighting gaps in that assembly. finding
- ★ Deep RNA-Seq across 8 tissues, 3 growth environments, and 2 light/density treatments in two genotypes generated 2.3 billion high-quality Illumina reads. method
- SNPtools, a variant-detection, genotype-scoring and visualization tool developed by the authors, refines an R500:IMB211 polymorphism list via a noise-reduction step that filters positions where allele-matching read majorities violate expectation. method
- SNP calls were experimentally validated by Sanger sequencing of 100 PCR-amplified 300-bp fragments, enabling estimation of true-positive, false-positive, and false-negative rates. method
- All SNPs, annotations, and predicted transcripts are made publicly browsable at http://phytonetworks.ucdavis.edu/. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk RNA-Seq (100-bp paired-end, Illumina TruSeq v1 libraries with custom barcoded adapters) | Brassica rapa genotypes R500 (var. trilocularis, Yellow Sarson) and IMB211 (rapid cycling); 8 tissues: root, internode, leaf, petiole, apical meristem, floral meristem, silique, seedling | Simulated sun (high R/FR, 2.0) vs simulated shade (low R/FR, 0.2) in growth chamber; dense vs nondense planting in greenhouse; field-grown controls | Transcript sequence reads used for SNP detection and transcriptome assembly/annotation | Illumina Genome Analyzer GAIIx; Illumina TruSeq v1 RNA Sample Preparation kit (RS-930-2002); RNeasy Plant Mini Kit (Qiagen); DNaseI (Qiagen); NanoDrop ND 1000; Analyst Plate Reader (LJL Biosystems) with SYBR Green I (Invitrogen) |
| Reference-based read alignment and SNP calling | R500 and IMB211 reads aligned separately to B. rapa Chiifu reference genome v1.2 (BRAD) | none | Sequence polymorphisms (R500 vs Chiifu, IMB211 vs Chiifu, deduced R500 vs IMB211); per-allele read counts in genotype files | BWA v0.6.1-r104 [k 1, l 25, n 0.02, e 15, i 10]; TopHat v1.4.1 for splice junctions; SAMtools v0.1.18; Picard; SNPtools v0.1.5 |
| SNP effect annotation | B. rapa genome annotation Brapa_gene_v1.2.gff (BRAD) | none | SNP categorization as intergenic, intronic, exonic (CDS, 5′-region, 3′-region) and as synonymous vs nonsynonymous | SnpEff v3.0 (default parameters) |
| PCR amplification + Sanger capillary sequencing (SNP validation) | Genomic DNA from 7-d-old seedlings of R500 and IMB211; plus RNA-Seq libraries and genomic DNA from recombinant inbred lines (RILs) segregating for four apparent false-positive SNPs | none | Verification of predicted polymorphisms; true-positive, false-positive rates, and false-negative rate within a 150-bp window around each focal SNP | ABI 3730 Capillary Electrophoresis Genetic Analyzer (UC Davis); Primer3Plus primer design; Axygen AxyPrep Mag PCR clean-up kit (MAG-PCR-CL-250); ebioX alignment software; Dellaporta DNA extraction (modified) |
| De novo transcriptome assembly | Pooled post-processed paired-end reads from 10 Illumina GAIIx lanes of B. rapa R500 | none | Assembly metrics as a function of k-mer length: total number of transcripts, N50, longest transcript length, average transcript length; novel transcripts after chimera and splice-variant removal | Velvet v1.2.07 + Oases v0.2.08 (k-mers 31, 35, 41, 45, 51, 55, 61; merged assemblies at k-mer 27 and 55) on XSEDE Lonestar; Trinity r2012-06-08 (k-mer 25, min contig 200, fragment length 500, 16 CPUs, butterfly HeapSpace 10G) on XSEDE Blacklight |
| Reference-based transcriptome assembly | High-quality reads from three Illumina GAIIx lanes (SRR1227842, SRR1228204, SRR1238058) aligned to the B. rapa reference genome | none | Assembled transcript models, run both with and without reference annotation to capture novel and native transcripts | TopHat v1.4.1 (default parameters) + Cufflinks v2.0.2 (default parameters) |
| Ka/Ks ratio computation with permutation testing | In silico R500 and IMB211 genomes generated by substituting all SNPs into the Chiifu reference; CDS extracted per gene, grouped by chromosome and by LF/MF1/MF2 sub-genome | none | Ka/Ks ratio per codon/gene; difference in median Ka/Ks for each chromosome vs rest of genome, compared against 1000 permutations (significance if |difference| exceeds 95% of permuted differences) | vcf2diploid v0.2.6 (AlleleSeq pipeline); gffread (Cufflinks package); KaKs calculator with 'MLWL' codon substitution method |
| Read quality control / pre-processing | Raw Illumina reads from all 126 pooled samples (66 growth chamber, 60 greenhouse+field) | none | Number and length of reads passing quality filtering, de-multiplexing by custom barcode, adapter removal, and paired-read extraction | FastX-toolkit (fastq_quality_filter [q 20, p 85]; fastx_trimmer [f 1,3; l 78]; fastx_barcode_splitter), custom Perl scripts, FastQC |
- – 330,995 SNPs identified in transcribed regions between R500 and IMB211 330,995 SNPs; average of one SNP per 200 bases
- – Reassembled deep RNA-Seq transcriptome identified protein-coding genes in B. rapa 44,239 protein-coding genes
- – Novel transcripts detected relative to current B. rapa gene models 3537 novel transcripts
- – Existing gene models received structural modifications 23,754 gene models modified
- – Annotated protein sequences changed as a result of re-annotation 3655 annotated proteins changed
- – Transcripts that could not be placed on the current genome assembly, indicating assembly gaps 780 unmapped transcripts
- – Sequencing across 16 GAIIx lanes plus one extra lane yielded high-quality reads after QC 3,354,028,456 total reads → 2,550,269,172 reads after quality control (~2.3 billion high-quality reads reported in abstract); 888 GB fastq
- – Average read length across all runs after quality control 98.7 bp average read length (per-run values 89 or 100 bp)
- count 330,995 (SNPs in transcribed regions between the two genotypes R500 and IMB211)
- other one SNP in every 200 bases (Average SNP frequency across transcribed regions)
- count 44,239 (Protein-coding genes in the reassembled B. rapa transcriptome)
- count 3537 novel transcripts; 23,754 gene models with structural modifications; 3655 annotated proteins changed; 780 unmapped transcripts (Annotation improvements relative to current B. rapa gene models)
- count 2.3 billion high-quality Illumina reads (2,550,269,172 post-QC of 3,354,028,456 total) (Total sequencing output; n = number of sequencing runs listed in Table 2)
- count 66 growth-chamber samples and 60 greenhouse/field samples (two pools at 20 nM), sequenced on eight lanes each (16 lanes total) plus one extra lane for failed libraries (Library pooling and sequencing design (Table 1, Table 2))
- mean 98.7 bp (Average read length after quality control across all runs (Table 2))
- count 100 fragments of 300 bases each; 150-bp window used for false-negative estimation; 4 apparent false-positive SNPs re-amplified from RILs (Sanger-based experimental validation of predicted SNPs)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study is primarily a bioinformatics pipeline paper (SNP discovery, genome annotation, transcriptome assembly) built on 2-3 biological replicates per tissue/genotype/treatment combination, with growth-chamber treatments assigned in a stated 'complete randomized design.' The one formal inferential test described is a permutation (randomization) test used to ask whether a chromosome's Ka/Ks ratio differs from the rest of the genome, using 1000 permuted datasets and a 95% empirical threshold. SNP-calling accuracy was assessed by Sanger-sequencing validation of 100 fragments, with true-positive, false-positive, and false-negative rates reported as point estimates. Overall, results are reported mainly as counts, rates, and ratios rather than as group means with a stated dispersion measure, exact p-values, or confidence intervals.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Permutation (randomization) test on Ka/Ks ratio differences by chromosome vs. rest of genome | Chromosome-level Ka/Ks ratio comparison (Ka/Ks ratio computation section) | 1000 permuted chromosome datasets, as stated | not stated |
-
Ka/Ks differences by chromosome were assessed with a permutation test using a 95% empirical threshold applied separately to each chromosome.↳ Could also: Apply a multiple-testing correction (e.g., Benjamini-Hochberg FDR or Bonferroni) across the set of chromosome-level tests — would control the family-wise error rate or false discovery rate when several chromosomes are evaluated simultaneously against the same null distribution
-
Biological replication for tissue/genotype/treatment sampling was 2-3 replicates per group (Table 1), without a stated power calculation.↳ Could also: Report a prospective power analysis or accompany key estimates with confidence intervals — would convey how much statistical certainty the small per-group sample sizes provide
-
SNP-calling true-positive, false-positive, and false-negative rates were estimated from a validation set of 100 fragments and reported as point estimates.↳ Could also: Report confidence intervals (e.g., Wilson score intervals) around these rate estimates — would communicate the precision of validation rates derived from a finite sample of fragments
-
Randomization was explicitly described for the growth-chamber sun/shade design but not stated for the greenhouse dense/nondense or field sampling designs.↳ Could also: State randomization (and blinding, where feasible) procedures for all three growth environments — would make experimental control comparable and transparent across growth chamber, greenhouse, and field conditions
-
The permutation-based Ka/Ks test used the 'MLWL' codon-substitution method without stated distributional assumptions.↳ Could also: A model-based, likelihood-based test of codon substitution (e.g., as implemented in PAML) with explicit assumptions — offers a parametric alternative for assessing selection pressure that could be compared against the permutation-based result
-
SNP effect categories (synonymous/nonsynonymous, genic region) were summarized as counts/proportions without a formal statistical comparison between categories.↳ Could also: A chi-square or Fisher's exact test comparing observed category proportions to an expected distribution — would provide a formal statistical test of whether SNP category proportions deviate from a null expectation
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The reproduction ran the authors' own published toolkit (SNPtools) end-to-end and recovered 296,705 SNP/indel calls against the paper's reported 330,995 — a ratio of 0.897 — with a plausible per-chromosome distribution (A03 = 54,256 highest, A08 = 15,040 lowest). Every material deviation sits on the input/availability side, not the authors' side: ~1/33rd of the sequencing depth (one SRA run per genotype instead of 33 pooled lanes across 8 tissues), Ensembl Plants Brapa_1.0 substituted for the unmirrored BRAD v1.2 reference, and no TopHat splice-aware pre-alignment. The only code changes were three mechanical portability patches (samtools version regex, Perl 5.24+ keys $scalar, Rscript shebang) with no algorithmic effect — a maintenance/bit-rot finding about the repository, not evidence against the paper. Recovering ~90% of the reported count from ~3% of the data supports rather than undermines the central claim, so q5/q6 are yellow (exact value not re-derivable, moderate gap) while q7 stays green.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.