Profiling Selective Packaging of Host RNA and Viral RNA Modification in SARS-CoV-2 Viral Preparations.
The main results reproduced, with only marginal, non-material deviations.
- ✓Same input data as the authors
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction, third-party-tool (P16) on the paper's own GSE182883 data, all compute on «our HPC». Pipeline: bowtie2 --local --no-unal vs a composite C.sabaeus ncRNA + hg38 tRNA + SARS-CoV-2 MN908947.3 reference; per-contig-class read counts. C1 (viral:rRNA read ratio) reproduces well: mean 9.86 vs reported ~9.5 (GEO-faithful read1 single-end) — though read-mate sensitive (read2 gives 0.31). C2: paper's headline enriched isoacceptor Glu(TTC) robustly confirmed (dominant in virion tRNA); the full isoacceptor set is not reproducible because virion tRNA reads are sparse. C3: SRP strongly enriched over tRNA in virion vs cell (785x; reported 150x) — direction/magnitude confirmed, exact fold off. KEY FINDING: the deposited VIRION small-RNA libraries are ~99% adapter-dimer (cell libraries from the same runs align at 72.5%), which is the principal limit on reproducing virion-based claims (C2-C6); this is a data-quality property, not a pipeline error. Two reproduction-critical fixes documented: tRNA reference U->T (RNA vs DNA alphabet) and adapter-trim-before-bbmerge. Nothing observed suggests fabricated values; the gap is virion data quality + an undeposited custom pipeline. All grades provisional, human-checkable (AUDIT.md).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 56assessed: 2026-06-21 ⛓ 6d2aae60efdf
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study asks whether SARS-CoV-2 virions, like retroviruses, selectively package specific host RNAs (tRNAs, tRNA fragments, SRP RNA) during viral assembly, and whether the SARS-CoV-2 genomic RNA itself carries detectable modifications.
- ★ SARS-CoV-2 viral preparations show selective enrichment of specific host tRNAs, tRNA fragments, and SRP RNA compared to uninfected VeroE6 cells finding
- ★ Different SARS-CoV-2 viral preparations (six distinct isolates) contain the same set of enriched host RNAs, suggesting a common packaging mechanism finding
- ★ A single SARS-CoV-2 particle likely contains up to one SRP RNA molecule and four tRNA molecules finding
- ★ tRNA modification levels (e.g., m1A58) differ between tRNAs packaged in viral preparations and their counterparts in VeroE6 host cells, suggesting modification-dependent packaging finding
- ★ SARS-CoV-2 genomic RNA contains uncharacterized candidate modification sites finding
- ★ tRNA Lys (TTT) is among the selectively packaged tRNAs in SARS-CoV-2 virions, paralleling its known role as the HIV-1 reverse transcriptase primer finding
- ★ Some enriched tRNAs in viral preparations (e.g., tRNA Glu (TTC)) are likely tRNA fragments rather than full-length tRNAs, based on read pileup patterns finding
- Demethylase (DM) treatment of small RNA-seq libraries removes Watson-Crick-face tRNA methylations that impede reverse transcription, enabling more quantitative tRNA abundance measurement method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| small RNA-seq (<200 nt, Illumina, with and without demethylase treatment) | VeroE6 cells and cell-free SARS-CoV-2 viral preparations (6 isolates) | SARS-CoV-2 infection/viral culture vs uninfected | tRNA, SRP RNA, rRNA abundance; tRNA modification mutation signatures | Illumina |
| large RNA-seq (>200 nt, size-selected, chemically fragmented) | cell-free SARS-CoV-2 viral preparations (6 isolates) | none (comparison across isolates) | viral genomic RNA SNPs, subgenomic RNA junction reads, SRP RNA/genomic RNA ratio, candidate RNA modifications | Illumina |
| tRNA mutation-signature/RT-stop modification analysis | VeroE6 cells and SARS-CoV-2 viral preparations | demethylase treatment vs untreated | presence/level of tRNA modifications (m1A58, m1G37, I34, m2,2G26, m1G9, m3C) | — |
- ▲ ~150-fold enrichment of SRP RNA relative to tRNA in viral preparations vs. cell samples ~150-fold
- ▲ >200-fold enrichment of SARS-CoV-2 genomic RNA over rRNA in viral preparation samples >200-fold
- – Estimated RNA content per SARS-CoV-2 particle: up to one SRP RNA and four tRNA molecules
- ▲ Six tRNA isoacceptor families significantly enriched across all six viral isolates: Glu(TTC), Lys(TTT), Leu(AAG), Ser(AGA), Ser(GCT), Ser(TGA)
- – Top 3 enriched tRNAs are Lys(TTT), Glu(TTC), Ser(GCT); top 3 depleted are Ile(AAT), Tyr(GTA), Asn(GTT)
- ▲ m1A58 mutation fraction higher in viral-preparation tRNA Leu(AAG) and tRNA Lys(TTT) than in cellular counterparts
- – SARS-CoV-2 genomic RNA to rRNA (18S+28S) read ratio averaged ~9.5, corresponding to a molar ratio of ~2 ratio ~9.5 (reads), ~2 (molar)
- – tRNA Glu(TTC) read pileup in viral preparations drops sharply in the anticodon loop, consistent with it being a 3′ half tRNA fragment
- fold_change ~150-fold (SRP RNA enrichment in viral preparations vs. VeroE6 cells)
- fold_change >200-fold (SARS-CoV-2 genomic RNA enrichment over rRNA in viral preparations)
- other ~9.5 (read ratio), ~2 (molar ratio) (SARS-CoV-2 genomic RNA to 18S+28S rRNA ratio in viral preparations)
- count up to 1 SRP RNA and 4 tRNA molecules per particle (estimated RNA content per SARS-CoV-2 virion)
- count 6 biological isolates (number of distinct SARS-CoV-2 viral isolates sequenced)
- count 3 biological replicates (number of uninfected VeroE6 cell replicates sequenced)
- other ≥50 read coverage (coverage filter applied for high-confidence modification site analysis)
- other >90% mutation fraction (threshold used to call SNPs relative to the Wuhan SARS-CoV-2 reference genome)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This exploratory transcriptomic study used small RNA-seq (<200 nt) and large RNA-seq (>200 nt) to compare host RNAs present in six cell-free SARS-CoV-2 viral preparations versus three uninfected VeroE6 cell cultures. The primary analytical strategy was descriptive and comparative: tRNA isoacceptor abundance fractions were assessed by subtracting viral-preparation values from the VeroE6 cell mean, and RNA modification levels were characterized by mutation fraction signatures in sequencing reads with a ≥50-read coverage filter. Results were communicated as proportions, fold-enrichments, read pileup profiles, and visual summaries (heatmaps, bar plots, box-and-whisker plots) without formal hypothesis tests or reported p-values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Subtraction-based enrichment (viral tRNA isoacceptor fraction minus mean VeroE6 cell tRNA fraction); no formal statistical test applied | Figure 1C — tRNA isoacceptor enrichment/depletion heatmap across all anticodon families | 6 viral preparations vs. 3 uninfected cell replicates | not stated |
| Descriptive comparison of mutation fractions at specific tRNA positions, with a ≥50-read coverage filter as inclusion criterion | Figure 4 — m1A58 and I34 modification fraction comparisons between viral preparations and cells | 6 viral preparations vs. 3 uninfected cell replicates | not stated |
| Fixed threshold SNP calling (>90% mutation fraction relative to Wuhan reference genome) | Figure 5A — single nucleotide polymorphisms across six viral isolates | 6 viral isolates | not stated |
| Read count ratio and molar ratio calculation (fold-enrichment; e.g., SARS-CoV-2 reads vs. 18S+28S rRNA reads; SRP RNA reads vs. viral genome reads) | Figures 5B–5C and main text; ~150-fold SRP enrichment over tRNA, ~200-fold SARS-CoV-2 enrichment over rRNA, average SARS-CoV-2/rRNA read ratio ~9.5 | 6 viral preparations | not stated |
-
Enrichment of tRNA isoacceptors in viral preparations was assessed by subtracting the viral tRNA fraction from the mean VeroE6 cell fraction and visualized as a heatmap, without a formal statistical test↳ Could also: Differential abundance could also be quantified using count-based RNA-seq frameworks such as DESeq2 or edgeR, which model count dispersion across biological replicates and produce per-feature adjusted p-values — Formal count-based testing would provide statistical uncertainty estimates and false-discovery-rate control across all tRNA families simultaneously, enabling readers to distinguish consistent enrichment from sampling variability — particularly relevant here given the small replicate numbers (n=6 vs. n=3)
-
Modification levels at specific tRNA positions (mutation fractions) were compared descriptively between viral preparations and cell controls using a ≥50-read coverage filter as the sole selection criterion↳ Could also: A two-sample test such as a Mann-Whitney U (non-parametric, appropriate for n=3 vs. n=6) or Welch's t-test could also be applied to per-replicate mutation fraction values for each modification site — Formal testing would quantify whether observed differences in mutation fractions exceed expected sampling variability across biological replicates and would allow reporting of effect sizes with associated uncertainty
-
Multiple tRNA isoacceptor families (~50+) and several modification sites were examined simultaneously without an explicit multiple-comparisons correction↳ Could also: A Benjamini-Hochberg false discovery rate (FDR) correction could also be applied across the family of comparisons — When many features are examined simultaneously, FDR control is a standard approach that allows readers to interpret how many reported enrichments are expected to be spurious at a given threshold — especially relevant for an exploratory discovery study such as this one
-
Variability across the six viral preparations and three cell replicates is displayed in Figure 2 with individual points and a mean bar, and in Figure 5D with box-and-whisker plots, but most other comparisons show fractions or pileups without any measure of dispersion↳ Could also: Standard deviation or 95% confidence intervals could also be added to bar or line summaries throughout — With n=3 to n=6 biological replicates, displaying SD or CI alongside means directly communicates within-group spread and the degree of overlap between conditions, making the consistency of enrichment patterns more interpretable
-
SNP calling used a fixed >90% mutation fraction threshold without a probabilistic variant-calling model↳ Could also: Dedicated viral variant callers such as iVar or LoFreq could also be applied, which model sequencing error rates as a function of depth and provide confidence metrics per variant — Probabilistic callers account for position-specific coverage depth and base-quality distributions, and can produce allele frequency estimates with uncertainty bounds — informative especially at lower-coverage genomic positions
-
tRNA isodecoder fractions were first pooled within isoacceptor families for the primary enrichment analysis (Figure 1C), then examined individually only for the six families identified as enriched↳ Could also: Isodecoder-level analysis could also be applied systematically to all isoacceptors from the start using a hierarchical or mixed-effects model that treats isodecoder identity nested within isoacceptor family — Body-sequence differences between isodecoders may independently influence packaging specificity, as the paper itself demonstrates for tRNA Glu(TTC); a systematic isodecoder-level analysis across all families would capture these effects without requiring a post-hoc subset selection step
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.