The SARS-CoV-2 subgenome landscape and its novel regulatory features.
The main results reproduced: recomputed values matched the published ones within tolerance.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (core). Registry accession SRP056612 is a harvest error (unrelated 2015 human RNA-seq); the real data is GSA CRA002508 (open, 15 runs) under PRJCA002477 — corrected and documented. Ran the paper's STAR pipeline (v2.7.10a; paper v2.7.2b) with the exact customised non-canonical-gap parameters on Vero E6 48h NGS (CRR127514 -> WIV04/MN996528). All 9 canonical SARS-CoV-2 subgenomic-RNA leader-body junctions (S,3a,E,M,6,7a,7b,8,N) were recovered EXACTLY, with leader donors at the TRS-L (~64-69) and biologically correct abundance ranking (N most abundant, ORF7b least) — a clean 1:1 reproduction of the qualitative/structural sgRNA landscape (Fig 3A). The precise '100 significant junctions' (Fig 1D) is graded PARTIAL: the authors' bespoke significance test was not reimplemented, so we report the upstream STAR junction set (4781 raw). No fabrication concern. NOT attempted: wet-lab, the significance toolkit, Nanopore minimap2 secondary metrics, RNA structure, proteomics, controlled human host data.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 initial assessment Score 50assessed: 2026-06-18 ⛓ 31ef6a94bc45
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates the global, dynamic landscape of SARS-CoV-2 subgenomic RNAs (sgRNAs) generated by template switching and asks what molecular rules (RNA-RNA pairing, TRS-dependent and TRS-independent mechanisms) govern this discontinuous transcription process, which had previously been unclear for coronaviruses including SARS-CoV-2.
- ★ Template switching in SARS-CoV-2 can occur bidirectionally, generating diverse subgenomes through successive template-switching events finding
- ★ The majority of template switches result from RNA-RNA interactions via 'seed' and 'compensatory' pairing modes, with terminal pairing status as a key determinant of efficacy mechanism
- ★ Two TRS-independent template switch modes also contribute to subgenome biogenesis mechanism
- ★ Constructed dynamic landscapes of SARS-CoV-2 subgenomes using integrated NGS short-read and Nanopore long-read poly(A) RNA-seq across multiple post-infection time points in two cell types method
- ★ Identified 433 distinct subgenomes clustered into 208 groups, classified into three types: leader/S-N (canonical), ORF1ab/S-N (novel), and S-N/S-N (novel internal) finding
- ★ Multi-switch sgRNAs (bi-switch and tri-switch) exist and share common junction sites, with their abundance correlated to parental single-switch sgRNA expression finding
- ★ Minimum free energy (MFE) of RNA-RNA base pairing correlates with template-switching efficiency (junction read counts) finding
- Developed a statistical toolkit to identify robust, reproducible template-switching junctions from NGS short reads while removing local background noise resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| NGS Illumina paired-end 150-bp poly(A) RNA-seq | Vero E6 cells (monkey), SARS-CoV-2 infected | SARS-CoV-2 infection, time course (6h, 12h, 24h, 48h) | template-switch junction counts/positions, gene expression levels | Illumina NGS |
| Nanopore direct RNA-seq (long-read) | Vero E6 cells (monkey), SARS-CoV-2 infected | SARS-CoV-2 infection, time course | full-length subgenome structures, poly(A) tail length, read length, junction confirmation | Nanopore MinION |
| NGS Illumina paired-end 150-bp poly(A) RNA-seq | Caco-2 cells (human), SARS-CoV-2 infected | SARS-CoV-2 infection, time course (4h, 12h, 24h) | template-switch junction counts/positions | Illumina NGS |
| Nanopore direct RNA-seq (long-read) | Caco-2 cells (human), SARS-CoV-2 infected | SARS-CoV-2 infection, time course | full-length subgenome structures, junction confirmation | Nanopore MinION |
| Northern blotting | Vero E6 cells, SARS-CoV-2 infected | SARS-CoV-2 infection | presence/size validation of sgRNAs | — |
| Computational RNA-RNA interaction / minimum free energy (MFE) prediction | in silico, SARS-CoV-2 genome sequence | none | base-pairing patterns (TRS-L/anti-TRS-B, UR-DR, UL-DL segments), MFE values correlated with junction read support | — |
- – 100 statistically significant NGS junctions and 141 significant Nanopore junctions identified in Vero E6 cells 48h post-infection, with 31 overlapping between platforms 100 / 141 / 31 junctions
- ▲ Canonical leader-body junctions represent the large majority of total junction counts in Nanopore reads 57.8% (99,548/172,107)
- – 433 distinct subgenomes identified in 208 clusters across three sgRNA classes 433 subgenomes / 208 clusters
- – Multi-switch (bi- and tri-switch) sgRNAs share common junction sites, e.g., 7 bi-switch and 4 tri-switch sgRNAs sharing junction 28,525–28,576 7 bi-switch, 4 tri-switch sgRNAs
- ▲ Strong positive correlation between multi-switch read counts and corresponding single-switch read counts for leader-type sgRNAs Spearman r=0.89
- ▲ N protein/sgRNA shows the highest junction expression level, with expression increasing from 5' to 3' direction of the genome
- ▲ Extensive RNA-RNA base pairing (7–12 consecutive bp) observed beyond the canonical 6-bp TRS-L/anti-TRS-B duplex for the 9 canonical sgRNAs 7-12 bp beyond 6 bp
- ▲ Fraction of SARS-CoV-2 reads increased over infection time course in both cell types, higher in Vero E6 than Caco-2 Vero E6: 0.1%→7%→50%→80-90%; Caco-2: 0.02%→1.2%→21%
- correlation Spearman r=0.89 (multi-switch reads vs. single-switch reads for leader-type sgRNAs)
- count 45,343 junctions (total junctions identified in combined NGS replicates, Vero E6 48h)
- count 100 significant NGS junctions; 141 significant Nanopore junctions; 31 overlapping (Vero E6 cells 48h post-infection)
- fold_change 57.8% (99,548/172,107) (proportion of canonical junction counts among total junction counts in Nanopore reads)
- count 433 subgenomes in 208 clusters (full-length sgRNAs identified in Vero E6 cells 48h post-infection)
- count 7 bi-switch and 4 tri-switch sgRNAs (sgRNAs sharing junction site 28,525–28,576)
- mean median poly(A) length ~50 nt; median Nanopore read length 1,248 nt (Nanopore reads, Vero E6 48h sample)
- other viral read fractions: Vero E6 0.1%, 7%, 50%, 80-90% (4h,6h,12h,24-48h); Caco-2 0.02%, 1.2%, 21% (4h,12h,24h) (fraction of SARS-CoV-2 reads over infection time course, Table S1)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study characterizes SARS-CoV-2 subgenomic RNA landscapes using hybrid NGS (Illumina short-read) and Nanopore long-read poly(A) RNA sequencing in Vero E6 and Caco-2 cells at multiple time points post-infection, with duplicate libraries per platform. Junction significance was determined via a custom statistical scoring approach that accounts for local background coverage around both junction sites; Spearman correlation and one-sided t-tests were subsequently used to assess relationships between RNA-RNA interaction features (e.g., minimum free energy) and template-switching read counts. Results are reported primarily as read counts, correlation coefficients, and p-values from group comparisons.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Custom statistical scoring (local background normalisation around junction sites) | Identification of significant junctions from NGS data (Figure 1D, Figures S2A–S2D) and Nanopore data (Figure S2E) | — | not stated |
| Spearman rank correlation | Correlation between multi-switch and single-switch read counts for leader-type sgRNAs (Figure 2G, r = 0.89); correlation between MFE and NGS read counts for 9 major leader-group sgRNA junctions (Figure 3F) | 9 major leader-group sgRNA junctions for Figure 3F; not stated for Figure 2G | not stated |
| One-sided t-test | Comparison of MFEs across sub-groups of leader-group sgRNA junctions stratified by NGS read count (Figure 3G) | Number of junctions per group stated in figure but not given in excerpt | not stated |
-
One-sided t-tests were used to compare MFE distributions across sgRNA junction groups stratified by read-count bins↳ Could also: A non-parametric Kruskal-Wallis test followed by Dunn's or Mann-Whitney U post-hoc tests with Benjamini-Hochberg FDR correction could also be applied across the same groups — MFE values in small-n groups may not be normally distributed; non-parametric alternatives do not assume normality, and a family-wise correction would explicitly control the false-discovery rate when comparing multiple group pairs simultaneously
-
The directional choice of one-sided t-tests for the MFE–read-count relationship implies an a priori hypothesis about the direction of effect↳ Could also: Two-sided t-tests or Mann-Whitney U tests could also be used, with directionality reported post-hoc if confirmed — Two-sided tests are the default in exploratory genomic work and do not require the direction of the effect to be specified in advance, making the inference more conservative and broadly reproducible
-
Spearman correlation was used to relate multi-switch read counts to single-switch read counts and to relate MFE to read counts↳ Could also: A negative-binomial or zero-inflated regression model could also be used, treating read counts as the outcome and MFE as a continuous predictor — Read counts from RNA-seq are overdispersed count data; a count-regression framework explicitly models this distribution and can yield confidence intervals and covariate adjustment alongside the association estimate
-
Junction significance was assessed with a custom local-background scoring approach described in STAR Methods↳ Could also: Established RNA splicing or chimeric-read significance frameworks (e.g., the statistical model in STAR aligner's chimeric-junction module, or a negative-binomial test as used in DESeq2/edgeR) could also be applied to score junction read counts against local background — Published tools with documented statistical properties and wide community validation allow readers to independently replicate the significance calls and compare results across studies
-
Replication is described only as 'duplicates'; whether these are biological or technical replicates is not stated in the available text↳ Could also: Explicit reporting of biological vs. technical replicates—and ideally formal biological replication with independent infections—could also accompany the current design — Distinguishing biological from technical replicates is important for interpreting reproducibility estimates and variance components, and is recommended by ENCODE and MIQE guidelines for sequencing experiments
-
Results are reported primarily as raw read counts and rank correlations, without confidence intervals around the correlation estimates↳ Could also: Bootstrap or Fisher-z-transformed 95% confidence intervals around Spearman r could also be reported alongside the point estimate — Confidence intervals convey the precision of the association estimate and are especially informative when n (number of junctions) is small, as in the 9-junction MFE correlation in Figure 3F
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33713597
Paper: Wang D et al. "The SARS-CoV-2 subgenome landscape and its novel regulatory features." Mol Cell 2021. PMID 33713597 · PMC7927579 · DOI 10.1016/j.molcel.2021.02.036.
What the paper does
Maps the SARS-CoV-2 subgenomic RNA (sgRNA) landscape from poly(A) RNA of infected cells, using NGS Illumina short reads + Nanopore MinION long reads, in Vero E6 (monkey, ATCC CRL-1586) and Caco-2 / Calu-3 (human) cells infected with SARS-CoV-2 WIV04 (IVCAS 6.7512). sgRNAs arise by TRS-mediated template switching, which appears in the data as leader–body junctions (and body–body junctions). The core computational result is detection + quantification of these junctions.
⚠️ Data accession correction (harvest error)
The registry/BRIEF lists sra:SRP056612. That is wrong — SRP056612 is an
unrelated Homo sapiens RNA-seq study from ~2015 (runs SRR1942xxx,
"Transcriptome or Gene expression"). It is NOT SARS-CoV-2 and NOT this paper.
The paper's real raw data (per its Data & Code Availability statement) is in the
China National Genomics Data Center / GSA:
- GSA: CRA002508 (open) + GSA-human: HRA000412 (controlled) under
project PRJCA002477, host
https://download.cncb.ac.cn/gsa/CRA002508/. We reproduce against CRA002508 (open, downloadable, 15 runs, md5s present).
Pipeline (from Methods → STAR is the in-scope tool, repo = alexdobin/STAR)
cutadapt v2.5— adapter removal; filter rRNA.STAR v2.7.2b— map clean reads to host genome (Vero E6 = Chlorocebus sabaeus Ensembl v99; Caco-2/Calu-3 = human hg38) with--sjdbScore 1 --outFilterMultimapNmax 20 --outFilterMismatchNmax 999 --outFilterMismatchNoverReadLmax 0.04 --alignIntronMin 20 --alignIntronMax 1000000 --alignMatesGapMax 1000000 --alignSJoverhangMin 8 --alignSJDBoverhangMin 1.STAR v2.7.2b— unmapped reads → virus genome (SARS-CoV-2 WIV04 = NCBI MN996528; SARS NC_004718.3; MERS NC_038294.1) with customised params to allow non-canonical gaps (template switches):--outFilterMultimapNmax 1 --alignSJoverhangMin 8 --outSJfilterOverhangMin 8 8 8 8 --outSJfilterCountUniqueMin 3 3 3 3 --outSJfilterCountTotalMin 3 3 3 3 --outSJfilterDistToOtherSJmin 0 0 0 0 --scoreGap -4 --scoreGapNoncan -4 --scoreGapATAC -4 --alignIntronMax 30000 --alignMatesGapMax 30000 --alignSJstitchMismatchNmax -1 -1 -1 -1. Uniquely-mapped reads kept. Junctions read fromSJ.out.tab= template-switch events.- Nanopore:
guppybasecall →NanoPlotQC →nanopolishpolyA →minimap2 v2.17 -ax splice -un -k14 --no-end-flt --secondary=noto combined host+virus genome → junctions from long reads. - A custom statistical "toolkit" then scores junctions for significance.
IN SCOPE (pipeline-derived, we attempt)
- R1 — STAR viral junction detection (NGS, Vero E6 48h, 2 reps). Re-run
STAR v2.7.2b on CRR127514/CRR127515 → recover the leader–body junctions of
the canonical sgRNAs from
SJ.out.tab. Core, low-hanging.- Claim: "100 significant junctions" (Fig 1D, red points; Table S2) in Vero E6 48h NGS. (Note: "significant" = their custom test; we reproduce the upstream STAR junction set + canonical-sgRNA recovery; raw junction count is a provisional comparison, not the significance-filtered 100.)
- Claim: 9 canonical sgRNAs observed in almost all samples (Fig 3A),
with leader TRS motif ACGAAC. Highly checkable: the 9 canonical
leader–body junctions should dominate the viral
SJ.out.tab.
- R2 (secondary, Nanopore) — canonical junctions = 57.8% (99,548/172,107) of total junction counts in full-length Nanopore reads (CRR127516/127517 via minimap2). Attempt if R1 lands and time permits.
- R3 (secondary, Nanopore QC) — median Nanopore read length 1,248 nt; 22 reads cover the whole ~30 kb genome (NanoPlot on Vero E6 48h Nanopore).
OUT OF SCOPE (not pipeline / not attempted)
- Wet-lab: cell culture
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.