In vivo structural characterization of the SARS-CoV-2 RNA genome identifies host proteins vulnerable to repurposed drugs.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED. icSHAPE in-vivo structure characterization of the SARS-CoV-2 genome (Sun et al., Cell 2021) from GSE153984, via STAR + icSHAPE-pipe-equivalent processing on «our HPC». All three in-scope pipeline-derived claims land: C1 genome icSHAPE coverage = 99.883% (matches >99.88%) -- confirmed BOTH from the authors' shipped virus-w50.shape AND independently regenerated from raw fastq (STAR->pysam RT-stop pileup, coverage>=20x = 99.883%, JOB B); C2 read depth mean 136M across 6 huh7 libs (~within-tol of 'about 150M'; in-vivo NAI reps avg 185M; JOB B trim counts matched ENA read_count exactly); C3 in-vivo inter-replicate RT-stop Pearson r = 0.9983 > 0.99 (exact, from raw regeneration; base-density r=0.9995; shipped reactivity proxy r=0.9924). No fabrication signal: the exact 99.883% coverage recurs independently from raw reads. C4-C7 (conserved elements, variable regions, covariant pairs, PrismNet host RBPs) are downstream / deep-learning / wet-lab and out of scope (not attempted). Dataset GSE153984 profiled: complete open deposit, 22 runs = 11 conditions x2 reps, delivers-promised, quality B (processed .shape profiles split into the authors' GitHub rather than GEO). Grades are provisional pending human sign-off in AUDIT.md.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-18 ⛓ 921338406c3c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether determining the in vivo (and in vitro) secondary structure of the SARS-CoV-2 RNA genome can reveal functional structural elements and host RNA-binding proteins that regulate viral infection, thereby nominating druggable targets among FDA-approved drugs.
- ★ icSHAPE was used to determine the in vivo and in vitro structural landscape of the SARS-CoV-2 RNA genome in infected Huh7.5.1 cells, plus UTR structures of six other coronaviruses method
- ★ The in vivo structural models validate several structural elements previously predicted in silico (e.g., Rangan et al. 5'UTR/3'UTR models) while revealing notable differences finding
- ★ Discovered structural features that affect the translation and abundance of subgenomic viral RNAs in cells finding
- ★ A deep-learning tool informed by the structural data predicted 42 host proteins that bind SARS-CoV-2 RNA resource
- ★ Antisense oligonucleotides (ASOs) targeting viral RNA structural elements dramatically reduced SARS-CoV-2 infection in liver- and lung-tumor-derived cells finding
- ★ FDA-approved drugs inhibiting the predicted SARS-CoV-2 RNA-binding proteins dramatically reduced SARS-CoV-2 infection in cells finding
- ★ The in vivo SARS-CoV-2 RNA structure differs substantially from in vitro refolded and purely theoretical structures, with the viral genome more single-stranded in vivo finding
- Covariation analysis across coronavirus genomes identified 170 co-variant base pairs, including a novel duplex between the 3'UTR and ORF10 finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| icSHAPE (in vivo RNA structure probing) | Huh7.5.1 human liver cancer cells infected with SARS-CoV-2 | NAI-N3 chemical probing | icSHAPE reactivity score per nucleotide (single-stranded vs base-paired) | deep sequencing / icSHAPE-pipe |
| icSHAPE (in vitro RNA structure probing) | SARS-CoV-2 RNA purified from infected Huh7.5.1 cells, refolded in vitro | NAI-N3 chemical probing | icSHAPE reactivity score per nucleotide | deep sequencing / icSHAPE-pipe |
| icSHAPE on in vitro transcribed UTRs | UTRs of SARS-CoV-2 (reference and mutant) and six other coronaviruses (e.g., SARS-CoV, MERS-CoV) | NAI-N3 chemical probing after in vitro transcription/refolding | icSHAPE reactivity score, comparative UTR structure | deep sequencing |
| Structure validation against reference RNAs | 18S rRNA, 28S rRNA, SRP RNA in Huh7.5.1 cells | none | AUC of icSHAPE scores vs known reference structures | — |
| Deep-learning RBP-RNA binding prediction | SARS-CoV-2 RNA (UTRs) structural/sequence data | none (computational) | predicted host RNA-binding proteins (42 candidates) | neural network model (Sun et al., 2021) |
| SARS-CoV-2 infection assay with antisense oligonucleotide treatment | cells derived from human liver and lung tumors | ASOs targeting RNA structural elements | SARS-CoV-2 infection level | — |
| SARS-CoV-2 infection assay with drug treatment | cells derived from human liver and lung tumors | FDA-approved drugs inhibiting predicted RNA-binding proteins | SARS-CoV-2 infection level | — |
| Covariation/phylogenetic analysis | deduplicated coronavirus genome sequences | none (computational) | co-variant base pairs supporting structural models | Infernal package |
- – icSHAPE scores obtained for more than 99.88% of nucleotides of the in vivo SARS-CoV-2 RNA genome 99.88%
- – Inter-replicate correlation very high: RPKM correlation >0.98, RT-stop correlation >0.99 r>0.98 / r>0.99
- ▲ High AUC for icSHAPE scores fitting known reference structures of 18S rRNA, 28S rRNA, and SRP RNA AUC=0.813 (18S), 0.804 (28S), 0.730 (SRP)
- – icSHAPE-derived structure model agreement with Rangan et al. theoretical models AUC=0.854 (5'UTR), 0.692 (3'UTR); sensitivity 0.945/0.913, PPV 1.0/0.824
- – In vivo and in vitro structural profiles of the SARS-CoV-2 genome are only moderately correlated, with 371 structurally variable regions identified genome-wide r=0.58; 371 regions
- – 170 co-variant base pairs identified across coronavirus genomes, including 6 in the 5'UTR, 12 in the 3'UTR, and 5 in a novel 3'UTR-ORF10 duplex 170 pairs (6+12+5 subsets)
- ▲ SARS-CoV-2 RNA genome appeared more single-stranded in vivo than in vitro
- – Deep-learning tool predicted 42 host proteins binding SARS-CoV-2 RNA, several validated physically/functionally 42 proteins
- correlation >0.98 (inter-replicate correlation of host transcriptome RPKM expression)
- correlation >0.99 (inter-replicate correlation of RT-stop from NAI-N3 modification on viral RNA genome)
- other AUC=0.813, 0.804, 0.730 (AUC of icSHAPE scores vs reference structures for 18S rRNA, 28S rRNA, SRP RNA)
- other AUC=0.854 (5'UTR), 0.692 (3'UTR) (AUC of icSHAPE scores vs Rangan et al. theoretical UTR models)
- other sensitivity 0.945/0.913, PPV 1.0/0.824 (sensitivity and PPV of in vivo structure model vs Rangan's model for 5'UTR/3'UTR)
- correlation 0.58 (Pearson correlation between in vitro and in vivo icSHAPE structural profiles)
- count 371 (structurally variable regions between in vivo and in vitro data across the whole genome)
- count 170 (co-variant base pairs identified across coronavirus genomes (6 in 5'UTR, 12 in 3'UTR, 5 in 3'UTR-ORF10 duplex))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper characterizes the in vivo and in vitro RNA secondary structure of the SARS-CoV-2 genome using icSHAPE, generating nucleotide-level reactivity scores across the ~30-kb genome. Quality control was assessed via Pearson correlation between library replicates, and structural accuracy was evaluated with ROC-AUC against known reference RNA structures. Structurally variable regions between in vivo and in vitro conditions were called using a binomial test and a permutation test, and structural model accuracy against a published theoretical model was quantified with sensitivity and positive predictive value (PPV).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson correlation coefficient | Inter-replicate quality control of icSHAPE libraries: RNA expression (RPKM) of host transcriptome and RT-stop sites on SARS-CoV-2 genome | ~150 million reads per replicate; exact replicate count not stated in excerpt | not stated |
| Area under the receiver operating characteristic curve (AUC/ROC) | Validation of icSHAPE reactivity scores against known reference structures (18S rRNA, 28S rRNA, SRP RNA) and against Rangan et al. theoretical models for SARS-CoV-2 5'UTR and 3'UTR | null | not stated |
| Sensitivity and positive predictive value (PPV) | Quantitative comparison of in vivo icSHAPE-derived structural model versus Rangan et al. theoretical model for SARS-CoV-2 5'UTR and 3'UTR | null | na |
| Binomial test | Identification of structurally variable regions (SVRs) between in vivo and in vitro icSHAPE profiles across the SARS-CoV-2 genome | null | not stated |
| Permutation test | Identification of structurally variable regions between in vivo and in vitro icSHAPE profiles, used in conjunction with the binomial test | null | not stated |
-
Structurally variable regions between in vivo and in vitro conditions were identified with a binomial test and permutation test, and 371 regions were reported as a count↳ Could also: A Benjamini-Hochberg false discovery rate (FDR) correction applied to the per-nucleotide test p-values could also be used to define the call set — With tens of thousands of positions tested genome-wide, an FDR threshold would quantify the expected proportion of false positives among the 371 calls and allow readers to calibrate confidence in individual calls
-
Pearson correlation coefficients were used to assess inter-replicate reproducibility of RPKM values and RT-stop counts↳ Could also: Spearman rank correlation or intraclass correlation coefficient (ICC) could also be used for the same purpose — Spearman correlation is robust to non-normality and outliers common in sequencing count distributions; ICC additionally captures absolute agreement and is widely used for assay reproducibility assessment
-
Structural model accuracy was summarized with AUC from a ROC curve against reference structures↳ Could also: Precision-recall AUC (PR-AUC) or Matthews correlation coefficient (MCC) could also be reported alongside ROC-AUC — When base-paired and single-stranded nucleotide classes are imbalanced (as is common in structured RNA), PR-AUC and MCC can provide a more complete picture of classifier performance than ROC-AUC alone
-
Sensitivity and PPV were reported as single point estimates when comparing icSHAPE-derived models to the Rangan et al. reference model↳ Could also: Bootstrap confidence intervals around sensitivity and PPV could also be reported — Confidence intervals would convey the uncertainty in these performance metrics, which is particularly useful given that the reference model itself carries inherent uncertainty
-
Covariation evidence was assessed using a covariation score derived from Infernal alignments across deduplicated coronavirus genomes↳ Could also: R-scape (RNA Structural Covariation Above Phylogenetic Expectation) could also be applied for the same validation — R-scape explicitly models the null distribution expected from phylogenetic relationships alone and computes a p-value per base pair, providing a significance-based framework for calling co-variant pairs rather than a score threshold
-
No dispersion measures (SD, SEM, or CI) were reported for the quantitative metrics presented (AUC, Pearson r, sensitivity, PPV)↳ Could also: Bootstrap or jackknife confidence intervals could be reported for each performance metric — Interval estimates around point metrics would allow readers to assess whether, for example, differences in AUC between 5'UTR (0.854) and 3'UTR (0.692) are statistically distinguishable given sampling variability
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33636127
Paper: Sun, Li, Ju et al. (2021) In vivo structural characterization of the SARS-CoV-2 RNA genome identifies host proteins vulnerable to repurposed drugs. Cell 184(7):1865-1883. DOI 10.1016/j.cell.2021.02.008.
Assay: icSHAPE (in vivo click SHAPE) of viral RNA in infected cells & in vitro. Data: GEO GSE153984 / BioProject PRJNA644648 / SRA SRP270837 — 22 SRA runs, 11 biosamples, HiSeq X Ten, library strategy OTHER/TRANSCRIPTOMIC. Code:
- Authors' analysis scripts: https://github.com/lipan6461188/SARS-CoV-2 (master, pushed 2020-10-01, no license)
- icSHAPE quantification: icSHAPE-pipe (Tsinghua gonglab) — STAR-based
- Aligners listed: STAR, Bowtie2; trimming Trimmomatic.
Data layout (key runs)
SARS-CoV-2 in-vivo genome structome = the huh7 samples:
- SARS2-huh7-D (DMSO background): SRR12169246 (45.7M), SRR12169247 (57.2M)
- SARS2-huh7-N (NAI-N3 in vivo) : SRR12169248 (163.8M), SRR12169249 (206.0M)
- SARS2-huh7-T (NAI-N3 in vitro): SRR12169250 (182.0M), SRR12169251 (160.5M) Other coronaviruses (MERS-HKU9, SARS-HKU5, NL63-HKU1, SARS2-T tissue) = smaller libs.
In scope (pipeline-derived, attemptable)
| id | reported result | paper loc | pipeline |
|---|---|---|---|
| C1 | icSHAPE scores 0-1 for SARS-CoV-2 genome; coverage >99.88% of nt | Results/STAR Methods | icSHAPE-pipe (STAR) |
| C2 | ~150M reads per library replicate (in vivo) | STAR Methods | read counting |
| C3 | inter-replicate Pearson r > 0.99 (viral icSHAPE) | Results | icSHAPE-pipe + corr |
| C4 | 37 conserved RNA structural elements | Results/Fig | RNAstructure + Infernal (downstream) |
| C5 | 371 structurally variable regions (whole-genome) | Results | downstream analysis |
| C6 | 170 co-variant pairs (6 in 5'UTR, 12 in 3'UTR) | Results | Infernal/R-scape |
| C7 | PrismNet: 42 host proteins (31 5'UTR, 34 3'UTR) | Results | PrismNet deep model |
Primary target (80% floor): C1, C2, C3 — recompute icSHAPE reactivity for the SARS-CoV-2 genome from raw reads and check coverage + inter-replicate correlation. These are the directly pipeline-derived, low-ambiguity outputs.
Out of scope (not pipeline / wet-lab / heavy bespoke)
- Drug efficacy (Nilotinib/Sorafenib/Deguelin), ASO knockdowns, pull-down validation — wet-lab.
- Full covariation/structure-element catalog (C4-C6) and PrismNet (C7) are downstream of C1 and require large bespoke multi-genome alignments / a trained DL model — attempted only if C1-C3 land and time permits; flagged as stretch.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.