Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

In vivo structural characterization of the SARS-CoV-2 RNA genome identifies host proteins vulnerable to repurposed drugs.

Cell · 2021
L1 95/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED. icSHAPE in-vivo structure characterization of the SARS-CoV-2 genome (Sun et al., Cell 2021) from GSE153984, via STAR + icSHAPE-pipe-equivalent processing on «our HPC». All three in-scope pipeline-derived claims land: C1 genome icSHAPE coverage = 99.883% (matches >99.88%) -- confirmed BOTH from the authors' shipped virus-w50.shape AND independently regenerated from raw fastq (STAR->pysam RT-stop pileup, coverage>=20x = 99.883%, JOB B); C2 read depth mean 136M across 6 huh7 libs (~within-tol of 'about 150M'; in-vivo NAI reps avg 185M; JOB B trim counts matched ENA read_count exactly); C3 in-vivo inter-replicate RT-stop Pearson r = 0.9983 > 0.99 (exact, from raw regeneration; base-density r=0.9995; shipped reactivity proxy r=0.9924). No fabrication signal: the exact 99.883% coverage recurs independently from raw reads. C4-C7 (conserved elements, variable regions, covariant pairs, PrismNet host RBPs) are downstream / deep-learning / wet-lab and out of scope (not attempted). Dataset GSE153984 profiled: complete open deposit, 22 runs = 11 conditions x2 reps, delivers-promised, quality B (processed .shape profiles split into the authors' GitHub rather than GEO). Grades are provisional pending human sign-off in AUDIT.md.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ 921338406c3c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether determining the in vivo (and in vitro) secondary structure of the SARS-CoV-2 RNA genome can reveal functional structural elements and host RNA-binding proteins that regulate viral infection, thereby nominating druggable targets among FDA-approved drugs.

Core claims
  • icSHAPE was used to determine the in vivo and in vitro structural landscape of the SARS-CoV-2 RNA genome in infected Huh7.5.1 cells, plus UTR structures of six other coronaviruses method
  • The in vivo structural models validate several structural elements previously predicted in silico (e.g., Rangan et al. 5'UTR/3'UTR models) while revealing notable differences finding
  • Discovered structural features that affect the translation and abundance of subgenomic viral RNAs in cells finding
  • A deep-learning tool informed by the structural data predicted 42 host proteins that bind SARS-CoV-2 RNA resource
  • Antisense oligonucleotides (ASOs) targeting viral RNA structural elements dramatically reduced SARS-CoV-2 infection in liver- and lung-tumor-derived cells finding
  • FDA-approved drugs inhibiting the predicted SARS-CoV-2 RNA-binding proteins dramatically reduced SARS-CoV-2 infection in cells finding
  • The in vivo SARS-CoV-2 RNA structure differs substantially from in vitro refolded and purely theoretical structures, with the viral genome more single-stranded in vivo finding
  • Covariation analysis across coronavirus genomes identified 170 co-variant base pairs, including a novel duplex between the 3'UTR and ORF10 finding
Experimental setups
Assay System Perturbation Readout Platform
icSHAPE (in vivo RNA structure probing) Huh7.5.1 human liver cancer cells infected with SARS-CoV-2 NAI-N3 chemical probing icSHAPE reactivity score per nucleotide (single-stranded vs base-paired) deep sequencing / icSHAPE-pipe
icSHAPE (in vitro RNA structure probing) SARS-CoV-2 RNA purified from infected Huh7.5.1 cells, refolded in vitro NAI-N3 chemical probing icSHAPE reactivity score per nucleotide deep sequencing / icSHAPE-pipe
icSHAPE on in vitro transcribed UTRs UTRs of SARS-CoV-2 (reference and mutant) and six other coronaviruses (e.g., SARS-CoV, MERS-CoV) NAI-N3 chemical probing after in vitro transcription/refolding icSHAPE reactivity score, comparative UTR structure deep sequencing
Structure validation against reference RNAs 18S rRNA, 28S rRNA, SRP RNA in Huh7.5.1 cells none AUC of icSHAPE scores vs known reference structures
Deep-learning RBP-RNA binding prediction SARS-CoV-2 RNA (UTRs) structural/sequence data none (computational) predicted host RNA-binding proteins (42 candidates) neural network model (Sun et al., 2021)
SARS-CoV-2 infection assay with antisense oligonucleotide treatment cells derived from human liver and lung tumors ASOs targeting RNA structural elements SARS-CoV-2 infection level
SARS-CoV-2 infection assay with drug treatment cells derived from human liver and lung tumors FDA-approved drugs inhibiting predicted RNA-binding proteins SARS-CoV-2 infection level
Covariation/phylogenetic analysis deduplicated coronavirus genome sequences none (computational) co-variant base pairs supporting structural models Infernal package
Key results
  • icSHAPE scores obtained for more than 99.88% of nucleotides of the in vivo SARS-CoV-2 RNA genome 99.88%
  • Inter-replicate correlation very high: RPKM correlation >0.98, RT-stop correlation >0.99 r>0.98 / r>0.99
  • High AUC for icSHAPE scores fitting known reference structures of 18S rRNA, 28S rRNA, and SRP RNA AUC=0.813 (18S), 0.804 (28S), 0.730 (SRP)
  • icSHAPE-derived structure model agreement with Rangan et al. theoretical models AUC=0.854 (5'UTR), 0.692 (3'UTR); sensitivity 0.945/0.913, PPV 1.0/0.824
  • In vivo and in vitro structural profiles of the SARS-CoV-2 genome are only moderately correlated, with 371 structurally variable regions identified genome-wide r=0.58; 371 regions
  • 170 co-variant base pairs identified across coronavirus genomes, including 6 in the 5'UTR, 12 in the 3'UTR, and 5 in a novel 3'UTR-ORF10 duplex 170 pairs (6+12+5 subsets)
  • SARS-CoV-2 RNA genome appeared more single-stranded in vivo than in vitro
  • Deep-learning tool predicted 42 host proteins binding SARS-CoV-2 RNA, several validated physically/functionally 42 proteins
Key statistics
  • correlation >0.98 (inter-replicate correlation of host transcriptome RPKM expression)
  • correlation >0.99 (inter-replicate correlation of RT-stop from NAI-N3 modification on viral RNA genome)
  • other AUC=0.813, 0.804, 0.730 (AUC of icSHAPE scores vs reference structures for 18S rRNA, 28S rRNA, SRP RNA)
  • other AUC=0.854 (5'UTR), 0.692 (3'UTR) (AUC of icSHAPE scores vs Rangan et al. theoretical UTR models)
  • other sensitivity 0.945/0.913, PPV 1.0/0.824 (sensitivity and PPV of in vivo structure model vs Rangan's model for 5'UTR/3'UTR)
  • correlation 0.58 (Pearson correlation between in vitro and in vivo icSHAPE structural profiles)
  • count 371 (structurally variable regions between in vivo and in vitro data across the whole genome)
  • count 170 (co-variant base pairs identified across coronavirus genomes (6 in 5'UTR, 12 in 3'UTR, 5 in 3'UTR-ORF10 duplex))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper characterizes the in vivo and in vitro RNA secondary structure of the SARS-CoV-2 genome using icSHAPE, generating nucleotide-level reactivity scores across the ~30-kb genome. Quality control was assessed via Pearson correlation between library replicates, and structural accuracy was evaluated with ROC-AUC against known reference RNA structures. Structurally variable regions between in vivo and in vitro conditions were called using a binomial test and a permutation test, and structural model accuracy against a published theoretical model was quantified with sensitivity and positive predictive value (PPV).

Replicationunclear Sample size~150 million reads per library replicate stated for sequencing depth; number of biological or technical replicates and sample sizes for cell infection assays not specified in this excerpt Groupsin vivo SARS-CoV-2 RNA structure vs. in vitro refolded viral RNA; SARS-CoV-2 and six additional coronavirus UTR structures Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Pearson correlation coefficient Inter-replicate quality control of icSHAPE libraries: RNA expression (RPKM) of host transcriptome and RT-stop sites on SARS-CoV-2 genome ~150 million reads per replicate; exact replicate count not stated in excerpt not stated
Area under the receiver operating characteristic curve (AUC/ROC) Validation of icSHAPE reactivity scores against known reference structures (18S rRNA, 28S rRNA, SRP RNA) and against Rangan et al. theoretical models for SARS-CoV-2 5'UTR and 3'UTR null not stated
Sensitivity and positive predictive value (PPV) Quantitative comparison of in vivo icSHAPE-derived structural model versus Rangan et al. theoretical model for SARS-CoV-2 5'UTR and 3'UTR null na
Binomial test Identification of structurally variable regions (SVRs) between in vivo and in vitro icSHAPE profiles across the SARS-CoV-2 genome null not stated
Permutation test Identification of structurally variable regions between in vivo and in vitro icSHAPE profiles, used in conjunction with the binomial test null not stated
Approaches that could also have been used
  • Structurally variable regions between in vivo and in vitro conditions were identified with a binomial test and permutation test, and 371 regions were reported as a count
    Could also: A Benjamini-Hochberg false discovery rate (FDR) correction applied to the per-nucleotide test p-values could also be used to define the call set — With tens of thousands of positions tested genome-wide, an FDR threshold would quantify the expected proportion of false positives among the 371 calls and allow readers to calibrate confidence in individual calls
  • Pearson correlation coefficients were used to assess inter-replicate reproducibility of RPKM values and RT-stop counts
    Could also: Spearman rank correlation or intraclass correlation coefficient (ICC) could also be used for the same purpose — Spearman correlation is robust to non-normality and outliers common in sequencing count distributions; ICC additionally captures absolute agreement and is widely used for assay reproducibility assessment
  • Structural model accuracy was summarized with AUC from a ROC curve against reference structures
    Could also: Precision-recall AUC (PR-AUC) or Matthews correlation coefficient (MCC) could also be reported alongside ROC-AUC — When base-paired and single-stranded nucleotide classes are imbalanced (as is common in structured RNA), PR-AUC and MCC can provide a more complete picture of classifier performance than ROC-AUC alone
  • Sensitivity and PPV were reported as single point estimates when comparing icSHAPE-derived models to the Rangan et al. reference model
    Could also: Bootstrap confidence intervals around sensitivity and PPV could also be reported — Confidence intervals would convey the uncertainty in these performance metrics, which is particularly useful given that the reference model itself carries inherent uncertainty
  • Covariation evidence was assessed using a covariation score derived from Infernal alignments across deduplicated coronavirus genomes
    Could also: R-scape (RNA Structural Covariation Above Phylogenetic Expectation) could also be applied for the same validation — R-scape explicitly models the null distribution expected from phylogenetic relationships alone and computes a p-value per base pair, providing a significance-based framework for calling co-variant pairs rather than a score threshold
  • No dispersion measures (SD, SEM, or CI) were reported for the quantitative metrics presented (AUC, Pearson r, sensitivity, PPV)
    Could also: Bootstrap or jackknife confidence intervals could be reported for each performance metric — Interval estimates around point metrics would allow readers to assess whether, for example, differences in AUC between 5'UTR (0.854) and 3'UTR (0.692) are statistically distinguishable given sampling variability
Software: icSHAPE-pipe · RNAstructure · Infernal · in-house deep-learning tool (Sun et al. 2021)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33636127

Paper: Sun, Li, Ju et al. (2021) In vivo structural characterization of the SARS-CoV-2 RNA genome identifies host proteins vulnerable to repurposed drugs. Cell 184(7):1865-1883. DOI 10.1016/j.cell.2021.02.008.

Assay: icSHAPE (in vivo click SHAPE) of viral RNA in infected cells & in vitro. Data: GEO GSE153984 / BioProject PRJNA644648 / SRA SRP270837 — 22 SRA runs, 11 biosamples, HiSeq X Ten, library strategy OTHER/TRANSCRIPTOMIC. Code:

  • Authors' analysis scripts: https://github.com/lipan6461188/SARS-CoV-2 (master, pushed 2020-10-01, no license)
  • icSHAPE quantification: icSHAPE-pipe (Tsinghua gonglab) — STAR-based
  • Aligners listed: STAR, Bowtie2; trimming Trimmomatic.

Data layout (key runs)

SARS-CoV-2 in-vivo genome structome = the huh7 samples:

  • SARS2-huh7-D (DMSO background): SRR12169246 (45.7M), SRR12169247 (57.2M)
  • SARS2-huh7-N (NAI-N3 in vivo) : SRR12169248 (163.8M), SRR12169249 (206.0M)
  • SARS2-huh7-T (NAI-N3 in vitro): SRR12169250 (182.0M), SRR12169251 (160.5M) Other coronaviruses (MERS-HKU9, SARS-HKU5, NL63-HKU1, SARS2-T tissue) = smaller libs.

In scope (pipeline-derived, attemptable)

id reported result paper loc pipeline
C1 icSHAPE scores 0-1 for SARS-CoV-2 genome; coverage >99.88% of nt Results/STAR Methods icSHAPE-pipe (STAR)
C2 ~150M reads per library replicate (in vivo) STAR Methods read counting
C3 inter-replicate Pearson r > 0.99 (viral icSHAPE) Results icSHAPE-pipe + corr
C4 37 conserved RNA structural elements Results/Fig RNAstructure + Infernal (downstream)
C5 371 structurally variable regions (whole-genome) Results downstream analysis
C6 170 co-variant pairs (6 in 5'UTR, 12 in 3'UTR) Results Infernal/R-scape
C7 PrismNet: 42 host proteins (31 5'UTR, 34 3'UTR) Results PrismNet deep model

Primary target (80% floor): C1, C2, C3 — recompute icSHAPE reactivity for the SARS-CoV-2 genome from raw reads and check coverage + inter-replicate correlation. These are the directly pipeline-derived, low-ambiguity outputs.

Out of scope (not pipeline / wet-lab / heavy bespoke)

  • Drug efficacy (Nilotinib/Sorafenib/Deguelin), ASO knockdowns, pull-down validation — wet-lab.
  • Full covariation/structure-element catalog (C4-C6) and PrismNet (C7) are downstream of C1 and require large bespoke multi-genome alignments / a trained DL model — attempted only if C1-C3 land and time permits; flagged as stretch.
Figures / tables: Fig 2
C1
Reported
>99.88% of SARS-CoV-2 nt have in-vivo icSHAPE scores
Reproduced
99.883% (29868/29903 nt non-NULL, NC_045512.2) from authors' shipped virus-w50.shape AND 99.883% independently from raw reads (JOB B, coverage >=20x)
exact
C2
Reported
average ~150M reads per library replicate
Reproduced
mean of 6 huh7 libs = 135.9M (D=45.7/57.2M, N=163.8/206.0M, T=182.0/160.5M); JOB B trim counts matched ENA exactly
within tolerance
C3
Reported
inter-replicate correlation of viral RT-stop exceeded 0.99
Reproduced
in-vivo N1 vs N2 RT-stop Pearson r = 0.9983 (raw regen, JOB B); base-density r = 0.9995; shipped reactivity-level r = 0.9924
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 95/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

242 k
tokens (I/O) · 15.2 M incl. cache
163 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.