Betacoronavirus-specific alternate splicing.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓Overall, the reproduction was clean
- 🟡The central claim did not (fully) hold under reproduction
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (Tier-1, exact). Pipeline STAR2.7.5c->RSEM1.3.3->HBA-DEALS. The paper's central result is Table 2 (per-dataset DGE+DAS gene counts) underpinning the 'betacoronavirus-specific alternate splicing' thesis. The analysis repo (TheJacksonLaboratory/covid19splicing) ships the final HBA-DEALS posterior tables (HBA-DEALS-covid19_output/*.txt) plus the FDR code (get_fdr_prob.R). Reimplementing their own get.fdr.prob (Bayesian FDR=0.05) in Python and counting DAS/DGE genes reproduces ALL 32 Table-2 values (16 datasets x DGE+DAS) EXACTLY. This proves Table 2 is fully derivable from the deposit and not fabricated. STILL IN PROGRESS: (a) Figures 4 (intron-retention enrichment p=0.016) and 5 (ribosomal-gene enrichment p=0.033/0.004) recomputation from the same outputs; (b) Tier-3 full upstream re-run (SRA->STAR->RSEM->HBA-DEALS MCMC) for one small betacorona dataset on «our HPC», which is stochastic (no seed) so expected to be close-not-identical. «our HPC»/VPN was intermittently unreachable during this work; Tier-1 needed no heavy compute. NOT attempted: wet-lab/external-database claims (out of scope).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-19 ⛓ 44e746919149
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-19
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusBetacoronaviruses (SARS-CoV-2, SARS-CoV, MERS) induce characteristic patterns of differential alternative splicing in host cells; the study tests whether genes differentially spliced during betacoronavirus infection share a distinct functional profile and structural/sequence features that differ from infection by other pathogens.
- ★ Genes showing differential alternative splicing in SARS-CoV-2 have a similar functional profile to those in SARS-CoV and MERS, affecting a diverse set of genes and biological functions related to virus biology. finding
- ★ Differentially spliced transcripts in betacoronavirus-infected cells are more likely to undergo intron retention, contain a pseudouridine modification, and have fewer exons than controls. finding
- ★ Viral load in clinical COVID-19 samples correlates with isoform distribution of differentially spliced genes. finding
- ★ A significantly higher number of ribosomal genes are affected by differential alternative splicing and gene expression in betacoronavirus samples. finding
- ★ Betacoronavirus differentially spliced genes are depleted for binding sites of RNA-binding proteins. finding
- ★ Betacoronavirus datasets show higher GO enrichment for 'translation' and 'RNA binding' among DAS genes than non-betacoronavirus datasets. finding
- A reproducible pipeline (fastp/STAR/RSEM) combined with HBA-DEALS Bayesian analysis quantifies differential expression and differential splicing across 17 datasets. method
- A gene ranking score quantifies the degree to which genes are more alternatively spliced in betacoronavirus samples than in control samples (Supplemental File S2). resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (differential expression and splicing analysis) | MRC5 cells infected with SARS-CoV (24 h) and MERS (24 h, 48 h) | viral infection vs mock | isoform/gene counts; differentially expressed (DGE) and differentially spliced (DAS) genes | STAR (v2.7.3a) alignment + RSEM (v1.3.1) quantification; HBA-DEALS |
| bulk RNA-seq | Calu-3 cells infected with MERS (24 h) | viral infection vs mock | DGE/DAS gene counts | snakemake pipeline (fastp, STAR, RSEM) |
| clinical RNA-seq | nasopharyngeal swab specimens (SARS-CoV-2-A), New York-Presbyterian/Weill Cornell | SARS-CoV-2 infection (100 infected vs 193 mock) | DGE/DAS; viral load vs isoform distribution correlation | COV-IRT pipeline: Trimmomatic v0.39, Kraken2, STAR v2.7.3a, RSEM v1.3.1; QIAsymphony/DSP Virus/Pathogen Mini Kit extraction |
| bulk RNA-seq | SARS-CoV-2-infected lung (SARS-CoV-2-B) and Nasal ECs (SARS-CoV-2-C) | viral infection vs mock | DGE/DAS gene counts | snakemake pipeline (fastp, STAR, RSEM) |
| bulk RNA-seq (transfection) | HEK293 cells | SARS-CoV-2 NSP1 and NSP2 transfection (24 h) | DGE/DAS gene counts | snakemake pipeline (fastp, STAR, RSEM) |
| bulk RNA-seq (control pathogens) | Huh7.5.1 (HCV), A549 (Zika, DENV), hNEC/lung (H3N2), Nasal ECs (Strep), Nasal scrape (RSV) | infection with non-betacoronavirus pathogens vs mock | DGE/DAS gene counts for comparison | snakemake pipeline (fastp, STAR, RSEM) |
| GO term enrichment analysis | all 17 datasets (computational) | none | enriched GO terms (BH-corrected p<0.05) among DAS/DGE genes | Ontologizer/phenol Java library (term-for-term, Benjamini-Hochberg) |
| transcript biotype / exon / RNA modification feature analysis | differentially spliced isoforms (computational) | none | proportion of retained-intron isoforms, mean number of exons, pseudouridine modification, RBP binding-site enrichment | Biomart (Ensembl transcript_biotype); GTF Homo_sapiens.GRCh38.91/100 |
- ▲ Of 1044 GO terms enriched in at least two datasets, 1025 showed high enrichment in at least one betacoronavirus dataset with respect to DAS 1025/1044 terms
- ▲ Differentially spliced transcripts in betacoronavirus infection more often undergo intron retention, contain pseudouridine, and have fewer exons than controls
- – Viral load in clinical COVID-19 samples correlated with isoform distribution of differentially spliced genes
- ▲ Higher number of ribosomal genes affected by DAS and DGE in betacoronavirus samples
- ▼ Betacoronavirus differentially spliced genes are depleted for RNA-binding protein binding sites
- ▲ Betacoronaviruses have higher enrichment scores for 'translation' (GO:0006412) and 'RNA binding' (GO:0003723) among DAS genes than non-betacoronaviruses
- ▲ MERS-48h dataset had the largest number of differentially expressed (7376) and differentially spliced (7394) genes 7376 DGE / 7394 DAS
- – 130 of the enriched terms had a median score below threshold 130 terms
- count 1044 GO terms enriched (BH p<0.05) in at least two datasets (GO enrichment across datasets)
- count 1025 terms highly enriched in at least one betacoronavirus dataset for DAS (betacoronavirus DAS enrichment)
- pvalue BH-corrected p < 0.05 (GO term enrichment / DAS enrichment threshold)
- other FDR threshold of 0.05; exclude genes/isoforms with probability of no-effect > 0.25 (HBA-DEALS differential analysis cutoffs)
- count SARS-CoV-2-A clinical dataset: 100 infected / 193 mock samples (clinical nasopharyngeal swab cohort)
- count At least 10 SARS-CoV-2 proteins bind 6 structural non-coding RNAs and 142 mRNAs (background on SARS-CoV-2 RNA binding)
- other ROPE: changes of 0.1 in log-expression and fold changes of 0.2 or less in isoform proportion (Region of Practical Equivalence in HBA-DEALS)
- count SARS dataset: 1218 DGE, 1441 DAS (MRC5 SARS-CoV 24h differential genes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper performed a multi-dataset survey of differential gene expression (DGE) and differential alternative splicing (DAS) across 16–17 RNA-seq experiments using the Bayesian hierarchical model HBA-DEALS, which estimates posterior probabilities of DAS/DGE per gene via MCMC sampling in Stan and controls the false discovery rate at 0.05 via the mean of posterior probabilities of no effect. Gene Ontology enrichment was conducted using a term-for-term approach with Benjamini-Hochberg correction (FDR ≤ 0.05). Cross-dataset comparisons were made descriptively using GO enrichment profiles, intron-retention proportions, exon-count summaries, and a custom additive gene-ranking score; a correlation between viral load and isoform distribution was also reported for the clinical SARS-CoV-2-A dataset.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| HBA-DEALS Bayesian hierarchical model (L-BFGS optimization + MCMC via Stan R interface) for differential gene expression and differential alternative splicing | All 16 RNA-seq datasets; DGE and DAS estimated simultaneously per dataset (counts in Table 2) | Varies by dataset: e.g., SARS 6 infected / 3 mock; MERS-24h 3/3; SARS-CoV-2-A 100 infected / 193 controls (Table 1) | not stated |
| Term-for-term Gene Ontology enrichment analysis with Benjamini-Hochberg correction | DAS gene sets from each dataset; 1044 GO terms enriched at BH-FDR ≤ 0.05 in ≥2 datasets summarized in Figs. 1–2 and Supplemental Figures S1–S3 | — | not stated |
| Correlation analysis (method not specified in text provided) between viral load and isoform distribution | Clinical nasopharyngeal swab dataset SARS-CoV-2-A (100 infected / 193 controls) | 100 infected samples stated | not stated |
| Custom additive gene-ranking score: sum of DAS posterior probabilities across betacoronavirus datasets minus sum across other-pathogen datasets | Betacoronavirus-specific DAS gene ranking (Supplemental File S2) | — | na |
| Proportion calculation (retained_intron biotype count / total biotype count from Ensembl Biomart) compared across datasets | Intron-retention analysis across all datasets | — | not stated |
-
Differential alternative splicing was detected using HBA-DEALS, a Bayesian hierarchical model that produces per-gene posterior probabilities of DAS without resolving individual splicing events↳ Could also: Event-level frequentist tools such as rMATS, DEXSeq, or MAJIQ could also quantify DAS at the level of specific exon-skipping, intron-retention, or splice-site usage events — Event-level methods provide mechanistic resolution (which exons or junctions change) and are widely used in the field, facilitating direct comparison with other published DAS studies and enabling downstream validation at the event level
-
GO enrichment used a term-for-term approach that tests each GO term independently after BH correction↳ Could also: Gene Set Enrichment Analysis (GSEA) or a parent-child GO method (e.g., Ontologizer parent-child union) could also be used — GSEA ranks all genes by a continuous score (e.g., DAS probability) rather than applying a binary threshold, potentially improving sensitivity for pathway-level signals; parent-child methods reduce inflation of enrichment due to GO term dependency structure
-
Cross-dataset comparisons were made descriptively by visually comparing enrichment profiles, proportions, and GO-term scores across independently analyzed datasets↳ Could also: A formal meta-analysis framework such as Fisher's combined p-value, a random-effects model, or cross-study Bayesian pooling could also integrate evidence across datasets — Meta-analytic methods provide a unified test statistic for consistency of DAS patterns across datasets and can explicitly model between-study heterogeneity arising from different cell types, experimental designs, and sequencing depths
-
The proportion of intron-retention isoforms among DAS transcripts was reported as a simple fraction for each dataset without a formal statistical comparison between dataset groups↳ Could also: A Fisher's exact test, chi-squared test, or logistic regression comparing intron-retention proportions between betacoronavirus and non-betacoronavirus datasets could also be applied — A formal test would provide a p-value and confidence interval for the between-group difference in intron-retention proportions, quantifying uncertainty and allowing assessment of statistical significance of the observed pattern
-
No measures of dispersion (SD, SEM, IQR, CI) were reported around summary statistics such as mean exon counts or mean DAS gene counts across datasets↳ Could also: Reporting SD or 95% confidence intervals around group-level summaries (e.g., mean number of DAS genes across betacoronavirus vs. other-pathogen datasets) could also be used — Dispersion measures would allow readers to assess within-group variability and the consistency of effects across heterogeneous datasets, and are especially informative given the wide range of sample sizes (n=3 to n=193) across experiments
-
Statistical power was not formally analyzed; datasets with as few as three samples per group were included alongside datasets with up to 193 samples↳ Could also: A post-hoc sensitivity analysis (e.g., simulation-based power curves or rarefaction subsampling) could also characterize the minimum detectable effect size for DAS given each dataset's sample size and sequencing depth — Such analysis would help contextualize low DAS counts in small datasets (e.g., HCV with 3/3 samples yielding 113 DAS genes vs. MERS-48h with 3/3 yielding 7394) as potentially reflecting power differences rather than true biological differences in splicing magnitude
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
All 32 DGE+DAS gene counts in Table 2 were reproduced exactly by re-deriving them from the authors' deposited HBA-DEALS output tables using their own get.fdr.prob() at FDR=0.05 — strong evidence the table is genuine and fully derivable, with no fabrication concern (q5 green). The deviation is effectively nil, so q3–q6 are green. The caveat is scope: this is a derivation from the authors' shipped intermediate tables, not a from-raw SRA→STAR→RSEM→HBA-DEALS rerun, and the enrichment figures (Fig 4 p=0.016, Fig 5 p=0.033/0.004) that actually establish specificity were not recomputed. Since Table 2 alone shows comparable DAS in non-beta viruses (DENV 1530, Zika 1103), the central 'betacoronavirus-specific' conclusion is only partially verified (q7 yellow), even though the reproduced numbers are 1:1 (q8 green).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.