Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Long-Read RNA Sequencing Identifies Polyadenylation Elongation and Differential Transcript Usage of Host Transcripts During SARS-CoV-2 In Vitro Infection

Front Immunol · 2022
L1 90/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 79% of all assessed papers rank 211 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED. Target = Table 1 (proportion of average viral reads at 24 hpi) across 4 cell lines x 3 ONT protocols = all 12 cells. Method: minimap2 v2.17 (paper's exact version; direct-RNA '-ax splice -uf -k14 --secondary=no', cDNA '-ax splice --secondary=no') vs combined host (Ensembl r100: human GRCh38 / green-monkey ChlSab1.1) + SARS-CoV-2 MT007544.1; viral reads = primary alignments on the viral contig, percentage of mapped reads (matches the paper denominator: Vero direct-RNA 74.4% vs reported 74%). 28 runs processed; observed read counts match the ENA run table 28/28. Result: 4 cells exact, 8 within-tol, overall reproduced, with the full biological pattern intact (Vero very high 45-74%, Calu-3/Caco-2 ~2-4%, A549 negligible ~0.01%). NOT attempted: Table 2 DESeq2 host DE and Table 5 DRIMSeq DTU (heavier host-pipeline stretches). NOT reproducible: poly(A) tail-length (Fig S5) — needs raw fast5 not in the deposit. Code repo npTranscript (third-party/authors' tool) describes the same minimap2 mapping; reproduction done with standard minimap2+samtools per the Methods parameters.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 90
    assessed: 2026-06-21 ⛓ ec36c21d8357
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Long-read (Oxford Nanopore) RNA sequencing can reveal host transcriptomic features during SARS-CoV-2 infection — differential gene expression, polyadenylation, and isoform/transcript usage — that are not well captured by standard short-read RNA-seq, across multiple in vitro epithelial cell line models over an infection time-course.

Core claims
  • Differential polyadenylation occurs in infected Calu-3 and Vero cells at a late time point (48 hpi) finding
  • Poly(A) tail lengths show increased length in the majority of differentially polyadenylated transcripts in Calu-3 and Vero cells finding
  • GO terms 'viral transcription' and 'translation' are significantly enriched among differentially polyadenylated genes in Calu-3 infection data finding
  • Ribosomal protein genes RPS4X and RPS6 are downregulated in expression while also showing elongated poly(A) tails during infection finding
  • Differential transcript usage is identified in Caco-2, Calu-3 and Vero cells, including in genes such as GSDMB and KPNA2 previously implicated in SARS-CoV-2 infection finding
  • ONT long-read sequencing (direct RNA, direct cDNA, PCR cDNA) enables combined measurement of differential gene expression, poly(A) tail length, and transcript/isoform usage in host cells during viral infection method
  • Vero cells show the highest viral burden among the four cell lines tested, consistent with their defective interferon type I response finding
  • There is a statistically significant overlap between genes downregulated in expression and genes with elongated poly(A) tails in Calu-3 at 48 hpi finding
Experimental setups
Assay System Perturbation Readout Platform
Direct RNA sequencing (ONT) Vero, Calu-3, Caco-2, A549 epithelial cell lines SARS-CoV-2 infection (MOI 0.1) vs mock gene expression, poly(A) tail length, transcript isoform usage ONT MinION/GridION, SQK-RNA002 kit, R9.4.1 flow cell
Direct cDNA sequencing (ONT) Vero, Calu-3, Caco-2, A549 epithelial cell lines SARS-CoV-2 infection (MOI 0.1) vs mock differential gene/transcript expression, poly(A)/poly(T) tail length (tailfindr) ONT MinION/GridION, SQK-DCS109 kit with EXP-NBD104/EXP-NBD114 barcoding, R9.4.1 flow cell
PCR cDNA sequencing (ONT) Vero, Calu-3, Caco-2, A549 epithelial cell lines SARS-CoV-2 infection (MOI 0.1) vs mock differential transcript usage ONT GridION, SQK-PCS109 + SQK-PBK004 kits, R9.4.1 flow cell
Poly(A) tail length analysis (nanopolish) Calu-3, Caco-2, Vero direct RNA-seq data SARS-CoV-2 infection vs mock median poly(A) tail length per gene nanopolish v0.13.2
Poly(A)/poly(T) tail length analysis (tailfindr) Calu-3, Caco-2, Vero direct cDNA-seq data SARS-CoV-2 infection vs mock median poly(A)/poly(T) tail length per gene; linear mixed-effects model of infection effect tailfindr v0.1.0, lmerTest v3.1-3
Differential expression analysis (DESeq2) Calu-3, Caco-2, Vero, A549 direct cDNA-seq counts (Featurecounts) SARS-CoV-2 infection vs mock across time points (0, 2, 24, 48 hpi) differentially expressed genes/transcripts (padj < 0.05) DESeq2, Featurecounts v2.0.0
Differential transcript usage analysis (DRIMSeq/StageR) Calu-3, Caco-2, Vero, A549 transcriptome-mapped counts (Salmon) SARS-CoV-2 infection vs mock significant genes/transcripts with differential usage (padj < 0.05) Salmon v0.13.1, DRIMSeq v1.16.1, StageR v1.10.0
GO/KEGG pathway enrichment analysis genes differentially expressed and differentially polyadenylated (Calu-3, other cell lines) SARS-CoV-2 infection vs mock enriched GO biological process terms and KEGG pathways multiGO shiny-app, hypergeometric test
Key results
  • Poly(A) tail lengths increased by up to ~101 nt in mean length among differentially polyadenylated transcripts in Calu-3 up to ~101 nt, padj = 0.029
  • GO terms 'viral transcription' and 'translation' were significantly enriched in Calu-3 differential polyadenylation gene set
  • RPS4X and RPS6 expression was downregulated alongside increased poly(A) tail length
  • Differential transcript usage detected for genes including GSDMB and KPNA2 across Caco-2, Calu-3 and Vero cells
  • Vero cells had the highest proportion of viral reads (~45%) at 24 hpi among the four cell lines ~45% of mapped reads
  • Hypergeometric test showed significant overlap (n=2) between downregulated genes (m=253) and poly(A)-elongated genes (k=13) in Calu-3 48 hpi against background N=15,426 n=2 overlap
  • Differential polyadenylation events identified in both Calu-3 and Vero cells specifically at the late 48 hpi time point
Key statistics
  • fold_change up to ~101 nt increase in mean poly(A) length (differentially polyadenylated transcripts in Calu-3)
  • pvalue padj = 0.029 (significance of mean poly(A) length increase)
  • count m = 253 (downregulated genes (padj<0.05) in Calu-3 48 hpi, DESeq2)
  • count k = 13 (genes with elongated poly(A) tails in Calu-3 48 hpi)
  • count n = 2 (overlap between downregulated and poly(A)-elongated genes)
  • count N = 15,426 (background gene count used in hypergeometric test)
  • other padj < 0.05 (significance threshold applied across DE, DP and DTU analyses)
  • other ~45% of reads mapped to virus (Vero cell viral burden at 24 hpi)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This in vitro SARS-CoV-2 infection study used Oxford Nanopore long-read RNA-seq across four epithelial cell lines and four time points (0, 2, 24, 48 hpi). Differential gene and transcript expression was tested with DESeq2; differential polyadenylation was assessed first with a Wilcoxon rank-sum test on nanopolish poly(A) medians and then with a linear mixed-effects model (lmerTest) on log-transformed tailfindr poly(T) lengths; differential transcript usage was analyzed with DRIMSeq followed by StageR stage-wise testing; and gene ontology and pathway enrichment was evaluated by hypergeometric test. Benjamini-Hochberg FDR correction was applied uniformly, with a padj < 0.05 significance threshold throughout.

Replicationmixed Sample sizeThree wells per condition (infected or mock) per time point per cell line in 6-well plates; direct RNA libraries prepared from RNA pooled across all three replicate wells (no independent sequencing replicates for that modality); direct cDNA and PCR cDNA libraries prepared from each replicate separately GroupsSARS-CoV-2-infected vs mock-infected cells; 4 cell lines (Vero, Calu-3, Caco-2, A549); 4 time points (0, 2, 24, 48 hpi) Pairingunpaired Randomization/blindingstated Dispersionnone Exact p-valuesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR (via R p.adjust); stage-wise FDR via StageR for differential transcript usage
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test (negative-binomial GLM) Differential gene and transcript expression: infected vs mock-infected cells per time point (0, 2, 24, 48 hpi) and across time points via interaction terms, in all four cell lines using direct cDNA data 3 biological replicates per condition (infected or mock) per time point per cell line (direct cDNA libraries) not stated
Wilcoxon rank-sum test (test of ranks) Comparison of overall median poly(A) tail lengths between control and infected cells per cell line using nanopolish direct RNA data; restricted to Ensembl IDs with more than one entry not stated
Linear mixed-effects regression (lmerTest v3.1-3) Main gene-level differential polyadenylation analysis: log-transformed tailfindr poly(T) lengths from direct cDNA replicates, with infection status as fixed effect; genes with at least 6 observations included log-transformation applied to address right-skewed distribution of raw lengths; further assumptions not stated
Pearson product-moment correlation Method validation: median poly(A) vs poly(T) lengths per gene from tailfindr compared across 2, 24, and 48 hpi in Caco-2, Calu-3, and Vero datasets not stated
Spearman rank correlation (cor.test in R) Method cross-validation: tailfindr poly(T)/poly(A) median lengths vs nanopolish poly(A) median lengths per gene in Calu-3 48 hpi control and infected datasets not stated
Hypergeometric test (multiGO) GO biological term and KEGG pathway enrichment for differentially expressed and differentially polyadenylated genes; tested against all genes in GO annotation database v100 as background not stated
Hypergeometric test (phyper in R) Overlap between downregulated genes (DESeq2, Calu-3 48 hpi, m=253) and genes with elongated poly(A) tails (lmerTest, k=13); observed overlap n=2; background N=15,426 N=15,426 background genes; m=253 downregulated; k=13 elongated poly(A); n=2 overlapping not stated
Dirichlet-Multinomial model (DRIMSeq v1.16.1) with stage-wise FDR (StageR v1.10.0) Differential transcript usage between control and infected conditions per cell line and time point, using Salmon transcript-level counts; filtered by min_samps_gene_expr=6, min_samps_feature_expr=3, min_gene_expr=10, min_feature_expr=10 not stated
Approaches that could also have been used
  • Direct RNA-seq replicates were pooled before library preparation, yielding a single library per condition-timepoint for that sequencing modality
    Could also: Keep individual replicate RNA samples separate through library preparation and sequencing to maintain independent biological replicates for direct RNA data as well — Separate replicates allow formal statistical modeling of between-replicate variance; pooling collapses biological variability into a single observation, precluding variance estimation for that data type and limiting inferential power for direct RNA-based analyses
  • Differential gene expression was analyzed with DESeq2
    Could also: edgeR (quasi-likelihood F-test or exact test) or limma-voom could also be applied to RNA-seq count data from long-read platforms — All three methods are widely benchmarked for count-based RNA-seq; edgeR and limma-voom handle small sample sizes and dispersion estimation differently and may offer complementary sensitivity or specificity, particularly for the low and variable read depths typical of long-read sequencing
  • Differential transcript usage was tested with DRIMSeq followed by StageR stage-wise testing
    Could also: DEXSeq or satuRn could also be used for differential transcript usage from count data — DEXSeq is a well-established method for exon/transcript-level differential usage; satuRn was designed specifically for transcript-level differential usage and has been benchmarked favorably for smaller sample sizes, which is relevant here given n=3 replicates per group
  • Analyses were performed separately per cell line and per time point
    Could also: A multi-factor model incorporating cell line as a covariate or as an interaction term could also be used within a single analysis — A joint model would allow formal testing of cell-line-by-infection interaction effects and could improve power by sharing variance estimates across cell lines, at the cost of requiring more complex model specification and assumption of comparable dispersion across lines
  • GO and KEGG enrichment was evaluated with a hypergeometric over-representation test against a fixed gene-list background
    Could also: Gene Set Enrichment Analysis (GSEA) or rank-based methods such as fgsea or camera using the full ranked gene list could also be applied — Rank-based enrichment methods use all expressed genes ordered by effect size or test statistic, avoiding the threshold-dependency inherent in over-representation tests and may detect coordinated but individually sub-threshold gene-set signals
  • Pearson correlation was used to compare median poly(A) and poly(T) lengths across tools and conditions
    Could also: Spearman rank correlation (which the authors also used in one comparison) could be applied consistently to all method-comparison analyses — Poly(A) length distributions are right-skewed and may contain outliers; Spearman correlation is more robust to these features and does not assume a linear relationship, making it a natural complement or alternative to Pearson for this type of measurement-method comparison
Software: DESeq2 (R/Bioconductor) · lmerTest (R) 3.1-3 · DRIMSeq (R/Bioconductor) 1.16.1 · StageR (R/Bioconductor) 1.10.0 · nanopolish 0.13.2 · tailfindr (R) 0.1.0 · Salmon 0.13.1 · Minimap2 2.17 · featureCounts (Subread) 2.0.0 · Guppy (ONT) 3.5.2 / 3.2.8 · Samtools 1.9 · ggplot2 (R) 3.3.4 · multiGO (in-house Shiny app) · R stats package (p.adjust, cor.test, phyper, dhyper)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35464437

Paper: Chang et al. 2022, Long-Read RNA Sequencing Identifies Polyadenylation Elongation and Differential Transcript Usage of Host Transcripts During SARS-CoV-2 In Vitro Infection, Front Immunol. DOI 10.3389/fimmu.2022.832223. PMCID PMC9019466.

Code: https://github.com/lachlancoin/npTranscript (latest commit c749797b39042a1cbaa14a2f5096844e3cb81684, 2025-05-08). Java (CIGAR/cluster analysis on BAMs via japsa) + R visualization. Apache-2.0.

Data: SRA BioProject PRJNA675370 — 152 Oxford Nanopore runs (MinION+GridION), 4 cell lines (Vero, Calu-3, Caco-2, A549) × timepoints (0/2/24/48 hpi) × {mock control, SARS-CoV-2 infected} × replicates, across 3 protocols: direct RNA (20 runs), cDNA (84), cDNA_PCR (48). Sample identity is encoded in library_name (e.g. Vero_24_Infected_RNA). Supplementary processed values on Figshare: 17139995 (DE), 16841794 (polyA), 17140007 (DTU).

Pipelines named in Methods

step tool (version) params
basecall Guppy 3.5.2 (3.2.8 for Vero PCR cDNA)
align minimap2 2.17 dRNA -ax splice -uf -k14 --secondary=no; cDNA -ax splice --secondary=no; transcriptome -ax map-ont
sort/index samtools 1.9
gene counts (cDNA→genome) featureCounts 2.0.0
tx counts (cDNA→transcriptome) Salmon 0.13.1
diff expr DESeq2 padj<0.05
read→Ensembl linking npTranscript (in-house)
polyA tail nanopolish 0.13.2; tailfindr 0.1.0
polyA stats lmerTest 3.1-3
diff transcript usage DRIMSeq 1.16.1 + stageR 1.10.0 padj<0.05
GO/KEGG multiGO

In scope (pipeline-derived → attempt)

  • R1 — Viral read fraction at 24 hpi (Data S1, abstract/Results): % of reads mapping to virus per cell line: Vero ~45%, Caco-2 ~2.3%, Calu-3 ~2%, A549 <0.01%. Pipeline: minimap2 → count reads mapping to SARS-CoV-2 (NC_045512.2) / total. Cheapest, clearest 1:1 — primary target. Uses direct-RNA 24hpi infected runs SRR13083402 (A549), SRR13089336 (Calu-3), SRR13089344 (Vero), SRR13089348 (Caco-2).
  • R2 — DESeq2 differential expression counts (Table 2): e.g. Calu-3 48hpi 371 up / 253 down; Vero 48hpi 210 up / 290 down; Caco-2 48hpi 8 up / 17 down (padj<0.05). Pipeline: minimap2(cDNA)→featureCounts→DESeq2. Heavier (needs host genomes + replicates). Stretch.
  • R3 — Poly(A) tail length (Results/Fig S5): up to ~101 nt mean (padj=0.029); Calu-3 48hpi twelve genes sig. increased. Pipeline: nanopolish/tailfindr on dRNA. Heavy. Stretch.
  • R4 — Differential transcript usage (Table 5): Calu-3 48hpi 16 isoforms; Caco-2 10; Vero 2 (DRIMSeq+stageR, padj<0.05). Stretch.

Out of scope (wet-lab / manual / not pipeline)

  • Virus culture, MOI, TCID50, plaque assays, RNA extraction, library prep.
  • multiGO GO/KEGG enrichment narrative (hosted webtool, not reproducibly scripted here).
  • Figure aesthetics / raincloud plots.

Strategy

Reach R1 fast (single run per cell line, viral genome only — no host genome needed for the fraction). Then escalate to R2 (DE counts) which is the next clearest tabular number, then R3/R4 as compute allows. All heavy compute on «our HPC»/«infra» via SLURM; downloads on front node. Honest 1:1; approximate paper values → within-tol grading.

Figures / tables: Table
T1_Vero_RNA
Reported
74%
Reproduced
74.38% (of mapped) / 70.10% (of total)
exact
T1_Vero_cDNA
Reported
45%
Reproduced
45.38% / 43.75%
exact
T1_Vero_PCR
Reported
55%
Reproduced
58.99% / 53.09%
within tolerance
T1_Calu_RNA
Reported
4%
Reproduced
3.87% / 3.65%
exact
T1_Calu_cDNA
Reported
2%
Reproduced
2.05% / 1.97%
exact
T1_Calu_PCR
Reported
3%
Reproduced
2.63% / 2.40%
within tolerance
T1_Caco_RNA
Reported
4%
Reproduced
4.31% / 4.02%
within tolerance
T1_Caco_cDNA
Reported
2%
Reproduced
2.32% / 2.21%
within tolerance
T1_Caco_PCR
Reported
3%
Reproduced
3.46% / 3.23%
within tolerance
T1_A549_RNA
Reported
0.02%
Reproduced
0.0155% / 0.0144%
within tolerance
T1_A549_cDNA
Reported
0.01%
Reproduced
0.0079% / 0.0075%
within tolerance
T1_A549_PCR
Reported
0.01%
Reproduced
0.0063% / 0.0057%
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 90/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

250.1 k
tokens (I/O) · 18.7 M incl. cache
95 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.