Time course profiling of host cell response to herpesvirus infection using nanopore and synthetic long-read transcriptome sequencing.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
IN PROGRESS. Reproducing LoRTIA transcript-isoform annotation (TSS/TES/intron/transcript counts) on ENA PRJEB33511 (55 ONT runs, 66.9 GB) aligned with minimap2 splice to bovine ARS-UCD1.2 + BoHV-1 JX898220.1. Setup (ref+env+index+data download) running on «our HPC» front1; LoRTIA per-run + combine to follow on SLURM. DE/k-means results (C6,C7) out of scope: R code not shipped in LoRTIA repo.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-18 ⛓ e1e1e0a9b2af
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-06-30
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether long-read sequencing (nanopore direct cDNA sequencing and Illumina-based synthetic long-read LoopSeq) can reveal time-resolved changes in host (bovine) gene expression and transcript isoform usage during BoHV-1 infection, including whether the virus induces transcriptional readthrough as reported for other alphaherpesviruses.
- ★ BoHV-1 infection causes substantial up- and down-regulation of host gene networks, including antiviral response and viral transcription/translation-associated genes finding
- ★ Nanopore and synthetic long-read (LoopSeq) sequencing identified a large number of novel bovine transcript isoforms finding
- ★ Viral infection causes differential expression and altered usage (TSS/TES/splicing) of host transcript isoforms finding
- ★ No increased rate of transcriptional readthrough was detected in BoHV-1-infected bovine cells, unlike what was reported for another alphaherpesvirus (HSV-1) finding
- ★ This is the first report of LoopSeq applied to eukaryotic transcriptome analysis resource
- ★ This is the first report of nanopore sequencing used for kinetic characterization of cellular transcriptomes resource
- LoRTIA software suite was used for transcript, TSS, TES, and intron annotation from long-read data method
- ★ FOS transcript degradation is regulated by retention of its third intron, and novel non-spliced/spliced FOS variants appear from the first hour of infection mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| direct cDNA sequencing (dcDNA-Seq, oligo(dT)-primed RT), long-read | bovine epithelial cells | BoHV-1 infection (time course: mock, 1,2,4,6,8,12 h p.i.) | host gene expression levels, transcript isoforms, TSS/TES/splicing usage | Oxford Nanopore MinION |
| amplified cDNA sequencing (random-oligonucleotide-primed RT), long-read | bovine cells | BoHV-1 infection | reads mapping to intergenic regions (transcriptional readthrough detection) | Oxford Nanopore MinION |
| synthetic long-read sequencing (LoopSeq) | bovine cells | none (transcript annotation) | bovine transcript identification and annotation | Illumina platform (Loop Genomics LoopSeq) |
| GO over-representation analysis (bioinformatic) | bovine host genes (from dcDNA-Seq dataset) | BoHV-1 infection | functional category enrichment of differentially expressed genes | PANTHER |
| protein-protein/gene network association analysis (bioinformatic) | bovine host genes (immediate response set) | BoHV-1 infection (1 h p.i. vs mock) | gene network associations among DE genes | STRING |
- – 686 of 8342 host genes showed significantly altered expression during the course of infection (FDR 0.01), clustering into 6 kinetic profiles
- – At 1 h p.i., 6 genes were significantly downregulated and 19 genes significantly upregulated relative to mock
- – Polyadenylated transcript length changed significantly starting at 2 h p.i., increasing most at 2 h and decreasing most at 12 h relative to mock +134 bps at 2h; -193 bps at 12h (mock mean 1365 bps)
- – No substantial reads mapped to intergenic regions despite comparable overall mapping rate, indicating no detectable transcriptional readthrough
- – 546 transcript isoforms detected with ≥10 reads in infected vs mock samples, including 130 alternatively spliced transcripts, 72 with downstream and 122 with upstream TESs, and 80 upstream/142 downstream TSS isoforms
- ▲ Novel non-spliced FOS variant and splice variants lacking the third intron detected from the first hour of infection
- ▼ A shorter 3'-UTR isoform of SOD1 detected in infected cells compared to mock
- count 11,025 TSSs, 21,317 TESs, 139,771 introns; 227,672 bovine transcripts (LoRTIA-based bovine transcript annotation)
- mean median transcript length 1678 nt (σ = 2386.5) (annotated bovine transcripts)
- count 686 significantly altered genes out of 8342 host genes (FDR 0.01) (differential expression across infection time course)
- fold_change mean transcript length change +134 bps at 2h p.i. (p<0.05) and -193 bps at 12h p.i. (p<0.05) vs mock mean 1365 bps (polyadenylated transcript length dynamics)
- mean TATA box to TSS distance 31.15 nt (σ=2.96); PAS to TES distance 25.35 nt (σ=8.26) (core promoter/terminator element positioning)
- count kinetic cluster sizes: n=53, 64, 82, 64, 88, 335 (six gene expression kinetic clusters)
- count GO functional categories: 296 metabolism, 257 transcription/RNA decay, 242 development, 187 immune response, 161 translation/protein folding, 61 viral transcription (over-representation analysis of 8342 reference genes)
- count 6 downregulated and 19 upregulated genes at 1 h p.i. (FDR 0.01) (immediate host response gene set)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This time-course study profiled bovine host-cell transcriptome responses to BoHV-1 infection across seven time points (mock plus 1–12 h p.i.) using ONT nanopore direct cDNA sequencing with three biological replicates per time point, supplemented by LoopSeq synthetic long-read sequencing. Differential expression (DE) analysis with an FDR ≤ 0.01 threshold identified significantly altered host genes, which were then clustered into six kinetic groups by relative expression profile. Over-representation analysis via PANTHER (FDR < 0.05) assigned functional GO categories, and transcript-length changes at each time point versus mock were assessed with p < 0.05 thresholds.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression analysis — specific package not named; output columns (logFC, logCPM, F-statistic, p-value, FDR) are consistent with a quasi-likelihood F-test framework (e.g., edgeR glmQLFTest) | Host gene expression changes across the 12 h infection time course; Mock vs. 1 h p.i. comparison reported in Table 1 | 8342 host genes with >10 transcripts in each of three biological replicates | not stated |
| Over-representation analysis (Fisher's exact or hypergeometric, via PANTHER software) | GO biological process and GO molecular function enrichment of DE genes and kinetic clusters (Fig. 4c, Supplementary Table S2) | 686 DE genes against a background of 8342 expressed genes | not stated |
| Statistical comparison of polyadenylated transcript length distributions at each time point vs. mock — specific test not named; p < 0.05 threshold stated | Transcript length changes at 1 h, 2 h, 4 h, 6 h, 8 h, 12 h p.i. versus mock (Fig. 3) | null | not stated |
| Expression profile clustering — algorithm not explicitly named | 686 DE host genes grouped into 6 kinetic clusters by relative expression profile (Fig. 4a, b) | 686 genes | na |
-
The DE analysis package is not named, though the tabular output (logFC, logCPM, F-statistic, FDR) is consistent with an edgeR quasi-likelihood framework applied to transcript count data↳ Could also: DESeq2 (negative-binomial Wald test) or limma-voom could also be applied to RNA count data; long-read-specific tools such as NanoCount or LIQA additionally model multi-mapping uncertainty inherent in full-length isoform reads — Explicitly naming the DE package lets readers assess model assumptions (negative-binomial dispersion estimation vs. mean-variance modelling) and reproduce results; long-read-aware tools address the higher multi-mapping rates typical of full-length isoform sequencing
-
Transcript-length distributions at each of six infected time points were compared to mock with p < 0.05, without a stated correction for the six simultaneous comparisons↳ Could also: A Benjamini-Hochberg FDR or Bonferroni correction across the six time-point comparisons could also be applied — Correcting across the family of six comparisons controls the probability of at least one spurious significant result, which is a standard step when the same null hypothesis (no length change vs. mock) is tested repeatedly
-
Gene expression profiles were clustered into six kinetic groups; the clustering algorithm and distance or similarity metric are not specified↳ Could also: Soft clustering (e.g., Mfuzz, which uses fuzzy c-means) or model-based clustering (mclust) designed for time-series expression data could also be used — Time-series-aware methods account for temporal ordering and measurement uncertainty; reporting the algorithm and any tuning parameters (e.g., number of clusters chosen) allows readers to assess cluster robustness and reproduce assignments
-
Transcript-length changes are summarized as mean (x̄) change in base pairs without an accompanying measure of spread↳ Could also: Reporting SD, SEM, or a 95% CI alongside the mean would also convey the variability of the length distribution at each time point — A dispersion measure contextualizes the biological magnitude of the change relative to within-group variability and is especially informative given the n = 3 biological replicates
-
Over-representation analysis used a binary DE/non-DE split (686 genes vs. 8342 background) with PANTHER↳ Could also: Gene-set enrichment analysis (GSEA) on the full ranked list of 8342 genes, or tools such as clusterProfiler or g:Profiler with regularly updated annotations, could also be used — GSEA avoids an arbitrary significance threshold by using the full ranked gene list, which can detect weaker but coordinated pathway signals; updated annotation databases (KEGG, Reactome) may capture pathway relationships not yet reflected in GO
-
The study treats each time-point replicate set as independent, without modelling any shared structure across time points within the DE analysis↳ Could also: A spline-based or polynomial time-series model (e.g., edgeR/limma with natural splines, or ImpulseDE2) could also test for temporal expression trajectories across all time points jointly — Time-aware models borrow information across adjacent time points, can formally test for trajectory shapes (transient vs. monotone), and reduce the number of pairwise tests needed to characterize kinetic gene clusters
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34244540 (LoRTIA on BoHV-1 / MDBK nanopore time course)
Paper: Moldovan et al. 2021, Sci Rep 11:14613. "Time course profiling of host cell response to herpesvirus infection using nanopore and synthetic long-read transcriptome sequencing." DOI 10.1038/s41598-021-93142-7.
IMPORTANT correction vs BRIEF: host is Bos taurus (MDBK cells), virus is Bovine Herpesvirus 1.1 (BoHV-1, Cooper isolate) — NOT human. The BRIEF title ("host cell response to herpesvirus") is a generic placeholder.
Code: https://github.com/zsolt-balazs/LoRTIA (v0.9.9, HEAD 80b98c1, 2019-08-20). Third-party-style toolkit by the same lab; valid per P16. Python3 pipeline: Samprocessor (adapter detect) -> Stats (poisson) -> Gff_creator (TSS/TES/intron GFF) -> Transcript_Annotator (isoform reconstruction); Sum_gffs combines samples. Data: ENA PRJEB33511 — 55 Oxford Nanopore MinION RNA-Seq runs (dcDNA + amplified cDNA), 52.4M reads, 66.9 GB fastq (already basecalled, Guppy 3.4.1 upstream);
- 6 Illumina MiSeq LoopSeq synthetic-long-read runs (27k reads, tiny). Reference: bovine GCF_002263795.1_ARS-UCD1.2 (genome fna + gff) + virus JX898220.1.
IN SCOPE (pipeline-derived, attempt to reproduce)
Pipeline: minimap2 -ax splice -Y -C5 to combined bovine+BoHV-1 reference -> LoRTIA
nanopore-cDNA mode (adapters -5 TGCCATTAGGCCGGG --five_score 16 --check_in_soft 15 -3 AAAAAAAAAAAAAAA --three_score 16 -s poisson -f True) per run -> combine.
| id | reported value | paper location |
|---|---|---|
| C1 TSS | 11,025 TSSs | Results, transcript annotation |
| C2 TES | 21,317 TESs | Results |
| C3 intron | 139,771 introns | Results |
| C4 transcripts | 227,672 bovine transcripts (median 1678 nt) | Results |
| C5 genes>10tx | 8342 host genes w/ >10 transcripts | DE analysis |
HARDER / likely OUT OF SCOPE (downstream R, not in LoRTIA repo)
- C6: 686 genes significantly altered (EdgeR, TMM, FDR<0.01) — R code NOT shipped.
- C7: 6 k-means clusters of expression profiles — not shipped.
- C8: 130 alternatively spliced transcripts; 546 isoforms >=10 reads.
- C9: transcript-length shifts (mock 1365 bp; 2h +134; 12h -193). These depend on a custom EdgeR/k-means analysis not in the repo -> if unreproducible record as no_code for that sub-result; primary target is C1-C4 (LoRTIA core output).
Datasets to profile
- PRJEB33511: nanopore long-read RNA-Seq (cDNA), Bos taurus + BoHV-1. n_reported:
21 conditions x3 reps for dcDNA (+ amplified + LoopSeq); n_observed: 55 ONT runs
- 6 Illumina. Profile N concordance, completeness, QC.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This reproduction is unfinished: ROOM_RESULT.json shows all four in-scope LoRTIA counts (11025 TSS / 21317 TES / 139771 introns / 227672 transcripts) as pending and agreement.json is not-run-yet, so nothing was actually compared. The shortfall is on our side (the SLURM annotation job had not completed) plus a code-availability gap on the authors' side — the EdgeR/k-means R code (C6=686 genes, C7=6 clusters) is not in the LoRTIA repo. Input data (ENA PRJEB33511) is publicly available, though sample selection is somewhat self-defined (61 runs observed vs 55 used). No discrepancy or fabrication is evident; the case is simply incomplete, hence uniformly yellow rather than red.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.