Identification of herpesvirus transcripts from genomic regions around the replication origins.
The main results reproduced, with only marginal, non-material deviations.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce, with one important data-identity caveat. The analysis code is split: LoRTIA (github.com/zsolt-balazs/LoRTIA, v0.9.9) is the actual transcript caller; Rlyeh (github.com/Balays/Rlyeh) is only an R viz helper (no driver) -> not a turnkey pipeline. We reproduced the documented pipeline (minimap2 -ax splice -Y -C5 --cs -> LoRTIA) on the PRV arm. DATA-IDENTITY CORRECTION (audit this): the room was auto-seeded with GSE79337 (an unrelated 2016 human SuperSeries) and an early read mislabelled PRJEB64684 as KSHV; ENA confirms PRJEB64684 = Suid herpesvirus 1 (Pseudorabies virus) strain Kaplan, 36 ONT direct-cDNA runs, so reference = PRV Kaplan KJ717942.1 and the KSHV '199 transcripts' number is a DIFFERENT dataset (not attempted). RESULTS («our HPC» «job», n031, exit 0): (1) 1:1 tool fidelity -- LoRTIA reproduces its CI test exactly for TSS (13/13) and TES (16/16) feature positions (Jaccard 1.0); transcript/intron callers diverge (0.84/0.50) with the scipy/pysam version, NOT minimap2 (pinned 2.26). (2) PRV replication kinetics reproduce the paper's central PRV finding -- viral read fraction rises monotonically 0.17%->71.1% over 1-12h (~427x) and PAA suppresses it 5-18x at every timepoint (PAA blocks viral DNA replication -> less late transcription; Fig 7 / Table 2). (3) The pooled PRV transcript catalog -- the deliverable that did not finish in the prior run -- now COMPLETES: 25 transcripts / 18 TSS / 457 TES / 2032 introns from 2,204,371 pooled reads on KJ717942.1, all human-auditable as GFF3. NOT attempted (out of scope / hard 20%): exact qRT-PCR fold-changes (wet-lab), the named origin transcripts (need supplement coordinates), the authors' multi-sample curation (under-specified), the other 7 viruses, and the KSHV count (different data). No fabrication concern: no reported value appeared underivable-in-principle from the shipped data/code.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 59assessed: 2026-06-15 ⛓ 74431cee9853
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether additional, previously undetected transcripts (lncRNAs and mRNA isoforms) exist proximal to or overlapping the replication origins (Oris) of herpesviruses across all three subfamilies, and whether these Ori-proximal regions show conserved patterns of transcriptional overlap linked to coordinated regulation of replication and transcription.
- ★ Herpesviruses display distinct patterns of transcriptional overlaps near or at the replication origins (Oris) finding
- ★ Novel lncRNAs and splice/length isoforms of mRNAs were discovered in Ori-proximal regions across nine herpesviruses from all three subfamilies finding
- ★ An intricate network of transcriptional overlaps exists within the examined Ori-proximal genomic regions finding
- ★ A 'super regulatory center' exists in alphaherpesvirus genomes that governs initiation of both DNA replication and global transcription through multilayered molecular interactions mechanism
- ★ TATA box promoter elements were identified within the Ori regions in all six examined alphaherpesviruses, with transcription start sites mapping close to them, suggesting these elements are functional finding
- ★ Very long 5' transcript isoforms of transcription regulator genes (us1 and icp4) overlap the OriS in BoHV-1, EHV-1, HSV-1, and SVV finding
- LoRTIA software applied to direct cDNA sequencing datasets was used as the primary pipeline to identify putative transcripts method
- A multi-technique validation pipeline (dcDNA-Seq across replicates, CAGE-Seq/RAMPAGE/dRNA-Seq for TSS/TES, qRT-PCR for transcript confirmation, dRNA-Seq for introns) was required for confident transcript annotation method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| direct cDNA sequencing (dcDNA-Seq) | nine herpesviruses (BoHV-1, EBV, HCMV, HSV-1, PRV, SVV, VZV, EHV-1, KSHV) in infected cells | none | transcript identification, TSS/TES mapping, splice variant detection (via LoRTIA software) | ONT / PacBio (RSII, Sequel) |
| direct RNA sequencing (dRNA-Seq) | same panel of herpesviruses | none | transcript detection/validation and intron annotation | ONT MinION |
| CAGE-Seq (Cap Analysis of Gene Expression sequencing) | VZV, EBV, EHV-1, KSHV | none | validation of transcription start sites (TSSs) | Illumina |
| qRT-PCR | PRV, BoHV-1, EHV-1, HSV-1, KSHV | none | validation of lncRNAs and longer mRNA isoforms | — |
| multi-time-point real-time RT-PCR (RT2-PCR) | PRV | time-course (lytic infection kinetics) | expression kinetics of the three most important lncRNAs | — |
| short-read RNA sequencing, amplified cDNA (SRS Illumina) | HSV-1, PRV, VZV, SVV, EBV | none | transcript detection and relative abundance quantification | Illumina |
| long-read amplified cDNA sequencing (LRS PacBio RSII/Sequel) | EBV, HCMV, HSV-1, PRV | none | full-length transcript identification | PacBio RSII / Sequel |
| synthetic long-read sequencing (LoopSeq) | BoHV-1 | none | transcript identification | Illumina (LoopSeq) |
- – Novel lncRNAs identified overlapping/proximal to OriS: OriS-RNA (BoHV-1), OriS-RNA1 (HSV-1), NOIR-1 (PRV, EHV-1, VZV, SVV), and NOIR-2 (PRV)
- – TATA box elements found within OriS regions in all six examined alphaherpesviruses, with transcript TSSs mapping close to them
- – Very long 5' transcript isoforms of us1 and icp4 overlap the OriS in BoHV-1, EHV-1, HSV-1, and SVV
- – Five distinct lncRNAs with unique TSSs/TESs (NOIR-1A, -1B, -1C, -1D, -1E) identified in the VZV OriS-proximal region
- – NOIR-1 transcripts expressed at moderate abundance while NOIR-2 shows very low abundance
- – In SVV, a novel long TSS variant of NOIR-1 overlaps the canonical ICP4 transcript, and both this variant and canonical NOIR-1 overlap the OriS
- ▼ Long (>5 kb) transcripts are underestimated or undetected by LRS due to library-preparation size bias, but RT2-PCR confirmed they are still expressed, at lower levels
- other approximately 72% (cited background survey: ~72% of mammalian ORC1 binding sites are associated with active promoters, more than half controlled by ncRNAs (not this study's own data))
- other abundance scale: 1=1–9 reads, 2=10–49, 3=50–199, 4=200–999, 5=>1000 reads (shading scale used in figures to indicate relative transcript abundance)
- count nine (number of herpesvirus species examined for Ori-proximal transcripts)
- count three (number of biological replicates required to confirm a dcDNA-Seq read before annotation)
- count five (number of distinct NOIR-1 lncRNA variants (NOIR-1A through -1E) identified in VZV)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a descriptive transcriptomic discovery study using long-read (LRS) and short-read sequencing (SRS) data—both newly generated and previously published—from nine herpesviruses across all three subfamilies. Transcript annotation relied on multi-method corroboration (dcDNA-Seq, dRNA-Seq, CAGE-Seq, and qRT-PCR) with detection requiring support across three biological replicates; no formal inferential statistical tests are described. Relative transcript abundance was reported in ordinal bins defined by raw read-count ranges rather than normalized expression metrics. Expression kinetics of selected lncRNAs were monitored by multi-time-point RT²-PCR.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Multi-method, multi-replicate corroboration threshold (transcript must appear in dcDNA-Seq across all three biological replicates, with TSS/TES support from CAGE-Seq, RAMPAGE, or dRNA-Seq, and confirmation by qRT-PCR) | Transcript identification and annotation across all nine herpesviruses | Three biological replicates stated as minimum detection criterion | not stated |
| Multi-time-point real-time RT-PCR (RT²-PCR) for expression kinetics monitoring and transcript validation | Expression kinetics of the three most important lncRNAs in PRV; lncRNA and mRNA isoform validation in PRV, BoHV-1, EHV-1, HSV-1, and KSHV | — | not stated |
-
Relative transcript abundance was summarized in discrete ordinal bins defined by raw read-count ranges (e.g., 1–9, 10–49, 50–199 reads)↳ Could also: Normalized expression metrics such as TPM (transcripts per million) or CPM could also be used to report relative abundance — Normalized metrics account for differences in sequencing depth and transcript length across libraries, enabling more direct quantitative comparison between samples and across viruses
-
Transcript detection was determined by a logical presence/absence criterion requiring support across all three biological replicates↳ Could also: Probabilistic models implemented in tools such as DESeq2, edgeR, or kallisto/sleuth could also be used to assess reproducibility and statistical confidence of transcript-level detection across replicates — Statistical frameworks quantify uncertainty in detection, accommodate variable sequencing depth, and yield calibrated confidence estimates rather than a binary threshold
-
Multi-time-point RT²-PCR expression kinetics were described without a stated formal statistical test comparing expression levels across time points↳ Could also: A repeated-measures ANOVA or linear mixed-effects model on log-transformed Ct values could also be applied to compare expression across kinetic phases — A formal model would quantify whether expression differences across IE, E, and L phases are statistically distinguishable while accounting for within-sample correlation across time
-
Transcriptional overlap patterns across nine herpesviruses were compared descriptively, qualitatively characterized as 'distinct patterns'↳ Could also: Permutation-based genome-wide overlap statistics (e.g., using BEDTools shuffle with a randomization framework) could also be used to assess whether observed overlap densities at Ori-proximal regions exceed chance expectation — Formal overlap statistics would provide a quantitative basis for claims that Ori-proximal regions are enriched for transcriptional complexity relative to the rest of the genome
-
Transcript annotations were validated by corroboration across multiple sequencing platforms and library types without a unified confidence score↳ Could also: An explicit evidence-weighting or confidence-scoring scheme (e.g., assigning tier levels based on number and type of supporting methods) could also be used to stratify annotation reliability — A transparent scoring system would make the hierarchy of evidence reproducible and allow readers to immediately assess how robustly any given transcript is supported
-
Size-biasing effects of LRS library preparation on long (>5 kb) transcript detection were acknowledged qualitatively, with RT²-PCR cited as evidence that long transcripts are expressed at lower levels↳ Could also: A correction model for length-dependent sequencing bias (e.g., as implemented in tools like NanoCORR or length-aware quantification approaches) could also be applied to the LRS data — Explicit bias correction would allow more accurate estimation of true abundance for long transcripts relative to short ones, rather than relying solely on orthogonal PCR-based evidence
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37773348
Paper: Torma et al. 2023, Sci Rep 13:16234. "Identification of herpesvirus transcripts from genomic regions around the replication origins." DOI 10.1038/s41598-023-43344-y · PMCID PMC10541914.
Nature of the study
A multi-virus re-analysis + new-data long-read transcriptomics study spanning 8 herpesviruses (HSV-1, PRV, VZV, EHV-1, BoHV-1, KSHV, EBV, HCMV) and several platforms (ONT MinION dRNA/dcDNA/cDNA, PacBio Iso-Seq, Illumina CAGE/RAMPAGE, LoopSeq). Goal: find novel transcripts overlapping the viral replication origins (OriS, OriL, OriLyt, OriP).
Pipelines named (per Methods)
- Alignment:
minimap2 -ax splice -Y -C5 -csto each viral reference. - Transcript identification: LoRTIA toolkit (github.com/zsolt-balazs/LoRTIA, v0.9.9) — calls TSS/TES/introns and reconstructs isoforms from long reads.
- Viz/quant helpers: Rlyeh R package (github.com/Balays/Rlyeh) — import alignments into R, coverage/end-distribution plots. No README, no driver script, no example pipeline → viz library, not a turnkey reproduction.
- Promoter/TATA: MotifFinder (Seqtools). TSS clustering: CAGEfightR. (downstream)
- qRT-PCR validation (Fig 6/7, Table 2): wet-lab → OUT OF SCOPE.
⚠️ DATA-IDENTITY CORRECTION (audit this)
The room was auto-seeded with GSE79337 (a 2016 human SuperSeries — unrelated to
this paper) and an early read mis-labelled PRJEB64684 as KSHV. Verified against
ENA: PRJEB64684 = Suid herpesvirus 1 (Pseudorabies virus) strain Kaplan — 36
ONT MinION direct-cDNA RNA-Seq runs, PK-15 cells, time course 1/2/4/6/8/12 h,
untreated vs PAA (3 reps). It is NOT KSHV, and the KSHV "199 transcripts" number
belongs to a different dataset. Reference therefore = PRV Kaplan KJ717942.1
(143,423 bp), and the in-scope target is the paper's PRV arm (Fig 7 / Table 2).
In scope (pipeline-derived, attempted)
- LoRTIA tool fidelity — re-run LoRTIA on its shipped synthetic test set and compare to the shipped expected outputs (deterministic). Anchor 1.
- PRV replication kinetics + PAA suppression — apply the documented pipeline (minimap2 -ax splice -Y -C5 --cs → LoRTIA) to all 36 PRJEB64684 ONT runs against PRV Kaplan KJ717942.1; compute per-timepoint viral read fraction (untreated vs PAA) to reproduce the central PRV finding (Fig 7 / Table 2). Anchor 2.
- Pooled PRV transcript catalog — pooled LoRTIA TSS/TES/intron/transcript call over all viral reads, to recover origin-associated transcripts (CTO, NOIR-1, AZURE). Anchor 3.
NOT this dataset (was mis-pinned)
- KSHV "199 transcripts (192 CAGE, 159 RAMPAGE)" — a different KSHV dataset, not PRJEB64684. Not attempted here.
Out of scope / not attempted (the hard ~20%)
- The other 7 viruses (each its own reference + heterogeneous platform params).
- CAGE/RAMPAGE Illumina confirmation and CAGEfightR clustering (separate data, STAR pipeline) — needed for the "192 / 159" sub-counts.
- The authors' multi-sample acceptance/curation logic ("accepted in ≥3 samples OR dcDNA+dRNA/CAGE validation"; dRNA "Gold Standard" intron rule). This manual integration on top of LoRTIA is under-specified in the repo (Rlyeh has no driver) → the exact integer 199 is unlikely to be hit; expect a partial grade.
- All wet-lab results (qRT-PCR fold-changes, Northern blots).
Data / code pointers
- Code: github.com/Balays/Rlyeh @ aeb8421 (viz) + github.com/zsolt-balazs/LoRTIA (caller, v0.9.9). Cloned on «infra» under reproductions/pmid-37773348/.
- Data: ENA PRJEB64684 — 36 ONT MinION RNA-Seq runs (ERR11788801…/ERR121316xx).
- Reference: NCBI GQ994935.1 (KSHV strain TREx, ~137 kb).
- «infra» work dir: «path»
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.