Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identification of herpesvirus transcripts from genomic regions around the replication origins.

Sci Rep · 2023
L1 59/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
59/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 19% of all assessed papers rank 925 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce, with one important data-identity caveat. The analysis code is split: LoRTIA (github.com/zsolt-balazs/LoRTIA, v0.9.9) is the actual transcript caller; Rlyeh (github.com/Balays/Rlyeh) is only an R viz helper (no driver) -> not a turnkey pipeline. We reproduced the documented pipeline (minimap2 -ax splice -Y -C5 --cs -> LoRTIA) on the PRV arm. DATA-IDENTITY CORRECTION (audit this): the room was auto-seeded with GSE79337 (an unrelated 2016 human SuperSeries) and an early read mislabelled PRJEB64684 as KSHV; ENA confirms PRJEB64684 = Suid herpesvirus 1 (Pseudorabies virus) strain Kaplan, 36 ONT direct-cDNA runs, so reference = PRV Kaplan KJ717942.1 and the KSHV '199 transcripts' number is a DIFFERENT dataset (not attempted). RESULTS («our HPC» «job», n031, exit 0): (1) 1:1 tool fidelity -- LoRTIA reproduces its CI test exactly for TSS (13/13) and TES (16/16) feature positions (Jaccard 1.0); transcript/intron callers diverge (0.84/0.50) with the scipy/pysam version, NOT minimap2 (pinned 2.26). (2) PRV replication kinetics reproduce the paper's central PRV finding -- viral read fraction rises monotonically 0.17%->71.1% over 1-12h (~427x) and PAA suppresses it 5-18x at every timepoint (PAA blocks viral DNA replication -> less late transcription; Fig 7 / Table 2). (3) The pooled PRV transcript catalog -- the deliverable that did not finish in the prior run -- now COMPLETES: 25 transcripts / 18 TSS / 457 TES / 2032 introns from 2,204,371 pooled reads on KJ717942.1, all human-auditable as GFF3. NOT attempted (out of scope / hard 20%): exact qRT-PCR fold-changes (wet-lab), the named origin transcripts (need supplement coordinates), the authors' multi-sample curation (under-specified), the other 7 viruses, and the KSHV count (different data). No fabrication concern: no reported value appeared underivable-in-principle from the shipped data/code.

💻 Code ↗ 🗄 Data: GSE79337

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 59
    assessed: 2026-06-15 ⛓ 74431cee9853
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether additional, previously undetected transcripts (lncRNAs and mRNA isoforms) exist proximal to or overlapping the replication origins (Oris) of herpesviruses across all three subfamilies, and whether these Ori-proximal regions show conserved patterns of transcriptional overlap linked to coordinated regulation of replication and transcription.

Core claims
  • Herpesviruses display distinct patterns of transcriptional overlaps near or at the replication origins (Oris) finding
  • Novel lncRNAs and splice/length isoforms of mRNAs were discovered in Ori-proximal regions across nine herpesviruses from all three subfamilies finding
  • An intricate network of transcriptional overlaps exists within the examined Ori-proximal genomic regions finding
  • A 'super regulatory center' exists in alphaherpesvirus genomes that governs initiation of both DNA replication and global transcription through multilayered molecular interactions mechanism
  • TATA box promoter elements were identified within the Ori regions in all six examined alphaherpesviruses, with transcription start sites mapping close to them, suggesting these elements are functional finding
  • Very long 5' transcript isoforms of transcription regulator genes (us1 and icp4) overlap the OriS in BoHV-1, EHV-1, HSV-1, and SVV finding
  • LoRTIA software applied to direct cDNA sequencing datasets was used as the primary pipeline to identify putative transcripts method
  • A multi-technique validation pipeline (dcDNA-Seq across replicates, CAGE-Seq/RAMPAGE/dRNA-Seq for TSS/TES, qRT-PCR for transcript confirmation, dRNA-Seq for introns) was required for confident transcript annotation method
Experimental setups
Assay System Perturbation Readout Platform
direct cDNA sequencing (dcDNA-Seq) nine herpesviruses (BoHV-1, EBV, HCMV, HSV-1, PRV, SVV, VZV, EHV-1, KSHV) in infected cells none transcript identification, TSS/TES mapping, splice variant detection (via LoRTIA software) ONT / PacBio (RSII, Sequel)
direct RNA sequencing (dRNA-Seq) same panel of herpesviruses none transcript detection/validation and intron annotation ONT MinION
CAGE-Seq (Cap Analysis of Gene Expression sequencing) VZV, EBV, EHV-1, KSHV none validation of transcription start sites (TSSs) Illumina
qRT-PCR PRV, BoHV-1, EHV-1, HSV-1, KSHV none validation of lncRNAs and longer mRNA isoforms
multi-time-point real-time RT-PCR (RT2-PCR) PRV time-course (lytic infection kinetics) expression kinetics of the three most important lncRNAs
short-read RNA sequencing, amplified cDNA (SRS Illumina) HSV-1, PRV, VZV, SVV, EBV none transcript detection and relative abundance quantification Illumina
long-read amplified cDNA sequencing (LRS PacBio RSII/Sequel) EBV, HCMV, HSV-1, PRV none full-length transcript identification PacBio RSII / Sequel
synthetic long-read sequencing (LoopSeq) BoHV-1 none transcript identification Illumina (LoopSeq)
Key results
  • Novel lncRNAs identified overlapping/proximal to OriS: OriS-RNA (BoHV-1), OriS-RNA1 (HSV-1), NOIR-1 (PRV, EHV-1, VZV, SVV), and NOIR-2 (PRV)
  • TATA box elements found within OriS regions in all six examined alphaherpesviruses, with transcript TSSs mapping close to them
  • Very long 5' transcript isoforms of us1 and icp4 overlap the OriS in BoHV-1, EHV-1, HSV-1, and SVV
  • Five distinct lncRNAs with unique TSSs/TESs (NOIR-1A, -1B, -1C, -1D, -1E) identified in the VZV OriS-proximal region
  • NOIR-1 transcripts expressed at moderate abundance while NOIR-2 shows very low abundance
  • In SVV, a novel long TSS variant of NOIR-1 overlaps the canonical ICP4 transcript, and both this variant and canonical NOIR-1 overlap the OriS
  • Long (>5 kb) transcripts are underestimated or undetected by LRS due to library-preparation size bias, but RT2-PCR confirmed they are still expressed, at lower levels
Key statistics
  • other approximately 72% (cited background survey: ~72% of mammalian ORC1 binding sites are associated with active promoters, more than half controlled by ncRNAs (not this study's own data))
  • other abundance scale: 1=1–9 reads, 2=10–49, 3=50–199, 4=200–999, 5=>1000 reads (shading scale used in figures to indicate relative transcript abundance)
  • count nine (number of herpesvirus species examined for Ori-proximal transcripts)
  • count three (number of biological replicates required to confirm a dcDNA-Seq read before annotation)
  • count five (number of distinct NOIR-1 lncRNA variants (NOIR-1A through -1E) identified in VZV)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a descriptive transcriptomic discovery study using long-read (LRS) and short-read sequencing (SRS) data—both newly generated and previously published—from nine herpesviruses across all three subfamilies. Transcript annotation relied on multi-method corroboration (dcDNA-Seq, dRNA-Seq, CAGE-Seq, and qRT-PCR) with detection requiring support across three biological replicates; no formal inferential statistical tests are described. Relative transcript abundance was reported in ordinal bins defined by raw read-count ranges rather than normalized expression metrics. Expression kinetics of selected lncRNAs were monitored by multi-time-point RT²-PCR.

Replicationbiological Sample sizeThree biological replicates required as a minimum detection criterion for transcript annotation; replicate counts for individual viruses and previously published datasets vary and are not uniformly stated GroupsOri-proximal transcriptomes of nine herpesviruses (HSV-1, PRV, EHV-1, BoHV-1, VZV, SVV, HCMV, EBV, KSHV) across alpha-, beta-, and gammaherpesvirus subfamilies Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Multi-method, multi-replicate corroboration threshold (transcript must appear in dcDNA-Seq across all three biological replicates, with TSS/TES support from CAGE-Seq, RAMPAGE, or dRNA-Seq, and confirmation by qRT-PCR) Transcript identification and annotation across all nine herpesviruses Three biological replicates stated as minimum detection criterion not stated
Multi-time-point real-time RT-PCR (RT²-PCR) for expression kinetics monitoring and transcript validation Expression kinetics of the three most important lncRNAs in PRV; lncRNA and mRNA isoform validation in PRV, BoHV-1, EHV-1, HSV-1, and KSHV not stated
Approaches that could also have been used
  • Relative transcript abundance was summarized in discrete ordinal bins defined by raw read-count ranges (e.g., 1–9, 10–49, 50–199 reads)
    Could also: Normalized expression metrics such as TPM (transcripts per million) or CPM could also be used to report relative abundance — Normalized metrics account for differences in sequencing depth and transcript length across libraries, enabling more direct quantitative comparison between samples and across viruses
  • Transcript detection was determined by a logical presence/absence criterion requiring support across all three biological replicates
    Could also: Probabilistic models implemented in tools such as DESeq2, edgeR, or kallisto/sleuth could also be used to assess reproducibility and statistical confidence of transcript-level detection across replicates — Statistical frameworks quantify uncertainty in detection, accommodate variable sequencing depth, and yield calibrated confidence estimates rather than a binary threshold
  • Multi-time-point RT²-PCR expression kinetics were described without a stated formal statistical test comparing expression levels across time points
    Could also: A repeated-measures ANOVA or linear mixed-effects model on log-transformed Ct values could also be applied to compare expression across kinetic phases — A formal model would quantify whether expression differences across IE, E, and L phases are statistically distinguishable while accounting for within-sample correlation across time
  • Transcriptional overlap patterns across nine herpesviruses were compared descriptively, qualitatively characterized as 'distinct patterns'
    Could also: Permutation-based genome-wide overlap statistics (e.g., using BEDTools shuffle with a randomization framework) could also be used to assess whether observed overlap densities at Ori-proximal regions exceed chance expectation — Formal overlap statistics would provide a quantitative basis for claims that Ori-proximal regions are enriched for transcriptional complexity relative to the rest of the genome
  • Transcript annotations were validated by corroboration across multiple sequencing platforms and library types without a unified confidence score
    Could also: An explicit evidence-weighting or confidence-scoring scheme (e.g., assigning tier levels based on number and type of supporting methods) could also be used to stratify annotation reliability — A transparent scoring system would make the hierarchy of evidence reproducible and allow readers to immediately assess how robustly any given transcript is supported
  • Size-biasing effects of LRS library preparation on long (>5 kb) transcript detection were acknowledged qualitatively, with RT²-PCR cited as evidence that long transcripts are expressed at lower levels
    Could also: A correction model for length-dependent sequencing bias (e.g., as implemented in tools like NanoCORR or length-aware quantification approaches) could also be applied to the LRS data — Explicit bias correction would allow more accurate estimation of true abundance for long transcripts relative to short ones, rather than relying solely on orthogonal PCR-based evidence
Software: LoRTIA

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
13
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

10.6084/m9.figshare.22339879.v1 DOI in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
ERP019579 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
ERP106430 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE128324 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE59717 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE97785 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
JX898220 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PRJEB64684 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37773348

Paper: Torma et al. 2023, Sci Rep 13:16234. "Identification of herpesvirus transcripts from genomic regions around the replication origins." DOI 10.1038/s41598-023-43344-y · PMCID PMC10541914.

Nature of the study

A multi-virus re-analysis + new-data long-read transcriptomics study spanning 8 herpesviruses (HSV-1, PRV, VZV, EHV-1, BoHV-1, KSHV, EBV, HCMV) and several platforms (ONT MinION dRNA/dcDNA/cDNA, PacBio Iso-Seq, Illumina CAGE/RAMPAGE, LoopSeq). Goal: find novel transcripts overlapping the viral replication origins (OriS, OriL, OriLyt, OriP).

Pipelines named (per Methods)

  • Alignment: minimap2 -ax splice -Y -C5 -cs to each viral reference.
  • Transcript identification: LoRTIA toolkit (github.com/zsolt-balazs/LoRTIA, v0.9.9) — calls TSS/TES/introns and reconstructs isoforms from long reads.
  • Viz/quant helpers: Rlyeh R package (github.com/Balays/Rlyeh) — import alignments into R, coverage/end-distribution plots. No README, no driver script, no example pipeline → viz library, not a turnkey reproduction.
  • Promoter/TATA: MotifFinder (Seqtools). TSS clustering: CAGEfightR. (downstream)
  • qRT-PCR validation (Fig 6/7, Table 2): wet-lab → OUT OF SCOPE.

⚠️ DATA-IDENTITY CORRECTION (audit this)

The room was auto-seeded with GSE79337 (a 2016 human SuperSeries — unrelated to this paper) and an early read mis-labelled PRJEB64684 as KSHV. Verified against ENA: PRJEB64684 = Suid herpesvirus 1 (Pseudorabies virus) strain Kaplan — 36 ONT MinION direct-cDNA RNA-Seq runs, PK-15 cells, time course 1/2/4/6/8/12 h, untreated vs PAA (3 reps). It is NOT KSHV, and the KSHV "199 transcripts" number belongs to a different dataset. Reference therefore = PRV Kaplan KJ717942.1 (143,423 bp), and the in-scope target is the paper's PRV arm (Fig 7 / Table 2).

In scope (pipeline-derived, attempted)

  1. LoRTIA tool fidelity — re-run LoRTIA on its shipped synthetic test set and compare to the shipped expected outputs (deterministic). Anchor 1.
  2. PRV replication kinetics + PAA suppression — apply the documented pipeline (minimap2 -ax splice -Y -C5 --cs → LoRTIA) to all 36 PRJEB64684 ONT runs against PRV Kaplan KJ717942.1; compute per-timepoint viral read fraction (untreated vs PAA) to reproduce the central PRV finding (Fig 7 / Table 2). Anchor 2.
  3. Pooled PRV transcript catalog — pooled LoRTIA TSS/TES/intron/transcript call over all viral reads, to recover origin-associated transcripts (CTO, NOIR-1, AZURE). Anchor 3.

NOT this dataset (was mis-pinned)

  • KSHV "199 transcripts (192 CAGE, 159 RAMPAGE)" — a different KSHV dataset, not PRJEB64684. Not attempted here.

Out of scope / not attempted (the hard ~20%)

  • The other 7 viruses (each its own reference + heterogeneous platform params).
  • CAGE/RAMPAGE Illumina confirmation and CAGEfightR clustering (separate data, STAR pipeline) — needed for the "192 / 159" sub-counts.
  • The authors' multi-sample acceptance/curation logic ("accepted in ≥3 samples OR dcDNA+dRNA/CAGE validation"; dRNA "Gold Standard" intron rule). This manual integration on top of LoRTIA is under-specified in the repo (Rlyeh has no driver) → the exact integer 199 is unlikely to be hit; expect a partial grade.
  • All wet-lab results (qRT-PCR fold-changes, Northern blots).

Data / code pointers

  • Code: github.com/Balays/Rlyeh @ aeb8421 (viz) + github.com/zsolt-balazs/LoRTIA (caller, v0.9.9). Cloned on «infra» under reproductions/pmid-37773348/.
  • Data: ENA PRJEB64684 — 36 ONT MinION RNA-Seq runs (ERR11788801…/ERR121316xx).
  • Reference: NCBI GQ994935.1 (KSHV strain TREx, ~137 kb).
  • «infra» work dir: «path»
Figures / tables: Fig 7
lortia_selftest_tss_tes
Reported
LoRTIA shipped synthetic CI expected TSS/TES/intron/transcript GFFs (deterministic tool-fidelity anchor)
Reproduced
TSS Jaccard 1.0 (13/13, positions identical); TES Jaccard 1.0 (16/16, identical); transcripts 0.836; introns 0.50
within tolerance
prv_replication_kinetics_paa
Reported
PRV transcription rises over the 1-12h time course and is strongly suppressed by PAA (Fig 7 / Table 2)
Reproduced
untreated PRV read fraction 0.17%->0.38%->5.0%->17.0%->35.3%->71.1% over 1-12h (~427x); untreated/PAA suppression ratio 5-18x at every timepoint (10.9/5.0/18.1/18.4/10.0/6.9)
partial
prv_transcript_catalog
Reported
origin-associated PRV transcripts (CTO near OriL; NOIR-1, AZURE near OriS) via minimap2 -> LoRTIA
Reproduced
pooled LoRTIA catalog over 2,204,371 primary-mapped PRV reads: 25 distinct transcripts, 18 TSS, 457 TES, 2032 introns on KJ717942.1 (the deliverable the prior run never finished; now complete and human-auditable)
partial
kshv_199_transcripts
Reported
199 KSHV transcripts (192 CAGE, 159 RAMPAGE)
Reproduced
NOT ATTEMPTED — different KSHV dataset (not PRJEB64684); out of scope
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 59/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

569.9 k
tokens (I/O) · 84.3 M incl. cache
275 min
runtime · 6.56 CPU-h
28.6 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine