Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Cell type differences in human cytomegalovirus transcription and epigenetic regulation with insights into major immediate-early enhancer-promoter control.

PLoS Pathog · 2025
90/100 PQI 97
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 79% of all assessed papers rank 211 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1. The 'code' repo (meierjl/EP3) is a UCSC Track Hub, not analysis code, so this is a P16 independent-tool reproduction: reran the GEO-specified standard DFF-ChIP pipeline (trim_galore 0.6.10 -> bowtie 1.2.2 -> bedtools genomecov -> bedGraphToBigWig) on the paper's own raw reads and correlated the regenerated viral pileup against the authors' deposited bigWig. This was a clean independent re-run after the «infra» workdir was reclaimed: envs, raw FASTQ, reference and pipeline were rebuilt from scratch and recomputed (SLURM «job»). BOTH D-NT2 Exp3 samples reproduce near-perfectly genome-wide: Pol II Rep1 Pearson 0.99645 / Spearman 0.997; H3K4me3 Rep1 Pearson 0.99970 / Spearman 0.99981. A consistency check shows the repo bigWigs are byte-identical to the GEO deposit (exact). A ~1.9x absolute-magnitude offset (per-base counting / normalisation convention, under-specified in the deposit) is flagged honestly; scale-invariant correlation confirms faithful shape reproduction. The deposited 'UMI-dedup' track matches the NON-dedup regenerated pileup best, indicating UMI dedup removed few fragments. AUDITABILITY: all six compared bigWigs are retained under reproduction/stage/ (1.5 MB each) with SHA256SUMS, and reproduction/verify_local.py re-derives every correlation locally on «host» (no «our HPC»), matching to all digits; the regenerated bigWig is byte-reproducible across two independent pipeline runs (deterministic). NOT attempted (out of scope, by design): exact UMI-dedup awk/bash (not shipped) and hg38+virus competitive mapping (virus-only used; cross-mapping negligible at r1.0); and the broader PRO-Seq/RNA-Seq subseries, all differential/metaplot figure panels, and all wet-lab assays (protein degradation, MIE promoter mutagenesis/function, morphology).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 90
    assessed: 2026-06-16 ⛓ 299e1b4bdb2b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Prompted by an initial observation of striking HCMV transcription differences between fibroblasts (HFF) and differentiated neural lineage NTera2 cells (D-NT2) in late-stage infection, the study asks how HCMV transcriptional regulation and chromatin architecture differ between these two cell types, and what viral (IE2, LTFs, cis-regulatory elements) and host (P-TEFb) factors govern this cell type-dependent control.

Core claims
  • Six viral promoters (UL5, UL72, EP3, UL57-AS, US16-AS, US30-S) are ≥50-fold more active in D-NT2 than in HFF at 96 h post-infection and are classified as viral long promoters. finding
  • In D-NT2, Pol II initiation and elongation at viral long promoters requires viral DNA synthesis and is independent of host P-TEFb, viral IE2, and LTF. finding
  • GC-box mutations in the MIE enhancer increase EP3 enhancer transcription, whereas mutations in CREB and NF-kB response elements reduce it. finding
  • GC-box mutations alter infected D-NT2 cell morphology and gene expression program without affecting viral MIE gene expression levels; CREB/NF-kB element mutations do not induce these changes. finding
  • LTF-driven promoters constitute a smaller proportion of the viral late promoter population in D-NT2 and are generally less active than in HFF. finding
  • HCMV genomes carry more nucleosomes in D-NT2 than in HFF, potentially restricting LTF access to viral promoters. finding
  • A TBP-IE2-nucleosome complex, containing more nucleosomes than the analogous complex in fibroblasts, occupies the MIEP transcription start site in D-NT2. finding
  • Host RNA Pol II, detected by DFF-ChIP Seq, occupies viral long promoters at positions coinciding with PRO-Seq signal, indicating host Pol II drives this unusual transcription mode. finding
Experimental setups
Assay System Perturbation Readout Platform
PRO-Seq D-NT2 and HFF cells, HCMV (Towne) infection flavopiridol (P-TEFb inhibitor) added final 1 h nascent RNA 5'-end density marking viral transcription start sites (TSS strength)
PRO-Seq D-NT2 cells, HCMV (Towne) infection phosphonoformic acid (PFA, viral DNA replication inhibitor); timepoints 12 h and 72 h pi nascent RNA at viral long promoters before vs after onset of viral DNA replication
DFF-ChIP Seq (Pol II) D-NT2 cells, HCMV (Towne) infection, 96 h pi none (native chromatin digestion) Pol II occupancy (40-65 bp DNA fragments) across viral genome/long promoters
DFF-ChIP Seq HCMV-infected cells (D-NT2/HFF), 48 h pi context none TBP-IE2-nucleosome complex occupancy at MIEP TATA box, crs, and TSS
Promoter mutation and function assays HCMV MIE enhancer (EP3) in infected D-NT2 GC-box, CREB, and NF-kB response element sequence mutations enhancer transcription activity, cell morphology, and gene expression program
Rapid viral protein degradation HCMV-infected D-NT2 and HFF targeted degradation of viral IE2 or LTF proteins dependence of viral promoter transcription on IE2/LTF
RNA-Seq HCMV-infected D-NT2 and HFF none/other (as stated in abstract) viral and host transcript levels
Key results
  • 124 viral TSSs differentially expressed between HFF and D-NT2 (≥4-fold difference, p≤0.005): 73 more active in D-NT2, 51 more active in HFF ≥4-fold, p≤0.005
  • Six promoters (UL5, UL72, EP3, UL57-AS, US16-AS, US30-S) selected as ≥50-fold more active in D-NT2 than HFF ≥50-fold
  • Viral long promoter (UL5, UL72, EP3) transcription in D-NT2 does not begin until after onset of viral DNA replication and is prevented by PFA
  • Pol II DFF-ChIP occupancy positions correspond to PRO-Seq signal locations at UL5, UL72, and EP3 long promoters
  • GC-box mutations increase EP3 enhancer transcription; CREB and NF-kB response element mutations reduce it
  • GC-box mutations alter infected D-NT2 cell morphology and gene expression program without changing viral MIE gene expression levels
  • D-NT2 infections yielded 32% fewer spike-in normalized HCMV genome-mapped reads and 37% lower HCMV-to-human read ratio vs HFF 32% fewer; 37% lower ratio
  • RNA4.9 promoter TSS read density is 30% lower in D-NT2 compared to HFF 30% lower
Key statistics
  • count 124 differentially expressed viral TSSs (TSSs with ≥4-fold difference and p≤0.005 between HFF and D-NT2 at 96 h pi)
  • fold_change ≥4-fold (threshold defining differentially expressed viral TSSs)
  • pvalue p≤0.005 (significance threshold (probability = 1 minus p-value) for differential TSS expression via NOISeq)
  • fold_change ≥50-fold (threshold for selecting six D-NT2-active promoters (UL5, UL72, EP3, UL57-AS, US16-AS, US30-S))
  • count 73 TSSs more active in D-NT2; 51 TSSs more active in HFF (split of the 124 differentially expressed viral TSSs)
  • fold_change 32% fewer spike-in normalized HCMV reads in D-NT2 vs HFF (total viral genome-mapped reads comparison)
  • fold_change 37% lower HCMV-to-human genome mapped read ratio in D-NT2 (normalization comparison between cell types)
  • other 2.5- to 3-fold more infectious units applied to D-NT2 (MOI 7.5) vs HFF (MOI 3) (viral inoculum adjustment to equalize infection/DNA levels at 96 h pi)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This functional genomics study used PRO-Seq, DFF-ChIP Seq, and RNA-Seq to compare HCMV transcription between human fibroblasts (HFF) and differentiated NTera2 neural lineage cells (D-NT2) at 96 h post-infection, with additional time-course and pharmacological perturbation conditions. Spike-in normalized nascent RNA read counts were compared between cell types using the NOISeq program, identifying differentially active viral transcription start sites (TSSs) at ≥4-fold difference with a probability threshold of ≥0.995 (reported as p-value ≤0.005). Results were primarily displayed as auto-scaled UCSC Genome Browser tracks and read count summaries across multiple independent experiments referenced in a supplementary table.

Replicationbiological Sample sizeMultiple independent experiments referenced (e.g., Exp 1, 2, 5 in S1 Table; 2 independent DFF-ChIP Seq experiments conducted months apart); exact biological replicate n per condition not stated in the excerpt GroupsHFF (fibroblasts) vs D-NT2 (differentiated neural lineage cells), both HCMV Towne-infected for 96 h; additional conditions include 12 h and 72 h time points, PFA (DNA replication inhibitor), and Flavopiridol (P-TEFb inhibitor) treatments Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated; NOISeq probability threshold (≥0.995) combined with fold-change filter (≥4-fold) applied across tested viral TSSs
Statistical tests used
Test Applied to n Assumptions
NOISeq differential expression analysis (non-parametric, count-based probability framework) Comparison of spike-in normalized viral TSS strength between HFF and D-NT2 PRO-Seq datasets (Fig 1D) Viral TSSs with ≥200 5'-end reads in either condition; 124 TSSs met the ≥4-fold and probability ≥0.995 thresholds; biological replicate n per condition not stated in excerpt not stated
Fold-change threshold filter (≥4-fold for differential display; ≥50-fold for promoter selection) Applied to PRO-Seq TSS comparisons to identify and select viral TSSs for in-depth study (Fig 1D) null na
Visual/qualitative comparison of auto-scaled genome browser tracks PRO-Seq and DFF-ChIP Seq signal at individual viral promoters across all main figures null na
Approaches that could also have been used
  • Viral TSS differential activity was assessed with NOISeq, a method suited for count data with few or no replicates
    Could also: DESeq2 or edgeR with explicit biological replicates could also be applied to PRO-Seq count data for differential TSS analysis — DESeq2 and edgeR use negative binomial dispersion models estimated from replicate data and produce FDR-adjusted p-values, formally quantifying uncertainty across the full family of tested features
  • Spike-in normalized PRO-Seq reads were additionally scaled for equivalence in total viral reads (excluding RNA4.9) before differential analysis
    Could also: Spike-in normalization alone without a secondary total-read adjustment is also a standard approach; TMM or RLE normalization on spike-in-scaled counts would be another option — Making the normalization strategy and its rationale explicit — and reporting sensitivity of differential calls to the chosen method — allows readers to assess how much the results depend on the particular scaling applied
  • Differences in Pol II occupancy at viral long promoters between cell types were presented as auto-scaled genome browser tracks for visual comparison
    Could also: Quantitative summaries (e.g., spike-in-normalized read counts or RPKM per region) with a measure of spread across replicates could also accompany genome browser views — Numerical summaries with dispersion enable readers to evaluate effect magnitude and inter-replicate consistency independently of the auto-scaling applied to browser tracks
  • A fold-change threshold (≥4-fold, ≥50-fold) was used alongside the NOISeq probability threshold to select viral TSSs for further study
    Could also: A volcano plot or MA-plot displaying FDR-adjusted values across all tested TSSs could also accompany the differential analysis — Displaying the full distribution of effect sizes and adjusted significance values gives readers context for where the chosen thresholds fall relative to the complete dataset
  • No formal sample size justification or power analysis was stated for the number of independent experiments conducted
    Could also: Reporting the observed coefficient of variation of spike-in-normalized viral read counts across independent experiments would also contextualize reproducibility — Documenting inter-replicate variability helps readers judge whether observed between-cell-type fold-changes are substantially larger than the assay's noise floor
  • Read counts and percentages were reported without measures of dispersion across replicates
    Could also: Reporting SD or range across independent experiments for key TSS-level read count values would also convey reproducibility of the quantitative results — Measures of spread allow assessment of whether between-cell-type differences consistently exceed within-condition variability, which is especially informative for absolute read count comparisons across separate experiments
Software: NOISeq · UCSC Genome Browser

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

FJ616285 ENA in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-40758707

Paper: Hu et al. 2025, PLoS Pathog 21(8):e1013374. "Cell type differences in human cytomegalovirus transcription and epigenetic regulation..." HCMV (Towne strain) in differentiated NTera2 (D-NT2) vs fibroblasts (HFF). Multi-assay: PRO-Seq, RNA-Seq, DFF-ChIP-Seq, + wet-lab (protein degradation, promoter mutation/function assays).

"Code" repo (github.com/meierjl/EP3): NOT analysis code. It is a UCSC Track Hub (hub.txt, genomes.txt, trackDb.txt + 256 binary .bw/.2bit/.bb files = processed pipeline OUTPUTS for genome-browser viewing). No scripts, no Snakefile, no pipeline source. So the authors' own analysis code is not shipped → exact-rerun of their code is impossible. Per brief rule P16, we reproduce by running the described standard pipeline (third-party tools) on the paper's own raw data and comparing to the deposited output.

Data: GEO GSE291633 [DFF-ChIP subseries] = BioProject PRJNA1234562 = SRA SRP569586. Raw paired-end FASTQ public on ENA/SRA; processed viral pileup bigWigs public as GEO per-sample supplementary files (and duplicated in the track-hub repo).

In scope (pipeline-derived, clearly specified, low-hanging) — ATTEMPTED

The GEO data_processing fields fully specify the DFF-ChIP track pipeline:

  1. adapter trim — trim_galore 0.6
  2. align trimmed reads to hg38 + FJ616285.1 (HCMV Towne) — bowtie 1.2.2
  3. dedup fragments by UMI tags — awk/bash (exact script NOT shipped)
  4. deduped BED → bedGraph → bigWig — bedtools genomecov + kentUtils bedGraphToBigWig Expected result = the deposited per-sample viral pileup bigWig (235,147 bp FJ616285.1).

Reproduction claim (the comparison): regenerate the viral pileup bigWig from raw FASTQ for 2 representative D-NT2 DFF-ChIP samples and correlate against the authors' deposited bigWig over the whole HCMV genome.

  • RU-claim-1: GSM8839001 / SRR32649842 — Exp3 Pol II Rep1
  • RU-claim-2: GSM8839005 / SRR32649833 — Exp3 H3K4me3 Rep1 Metric: genome-wide Pearson + Spearman of per-base coverage (mine vs deposited); grade exact/within-tol/partial/mismatch on r. Also: consistency check that GEO-supplementary bigWig == repo track-hub bigWig (same deposit).

Known deviations (the hard ~20%, deliberately not chased)

  • No hg38 co-alignment. We align to FJ616285.1 only. HCMV (235 kb) vs human cross- mapping under bowtie1 is negligible; deposited tracks are virus-only. Skipping the human index avoids a multi-GB bowtie1 build for ~no effect on the viral track.
  • UMI dedup approximated. The authors' awk/bash UMI logic is not in the repo, and UMI position in the SRA reads is undocumented. We dedup by fragment coordinates (and report a no-dedup variant too). This shifts peak magnitudes slightly, not the genome-wide shape.
  • Exact bowtie param flags (seed/mismatch) beyond version are unspecified → tool defaults.

Out of scope (not attempted, why)

  • PRO-Seq / RNA-Seq subseries pipelines and all differential/metaplot/quantitative figure panels — broad, multi-step, many panels are browser screenshots or wet-lab.
  • All wet-lab results (rapid protein degradation, promoter mutagenesis, morphology).
  • Exact peak-calling / TSS / nucleosome-occupancy downstream analyses.
RU-claim-1
Reported
Deposited HCMV viral pileup bigWig GSM8839001_Exp3-Pol-II-Rep1-virus.bw (FJ616285.1, 235147 bp), D-NT2 Native DFF-ChIP Exp3 Pol II Rep1 (SRA SRR32649842)
Reproduced
Regenerated from raw FASTQ via trim_galore 0.6.10 -> bowtie 1.2.2 (FJ616285.1) -> proper-pair fragment genomecov -> bigWig; genome-wide raw-pileup Pearson 0.99645 (1bp) / 0.99870 (100bp), Spearman 0.99700. 43,038,429 read-pairs, 31.19% aligned to virus. Deposited total ~1.96x my raw total (under-specified counting convention, flagged). Regenerated bigWig SHA 51394eca... byte-identical to a prior independent run -> deterministic.
within tolerance
RU-claim-2
Reported
Deposited HCMV viral pileup bigWig GSM8839005_Exp3-H3K4me3-Rep1-virus.bw, D-NT2 Native DFF-ChIP Exp3 H3K4me3 Rep1 (SRA SRR32649833)
Reproduced
Same validated pipeline; genome-wide raw-pileup Pearson 0.99970 (1bp) / 0.99973 (100bp), Spearman 0.99981. 39,990,976 read-pairs, 58.37% aligned to virus. Deposited total ~1.86x my raw total (same counting-convention offset).
within tolerance
RU-check-3
Reported
The GitHub track-hub bigWigs (the 'code' repo meierjl/EP3) and the GEO-deposited supplementary bigWigs for the same samples are the same deposit
Reproduced
BYTE-IDENTICAL SHA256 for both samples, re-verified fresh from pinned commit bab5eea (repo Exp3-PolII-Rep1_virus.bw == GEO GSM8839001 3c7d9edc...; repo Exp3-H3K4me3-Rep1_virus.bw == GEO GSM8839005 39969502...).
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

534.7 k
tokens (I/O) · 41 M incl. cache
150 min
runtime · 2.31 CPU-h
7.3 GB
peak RAM
1
HPC jobs
hummel
machine