Spatially clustered loci with multiple enhancers are frequent targets of HIV-1 integration.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH: yes for the self-contained maja/ R analyses, which ship their own derived data objects (genidf.Robj, is.Robj, IS.txt, roc.Robj). Mostly 1:1. Reproduced from shipped data on «our HPC» (R 4.3.3 + GenomicRanges 1.54.1), repo @ 5c35ee8: C1 total integration sites 13,550 EXACT (in-vitro 4031 and patient 9519 subtotals both exact); C3 non-RIG protein-coding genes 13,140 EXACT; C4 super-enhancer-overlapping RIGs count 564 EXACT (reported %34.22 uses denominator 1648; shipped RIG total is 1605 -> 35.14%); C5 ROC Fig 1B directionality 9/9 marks match (enriched/ns/depleted). PARTIAL/DIFFERENT: C2 exact RIG total 1648 not re-derivable because the Brady-redefinition step needs two inputs absent from the repo (GRCh37AllGenes.Robj, IntegrationSitesInActivatedALLCells_bushman.txt); shipped genidf yields 1605 -- the 43-gene gap reconciles consistently with C4 (all added RIGs are non-SE). NOT ATTEMPTED: Hi-C .cool generation + A1/A2/B1/B2/AB sub-compartment clustering (Fig 4, clusters/ branch, large GSE122958 cool) -- deferred heavy last-20%; TableM1 %-IS-in-genes (blocked, un-shipped inputs); wet-lab + ngsplot metagene (non-pipeline); full independent ROC recompute (controls object structure mismatch). No fabrication indicators: the hardest-to-fake counts reproduce to the digit.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 77assessed: 2026-06-14 ⛓ 14e7b69acd83
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhat are the genomic and spatial nuclear features of genes recurrently targeted by HIV-1 integration in CD4+ T cells, and is the transcriptional activity of these genes the determinant of integration, or is it their spatial organization relative to super-enhancers?
- ★ HIV-1 recurrently integrates into genes that are proximal to super-enhancer (SE) genomic elements in both patients and in vitro T cell cultures. finding
- ★ Recurrent integration genes (RIGs) are proximal to SEs irrespective of their transcriptional levels, and disruption of SE activity (JQ1) does not alter HIV-1 integration patterns. finding
- ★ HIV-1 insertion hotspots cluster together in 3D nuclear space and preferentially contact super-enhancers, occupying the same 3D sub-compartment. finding
- ★ Spatial clustering of targeted genes together with their transcriptional activity are the major determinants of HIV-1 integration, with SEs contributing indirectly via genome reorganization during T cell activation. mechanism
- RIGs were defined as genes with ≥1 HIV-1 integration in at least 2 of 8 datasets, yielding 1648 RIGs, used to compare chromatin/expression features against non-RIGs. method
- RIGs show higher H3K27ac, H3K4me1, H3K4me3, BRD4, MED1, H3K36me3, and H4K20me1 and lower H3K27me3/H3K9me2 relative to non-RIGs. finding
- A high-resolution Jurkat Hi-C dataset (~1.5 billion contacts) was generated and used to segment the genome into spatial sub-compartments (A1, A2, B1, B2, AB). resource
- HTLV-1 integration sites are not enriched in SE marks whereas MLV shows strong enrichment, distinguishing HIV-1's SE-proximal bias. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| HIV-1 integration site mapping (compiled datasets + inverse PCR + linear amplification-mediated PCR) | activated primary CD4+ T cells (in vitro infection) and HIV-1 patient samples | HIV-1 infection | genomic location/frequency of proviral integration sites | — |
| ChIP-Seq | primary CD4+ T cells (activated) | none | genomic levels of H3K27ac, H3K4me1, H3K4me3, BRD4, MED1, H3K36me3, H4K20me1, H3K9me2, H3K27me3; SE identification | — |
| RNA-Seq | CD3/CD28-activated primary CD4+ T cells | none (baseline expression) | transcript abundance (regularized log read counts) of protein-coding genes | — |
| RNA-Seq (transcriptional profiling) | activated CD4+ T cells | JQ1 (BET/BRD4 bromodomain inhibitor) | differential expression of SE-proximal vs non-SE genes, RIGs vs non-targeted | — |
| HIV-1 integration site mapping by inverse PCR | CD4+ T cells | JQ1 treatment vs control | HIV-1 insertion profiles (38,964 sites mapped) | — |
| Hi-C | uninfected Jurkat lymphoid T cells | none | genome-wide chromosomal contact frequencies; TADs, loop domains, A/B compartments, 3D sub-compartments; inter-chromosomal contact density | — |
| RNA and protein quantification (MYC control for JQ1) | CD4+ T cells | JQ1 treatment | MYC RNA and protein levels | — |
- ▲ The more datasets a RIG is found in, the closer it lies to super-enhancers on average.
- ▲ 21.4% of protein-coding genes targeted by HIV-1 are in the top 10% most expressed genes vs 6.07% of non-targeted genes. 21.4% vs 6.07%
- ▲ 19.05% of silent RIGs have a proximal SE, versus only 1.5% of silent genes never targeted by HIV-1. 19.05% vs 1.5%
- – 2584 SEs identified, intersecting 564 RIGs (34.22%). 564/1648 = 34.22%
- ▲ Loci most targeted by HIV-1 engage in stronger inter-chromosomal Hi-C contacts with each other than non-targeted loci, and SEs cluster with HIV-1 hotspots in 3D.
- – JQ1 treatment does not alter HIV-1 insertion biases at chromosome scale nor spatial localization of provirus/RIGs.
- ▲ Insertion rate per chromosome is similar between primary T and Jurkat cells, with ~3-fold increase on chromosomes 17 and 19. ~3-fold (chr17, chr19)
- – HTLV-1 insertion sites not enriched in SE marks; MLV strongly enriched in all SE marks.
- count 4031 HIV-1 integration sites from in vitro activated primary CD4+ T cells (in vitro infection integration sites assembled)
- count 9519 insertion sites from 6 HIV-1 patient studies (patient-derived integration sites)
- count 10,735 integrations in gene bodies (77% patient avg, 84% in vitro avg) targeting 5601 genes (integrations within gene bodies)
- count 1648 RIGs (recurrent integration genes defined (≥1 integration in ≥2 of 8 datasets))
- count 13,140 non-RIG protein-coding genes (comparison set without HIV-1 insertions)
- pvalue <2.2 x 10^-16 (Wilcoxon rank-sum test of mRNA abundance difference for genes without HIV integrations and genes on one list (SE vs no SE))
- pvalue 3.7 × 10^-12 (Wilcoxon rank-sum test of mRNA abundance difference for RIGs (SE vs no SE))
- count ~1.5 billion informative Hi-C contacts (Jurkat Hi-C dataset depth)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a genomics/computational study characterizing HIV-1 integration site preferences in CD4+ T cells, integrating published and newly generated integration-site datasets with ChIP-Seq, RNA-Seq, and Hi-C data. Group comparisons of genomic/expression features between recurrently targeted genes (RIGs) and non-targeted genes were assessed mainly with the non-parametric Wilcoxon rank-sum test, enrichment of chromatin marks at integration sites was quantified using ROC area analysis against distance-matched control sites, and distributions were displayed primarily as box and violin plots. Exact p-values were reported for the key expression comparisons.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon rank-sum (Mann-Whitney) test | differences in median mRNA abundance between genes with vs without a proximal super-enhancer, across gene groups (Fig. 2c) | expression averaged over three RNA-Seq replicates; gene-level n not stated numerically | not stated |
| ROC area (area under ROC curve) analysis | co-occurrence/enrichment of integration sites with each ChIP-Seq epigenetic mark vs distance-matched control sites (Fig. 1b) | — | na |
| Inter-chromosomal Hi-C contact density comparison across gene classes (box-plot distributions) | contact strength among active/silent, HIV/No-HIV, SE/No-SE gene aggregates (Fig. 3d, e) | — | not stated |
-
Median mRNA abundance between gene groups was compared with the non-parametric Wilcoxon rank-sum test.↳ Could also: A permutation/bootstrap test, or a generalized linear model on counts (e.g., negative-binomial regression as in DESeq2) with expression group and SE status as covariates. — A modeling approach would also estimate effect sizes and let one jointly account for expression level and SE proximity, complementing the rank-based two-group test.
-
Enrichment of chromatin marks at integration sites was quantified using the ROC area method against distance-matched control sites.↳ Could also: Logistic regression or a precision-recall (PR) analysis could also summarize the same association. — PR curves are often informative under class imbalance, and a regression framework would allow simultaneous adjustment for several genomic covariates while still yielding an enrichment estimate.
-
Multiple group comparisons (e.g., several gene groups in Fig. 2c) were each reported with a p-value.↳ Could also: A family-wise or false-discovery-rate adjustment (e.g., Benjamini-Hochberg, Bonferroni, or a Kruskal-Wallis test with post-hoc comparisons) could also be applied across the related comparisons. — An explicit multiplicity adjustment controls the error rate across the family of tests and is commonly reported when several related comparisons are made.
-
Spread of distributions was conveyed primarily through box plots and violin plots.↳ Could also: Reporting accompanying effect-size measures (e.g., Cliff's delta, rank-biserial correlation, or median differences with 95% confidence intervals) could also be included. — Effect sizes and CIs convey the magnitude and precision of differences, which complements significance values, especially given very small p-values on large gene sets.
-
Integration-site enrichment used control sites matched on distance to the nearest gene.↳ Could also: Multiple matched control sets or a covariate-balanced random-control resampling scheme could also be used. — Repeated random control draws would provide an empirical null distribution and a sense of variability around the enrichment estimates.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
MLV integration sites are strongly enriched for all super-enhancer chromatin marks, while HTLV-1 integration sites show no enrichment for SE marksChIP-seq human cd4-t-cell mixed 2019×1papers★ This paper is the founder (earliest)
-
34.22% of HIV-1 recurrently integrated genes (RIGs) intersect with super-enhancers identified in activated CD4+ T cells (564 of 1648 RIGs)ChIP-seq human cd4-t-cell 2019×1papers★ This paper is the founder (earliest)
-
Silent HIV-1-targeted genes are enriched for proximal super-enhancers compared to silent non-targeted genes (19.05% vs 1.5%)ChIP-seq human cd4-t-cell up 2019×1papers★ This paper is the founder (earliest)
-
HIV-1 integration hotspots show stronger inter-chromosomal contacts with each other than non-targeted loci, and super-enhancers spatially co-cluster with HIV-1 hotspots in 3D nuclear spaceHi-C human jurkat up 2019×1papers★ This paper is the founder (earliest)
-
HIV-1 integration rate is approximately 3-fold elevated on chromosomes 17 and 19 in both primary CD4+ T cells and Jurkat cellsother human cd4-t-cell up 2019×1papers★ This paper is the founder (earliest)
-
JQ1 (BRD4/BET bromodomain inhibitor) does not alter HIV-1 integration bias at chromosome scale or the spatial localization of RIGsother human cd4-t-cell none 2019×1papers★ This paper is the founder (earliest)
-
HIV-1 integration sites that recur across more independent datasets are located progressively closer to super-enhancers in activated CD4+ T cellsother human cd4-t-cell up 2019×1papers★ This paper is the founder (earliest)
-
HIV-1-targeted protein-coding genes (RIGs) are enriched among the top 10% most highly expressed genes compared to non-targeted genes (21.4% vs 6.07%)RNA-seq human cd4-t-cell up 2019×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Counts reproduce to the digit (13,550 with exact 4031/9519 split; 13,140 non-RIGs; 564 SE-overlapping RIGs) and Fig 1B ROC directionality matches 9/9, so the core claim holds and there is no fabrication signal. The only real deviation is the exact RIG total 1648 vs shipped 1605 — not re-derivable because two redefinition inputs (GRCh37AllGenes.Robj, bushman activated-cell IS) were never shipped, also explaining the C4 % drift (34.22% uses 1648 as denominator vs 35.14% on 1605). The gap is on the authors'/deposit side (incomplete repo, value not derivable from shared data) but is small (+43 genes, all non-SE) and reconciles consistently rather than contradicting, so overall a yellow-quality reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.