Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Progressive transformation of the HIV-1 reservoir cell profile over two decades of antiviral therapy.

Cell Host Microbe · 2023
L1 80/100 PQI 93
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 56% of all assessed papers rank 484 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH? Yes for the code, NO for the data. The analysis code (BWH-Lichterfeld-Lab/Intactness-Pipeline, the authors' own in-house Python proviral-intactness classifier) is public, documented, and runnable. The paper's HEADLINE pipeline result — 496 near-full-length HIV-1 proviral genomes classified as 96 genome-intact / 400 defective — is NOT reproducible 1:1 because the authors explicitly restrict the input sequences ('Due to study participant confidentiality concerns, viral sequencing data cannot be publicly released'). The cited GEO accessions (GSE168337 Hi-C, GSE144334 RNA-seq) are re-used from other papers and are not the intactness-pipeline input. Per Brief rule P16 (running an existing tool on suitable data is a valid reproduction), we reproduced the pipeline AS A TOOL: built its environment on «our HPC» and functionally validated each classification branch on deterministic public control sequences derived from the pipeline's own HXB2 reference, with unambiguous expected calls. RESULT: 3/3 defective-control branches reproduce EXACTLY (Large Deletion, Hypermut, NonHIV) and the intact control is correctly identified as a genome-intact candidate at every local stage. WHAT WE DID NOT ATTEMPT / 80-20 SKIP: (1) the exact 96/400 split (input restricted); (2) the premature-stop-codon (PSC) refinement of intact candidates, which depends on the LANL HIV GeneCutter web service via MechanicalSoup==0.11.0 and hangs against the present-day site. REPRODUCIBILITY NOTES: the code only runs after era-faithful version pins (no code edits) — biopython=1.79 (>=1.80 removed SeqFeature(strand=)) and PyPDF2==1.26.0 (>=3.0 removed PdfFileMerger), both crashes in the cosmetic PDF alignment-view step. NO fabrication concern: the restricted-data status is openly stated by the authors and the classifier behaves correctly on every public control. Verdict = PARTIAL (tool reproduced + functionally validated; exact paper counts data_restricted).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 80
    assessed: 2026-06-15 ⛓ 7fff021ab151
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The authors hypothesized that selection of viral reservoir cells with features of deep viral latency, as seen in elite controllers, also occurs in antiretroviral therapy (ART)-treated individuals, and that extended (~two decades) durations of suppressive ART make these selection forces more visible in the HIV-1 reservoir integration site profile.

Core claims
  • After long-term ART, intact HIV-1 proviruses are predominantly integrated in heterochromatin locations, most prominently centromeric satellite/micro-satellite DNA. finding
  • This heterochromatin-biased integration site profile results from longitudinal immune-mediated selection that preferentially eliminates intact proviruses in transcriptionally permissive locations while preserving those in repressive chromatin. mechanism
  • Intact proviruses integrated in satellite DNA/heterochromatin are transcriptionally silent and difficult to reactivate, consistent with a state of deep latency. finding
  • No similar selection process toward heterochromatin is observed for defective proviruses. finding
  • Intact proviruses are also mainly integrated in heterochromatin in post-treatment controllers, suggesting reduced rebound competence. finding
  • MIP-seq enables paired assessment of individual proviral sequences and their chromosomal integration sites. method
  • PRIP-seq enables parallel ex vivo assessment of HIV-1 RNA, integration site, and proviral sequence from single infected cells. method
  • Large clones of intact proviruses persist in KRAB-ZNF genes on chromosome 19, a heterochromatic, H3K9me3-rich region. finding
Experimental setups
Assay System Perturbation Readout Platform
FLIP-seq (near full-length individual proviral sequencing) PBMC from long-term ART-treated individuals, m-ART individuals, and elite controllers none (cross-sectional clinical cohorts) frequency of genome-intact vs defective proviral genomes, clonality, genetic distance, HLA-associated variants
MIP-seq (matched integration site and proviral sequencing) with ISLA PBMC/CD4 T cells from LT-ART individuals none paired proviral sequence and chromosomal integration site coordinates phi29-catalyzed multiple displacement amplification
PRIP-seq (parallel HIV-1 RNA, integration site, and proviral sequencing) single HIV-1-infected cells from study participant LT03 (and prior datasets) physiological in vivo reactivation signals (ex vivo) proviral transcriptional activity linked to integration site
Quantitative viral outgrowth assay (qVOA) isolated CD4 T cells (15.7 million) from 5 participants (LT02, LT03, LT06, LT07, LT08) in vitro stimulation with non-physiological reactivation agents + high-dose exogenous cytokines infectious viral particle production after 14/21 days culture
RNA-seq (reference dataset alignment) primary CD4 T cell subsets (total, effector-memory, central-memory) none host gene transcription / distance of integration sites to TSS
Hi-C sequencing (reference dataset alignment) CD4 T cell reference dataset none 3D chromatin compartments A/B and subcompartments B2/B4
ChIP-seq and ATAC-seq (reference dataset alignment) primary CD4 T cells / ROADMAP consortium reference data none H3K9me3 repressive histone marks and chromatin accessibility around integration sites
Key results
  • Frequencies of intact proviruses in LT-ART participants were significantly higher than in elite controllers, but the intact:defective ratio and intra-individual genetic distance were similar between LT-ART and ECs.
  • Intact proviruses in LT-ART individuals were strongly biased toward centromeric/peri-centromeric satellite/micro-satellite DNA, with such integration observed in 6/8 participants. 32.26% of independent intact proviruses
  • A substantial fraction of intact proviruses were integrated in KRAB-ZNF genes, mostly on chromosome 19. 23.08% of independent intact proviruses
  • Only a small proportion of intact proviruses were located in active transcription units other than ZNF genes, almost exclusively from one participant (LT06). 19.35% of independent intact proviruses
  • The large intact proviral clone integrated in centromeric satellite DNA of chromosome 18 in LT03 was completely transcriptionally silent across all 5 member sequences, whereas genic-integrated defective proviruses (e.g., in SMURF2) were transcriptionally active. 0/5 member sequences active
  • Viral outgrowth was negative for an estimated 307 intact proviruses after 14 days; only 2 proviruses (LT02, LT03) produced outgrowth after 21 days. 2 of ~307 proviruses
  • Intact proviruses from LT-ART individuals were enriched in 3D heterochromatin compartments B2 and B4 and at marked distance from host TSS, resembling the EC landscape.
  • Longitudinal analysis showed intact proviruses at early ART (T1) were frequently in introns of highly expressed genes, shifting toward centromeric satellite/micro-satellite DNA and larger clones over time.
Key statistics
  • count 496 individual proviral genomes (96 intact, 400 defective) (FLIP-seq total proviral genomes from LT-ART individuals)
  • other median 20 years (range 18–23) (duration of continuous suppressive ART in LT-ART cohort)
  • count 31 independent intact provirus integration sites (MIP-seq cross-sectional analysis in LT-ART individuals)
  • count 32.26% (proportion of independent intact proviruses in centromeric/peri-centromeric satellite/micro-satellite DNA)
  • count 23.08% (proportion of independent intact proviruses integrated in ZNF genes)
  • count 19.35% (proportion of independent intact proviruses in active transcription units other than ZNF genes)
  • count ~307 genome-intact proviruses from 15.7 million CD4 T cells (qVOA across participants LT02, LT03, LT06, LT07, LT08)
  • other median 9 years (ART duration of moderate-treatment (m-ART) comparison cohort)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This observational cohort study compared HIV-1 proviral reservoir characteristics (frequencies, integration site profiles, clonality, transcriptional activity) across three groups: long-term ART-treated individuals (LT-ART, n=8, ~20 years), moderately ART-treated individuals (m-ART, median 9 years), and elite controllers (ECs). Cross-sectional group comparisons used FDR-adjusted nonparametric or categorical tests; a longitudinal sub-analysis (n=5 LT-ART participants with archived samples) tracked integration site evolution over time. Results were reported primarily as medians with interquartile ranges and proportions, displayed in box-and-whisker plots and pie charts.

Replicationbiological Sample sizeEight LT-ART participants described by recruitment criteria (continuous suppressive ART for median 20 years, range 18–23; ≤2 viremia blips); no formal power calculation or sample size justification stated in available text GroupsLT-ART (n=8) vs. m-ART (median 9 years ART) vs. elite controllers (ECs); longitudinal sub-analysis in n=5 LT-ART participants at up to three time points Pairingmixed Randomization/blindingnot stated DispersionIQR Effect sizesno Confidence intervalsno Multiplicity correctionFDR adjustment (method not further specified; applied to Kruskal-Wallis and Fisher's exact tests); chi-square tests stated as adjusted for multiple comparison testing
Statistical tests used
Test Applied to n Assumptions
FDR-adjusted two-sided Kruskal-Wallis nonparametric test Comparisons of proviral frequencies, proportions of intact proviruses, genetic distances, clonal proportions, and integration site features across LT-ART, m-ART, and EC groups (Figures 1A–C, 1E, 1G, 2C, 2D) n varies by panel: n=8 study subjects for participant-level measures; n reflects number of viral sequences or integration sites for sequence-level measures (e.g., 496 total proviruses for FLIP-seq analyses, 31 intact integration sites for MIP-seq) not stated
FDR-adjusted Fisher's exact test Categorical comparisons in Figures 1A–H, used 'as appropriate' alongside Kruskal-Wallis n reflects number of viral sequences or study subjects depending on panel not stated
Chi-square test adjusted for multiple comparison testing Proportions of intact proviruses with specific integration site features across groups (Figures 2B–2D) n reflects number of integration sites not stated
Maximum-likelihood phylogenetic analysis Reconstruction of clonal relationships among intact and defective proviruses (Figures 2A, 3A) 496 total proviral sequences from LT-ART individuals (96 intact, 400 defective); additional sequences from m-ART and EC comparators na
Approaches that could also have been used
  • Continuous outcomes (proviral frequencies, genetic distances, clonal proportions) were compared across three groups with the Kruskal-Wallis test
    Could also: A one-way ANOVA with a post-hoc correction (e.g., Tukey HSD or Dunnett's test) could also have been applied to three-group comparisons if normality were plausible — Parametric post-hoc tests yield pairwise confidence intervals and are more powerful under normality; Kruskal-Wallis, as used here, is the more conservative and robust choice when normality is uncertain with small n per group
  • The longitudinal sub-analysis (n=5 participants, up to three time points) was presented descriptively with integration site tracking per participant
    Could also: A linear mixed-effects model or a repeated-measures nonparametric approach (e.g., Friedman test) could also formally test within-person change over time while accounting for the repeated-measures correlation structure — Formal longitudinal modeling would provide a statistical test of the direction and rate of integration site profile change over time, separating within-person trends from between-person variability
  • Dispersion for continuous variables was reported as interquartile range in box-and-whisker plots
    Could also: Standard deviation or 95% bootstrap confidence intervals around the median could also convey spread and uncertainty around the central estimate — With small per-group n (e.g., n=8 LT-ART participants), confidence intervals would additionally communicate the precision of group estimates, aiding interpretation of apparent between-group differences
  • Categorical integration site features (e.g., proportions in heterochromatin vs. euchromatin) were compared with chi-square tests and Fisher's exact tests
    Could also: Logistic regression or Poisson regression models could also have been used to compare proportions while simultaneously adjusting for potential covariates (e.g., ART duration as a continuous variable, CD4 count) — Regression approaches accommodate continuous predictors and potential confounders, which may be informative given that ART duration varied continuously across participants
  • Phylogenetic relationships among proviral sequences were inferred by maximum-likelihood methods
    Could also: Bayesian phylogenetic inference (e.g., BEAST) could also have been applied — Bayesian methods additionally yield posterior distributions on tree topology and branch lengths, and can incorporate molecular clock models to estimate timing of clonal expansions—directly relevant to the paper's hypothesis about longitudinal selection
  • FDR correction was applied to the family of three-group comparisons; the specific FDR method (e.g., Benjamini-Hochberg) was not stated
    Could also: Bonferroni or Holm-Bonferroni correction could also have been applied to the same family of tests — FDR control (as used here) tolerates a specified proportion of false positives among rejected nulls; Bonferroni-type methods control the family-wise error rate more strictly, which may be preferable when any single false positive has high consequence; naming the specific FDR procedure aids reproducibility
Software: Not stated in available text

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE144334 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE168337 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36596305

Paper: Lian et al. 2023, Cell Host & Microbe. "Progressive transformation of the HIV-1 reservoir cell profile over two decades of antiviral therapy." PMID 36596305 / PMC9839361 / DOI 10.1016/j.chom.2022.12.002.

Code: https://github.com/BWH-Lichterfeld-Lab/Intactness-Pipeline (authors' own "in-house proviral intactness bioinformatic pipeline in Python"). Data accessions cited: GSE168337 (Hi-C, 3 CD4 samples — from ref. 8) and GSE144334 (RNA-seq — from ref. 9). Both are re-used from earlier papers, not the core data of this study.

The central pipeline-derived result

The paper's headline computational result is the intactness classification of near-full-length HIV-1 proviral sequences: of 496 single proviral genomes, 96 genome-intact and 400 defective, with defects categorized as large deletions (<8000 bp), out-of-frame indels, premature stop codons (PSC), internal inversions, hypermutation, and packaging-signal/5' defects. This is produced by the Intactness-Pipeline.

In scope vs out of scope

Result Pipeline In scope? Note
Intactness classification (96 intact / 400 defective of 496) Intactness-Pipeline Code in scope; INPUT DATA restricted Proviral sequences "cannot be publicly released" (participant confidentiality). Exact counts NOT reproducible.
Integration-site / chromatin (Hi-C) analysis external (ref. 8 / GSE168337) out data from a different paper; not this paper's pipeline
scRNA-seq reservoir cell profiling external (ref. 9 / GSE144334) out data from a different paper
Wet-lab (QVOA, FISH-flow, ddPCR, integration-site PCR) none out experimental, not pipeline-derived

Decision

The one in-scope, paper-specific pipeline is the Intactness-Pipeline. Its code is fully public and runnable, but the input sequences are restricted (data_restricted), so the exact reported counts (96/400) cannot be reproduced 1:1.

Reproduction approach (per Brief rule P16 — running an existing tool on suitable data is valid): build the pipeline environment on «our HPC» and functionally validate the classifier on deterministic public control sequences derived from the pipeline's own HXB2 reference, with unambiguous expected calls — one per major decision branch the paper relies on:

  • full HXB2 → Intact / Inferred Intact
  • HXB2 with a 3 kb internal deletion → Large Deletion
  • HXB2 with APOBEC3G/F G→A signature → Hypermut
  • deterministic random ACGT → NonHIV

This reproduces the pipeline as a runnable, correct tool (the part that is public) while honestly recording that the exact 96/400 split is not reproducible because the input is confidentiality-restricted.

intactness_split
Reported
496 proviral genomes: 96 genome-intact, 400 defective
Reproduced
not attempted on paper data — near-full-length HIV-1 sequences are confidentiality-restricted ('viral sequencing data cannot be publicly released')
partial
branch_large_deletion
Reported
defective class: large deletion (<8000 bp aligned)
Reproduced
CTRLdel3kb (HXB2 with 3 kb internal deletion, aligned 5958 bp) -> Final Call = 'Large Deletion'
exact
branch_hypermut
Reported
defective class: APOBEC-mediated hypermutation
Reproduced
CTRLhypermut (HXB2 + G->A in GG/GA contexts) -> Hypermut?=Yes -> Final Call = 'Hypermut'
exact
branch_nonhiv
Reported
non-HIV / non-mapping contigs excluded as NonHIV
Reproduced
CTRLnonhiv (random ACGT, 5000 bp) -> Is HIV?=No -> Final Call = 'NonHIV'
exact
branch_intact
Reported
genome-intact provirus call
Reproduced
CTRLintact (full HXB2): Is HIV?=Yes, Large Deletion=No, Inversion=No, Hypermut=No (p=1.0), Primer=Yes -> candidate-intact; final 'Intact' call blocked at the premature-stop-codon step (LANL GeneCutter web service hangs)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

The paper's headline intactness result (496 proviral genomes → 96 intact / 400 defective) is not reproducible 1:1 purely because the input near-full-length HIV-1 sequences are confidentiality-restricted by the authors — an openly stated, legitimate restriction, not a defect or fabrication concern. The public CODE (Intactness-Pipeline) was reproduced as a tool: the Large Deletion, Hypermut and NonHIV branches reproduce exactly on deterministic HXB2-derived controls, and the intact control correctly reaches candidate-intact at every local stage. The only residual gaps are on the data-availability side (restricted sequences → counts not derivable) and dependency/external-service decay (LANL GeneCutter hang; era-faithful pins for cosmetic PDF steps). Net: a solid, explainable partial — the tool is correct and runnable, but the central counts cannot be independently checked.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

216.5 k
tokens (I/O) · 19.4 M incl. cache
24 min
runtime · 0.01 CPU-h
1.4 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine