Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Depletion of HIV reservoir by activation of ISR signaling in resting CD4+T cells.

iScience · 2022
L1 89/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
89/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 77% of all assessed papers rank 246 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

RE-QUEUE RESOLVED with full pipeline compute. Single in-scope pipeline (10x scRNA-seq; authors' own repo jeremymsimon/Li_HIV_HA15 @4108fa2 on GEO GSE210824). Unlike the prior pass (which only read the authors' deposited labels), this pass INDEPENDENTLY RE-EXECUTED the entire seurat.R pipeline on «our HPC» (Seurat 4.4.0) from the raw deposited alevin-fry GEX+HTO count matrices: USA-mode load (scRNA=S+A), 10x 3M-february-2018 barcode translation (reproduced the authors' 9951 GEX∩HTO intersect EXACTLY; raw intersect is only 2), HTODemux, miQC+count QC, SCT v2 per-sample, integrate 5000 features, PCA(300)/UMAP/FindClusters(res=0.25,algo=2), DESeq2 diff-abundance. Results: 9951 common cells (exact), 5292/5312 QC cells (99.6%, within-tol; deposited idents confirm 5312), 8 clusters (EXACT), clusters 0-3 = 83.6% (exact), 6-sample demux with sizes within ~6 cells (HA15_Rep3 exact), and the key Cluster5 HA15-depletion reproduced BOTH deterministically (-0.834435/0.0276248, matching the authors' embedded table to every printed digit) AND independently in the re-run (Cluster5 -0.808/0.0271, sig). Adjusted Rand Index between the independent re-run clustering and the authors' deposited labels = 0.892 (strong). The ~20-cell QC difference and <=6-cell per-cluster differences are fully explained by RNG in HTODemux (kmeans) and miQC (flexmix), for which the authors set no global seed. NOT attempted: alevin-fry-from-FASTQ (deposited matrices are the documented inputs) and Fig 7B Distinct/gProfiler2/ferroptosis (qualitative only, no pinnable numbers). No fabrication indicators. Verdict: 1:1 reproduced, described well enough. Grades provisional pending human audit.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-15 ⛓ 822550d8f962
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether activation of the integrated stress response (ISR)/ATF4 signaling pathway can reverse HIV latency and selectively deplete HIV-infected (reservoir) CD4+ T cells, including in resting CD4+ T cells from ART-suppressed people living with HIV.

Core claims
  • Activation of ISR/ATF4 signaling (via HA15) reverses HIV latency in Jurkat models of HIV latency (J-Lat A1, 2D10) finding
  • Prolonged ISR/ATF4 activation reduces the frequency of HIV+ (GFP+) cells in the primary CD4+ T cell model of latency after an initial transient increase in HIV transcription finding
  • ISR activation selectively induces apoptosis/cell death in HIV+ cells with minimal effect on HIV-negative CD4+ T cells finding
  • Ex vivo ISR/ATF4/CHOP activation reduces HIV proviral DNA in resting CD4+ T cells from ART-suppressed PLWH finding
  • Ex vivo ISR/ATF4 activation reduces replication-competent HIV (viral outgrowth) in resting CD4+ T cells from ART-suppressed PLWH finding
  • Blocking the ISR/ATF4 negative feedback loop (eIF2/GADD34/CReP) with ritonavir enhances HA15-induced reduction of HIV+ cells finding
  • HA15-induced cell death occurs predominantly in HIV RNA+ cells, linking HIV RNA induction to apoptosis-driven reservoir depletion mechanism
  • ATF4 expression correlates with reciprocal decline of SIV viral loads in intestinal tissue of acutely infected rhesus macaques in vivo finding
Experimental setups
Assay System Perturbation Readout Platform
Flow cytometry (GFP%, viability) Jurkat models of HIV latency (J-Lat A1, 2D10 cells) HA15 (ISR/ATF4 agonist) GFP+ percentage (HIV transcription), cell viability
Flow cytometry (GFP%, Annexin V, cleaved PARP1) and qPCR Primary CD4+ T cell model of HIV latency (pNL4.3Δ6-GFP infected) HA15 GFP+ cells, HIV RNA levels, apoptosis markers
Trypan blue staining and flow cytometry (Annexin V) HIV-negative primary CD4+ T cells HA15 total cell number, viability, apoptosis
Western blot Primary CD4+ T cell model of latency and HIV-negative primary CD4+ T cells HA15 ATF4, CHOP, cleaved Caspase3, cleaved LC3B protein expression
Trypan blue staining and flow cytometry (Annexin V) Resting CD4+ (rCD4+) T cells from ART-suppressed PLWH HA15 total cell number, viability, apoptosis
RT-qPCR Patient primary CD4+ T cells (PLWH on ART) HA15 ATF4 and CHOP mRNA expression
Droplet digital PCR (ddPCR) Resting CD4+ T cells from ART-suppressed PLWH HA15 HIV env proviral DNA levels
Quantitative viral outgrowth assay (QVOA) with RT-ddPCR/ddPCR Resting CD4+ T cells from PBMCs of ART-suppressed PLWH HA15 followed by PHA reactivation cell-associated HIV gag RNA/DNA, cell-free HIV gag RNA in outgrowth culture
Key results
  • HA15 increased the percentage of GFP+ cells in Jurkat J-Lat A1 and 2D10 latency models up to 50% GFP+ cells (2D10)
  • GFP+ cell frequency in the primary CD4+ T cell latency model decreased after HA15 treatment 51% reduction at day 2, 71% reduction at day 4
  • HIV RNA transiently increased 12h post-HA15 then returned to baseline by 24h
  • Apoptosis markers increased in the primary CD4+ T cell latency model after HA15 Annexin V+ up to 26.8%; cleaved-PARP1+ up to 29.55% at 40 μM HA15
  • HIV-negative CD4+ T cells showed no reduction in total cell number and minimal apoptosis after HA15
  • HIV proviral DNA reduced in rCD4+ T cells from ART-suppressed PLWH after HA15 up to 19.27-fold reduction
  • Viral outgrowth assay showed reduced cell-associated and cell-free HIV RNA/DNA after HA15 up to 340-fold (outgrowth RNA), 2378-fold (outgrowth DNA), 1917-fold (supernatant HIV)
  • Co-treatment of HA15 with ritonavir further reduced GFP+ cells compared to HA15 alone 47% reduction (HA15 alone) vs 80% reduction (HA15+ritonavir)
Key statistics
  • fold_change 51% reduction in GFP+ cells (primary CD4+ T cell latency model, 2 days post-HA15)
  • fold_change 71% reduction in GFP+ cells (primary CD4+ T cell latency model, 4 days post-HA15)
  • count 26.8% Annexin V+ cells (apoptosis induction in primary latency model after HA15)
  • count 29.55% cleaved-PARP1+ cells (primary latency model, 40 μM HA15)
  • fold_change up to 19.27-fold reduction in HIV DNA (rCD4+ T cells from ART-suppressed PLWH, ddPCR after HA15)
  • fold_change up to 340-fold reduction in outgrowth HIV RNA (viral outgrowth assay in rCD4+ T cells from PLWH)
  • fold_change up to 2378-fold reduction in outgrowth HIV DNA (viral outgrowth assay in rCD4+ T cells from PLWH)
  • pvalue p=0.0054 (comparison of HIV RNA+/cleaved PARP1+ vs HIV RNA-/cleaved PARP1+ cell percentages)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combined in vitro Jurkat latency models, a primary CD4+ T cell model of HIV latency, ex vivo resting CD4+ T cells from ART-suppressed PLWH, and single-cell RNA-seq to investigate ISR/ATF4 activation as an HIV reservoir clearance strategy. Quantitative outcomes (GFP+ cell frequency, HIV DNA/RNA, apoptosis markers, ATF4/CHOP expression) were compared between HA15-treated and control conditions using one-way ANOVA and two-tailed t-tests; a linear regression assessed the ATF4–SIV RNA correlation in macaque tissue. Results from patient samples (n=4–5 donors) and in vitro replicates (n=3) are reported alongside fold-reduction summaries.

Replicationmixed Sample sizen=3 biological replicates for most in vitro primary cell experiments; n=4 PLWH donors for rCD4+ T cell viability/apoptosis (Figure 3A–C); n=3–5 PLWH donors for HIV DNA/RNA and VOA experiments (Figures 3G, 4); n=5 rhesus macaques for SIV correlation GroupsHA15 (20 or 40 µM) vs. DMSO control; HA15 alone vs. HA15 + Ritonavir or Sephin1; HIV-latently-infected vs. HIV-negative primary CD4+ T cells; rCD4+ T cells from PLWH vs. HIV-negative donors Pairingunclear Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
one-way ANOVA Comparison of HA15 dose/time-point groups for GFP+ cell frequency, cell viability, and apoptosis in primary CD4+ T cell latency model and drug-combination experiments (Figures 1, 5) n=3 not stated
two-tailed t-test HA15 vs. DMSO control comparisons for GFP+ frequency, HIV RNA, Annexin V, cleaved PARP1, ATF4/CHOP expression, and HIV RNA+/cleaved PARP1+ vs. HIV RNA−/cleaved PARP1+ cells (Figures 1, 3, 6) n=3 (in vitro biological replicates); n=3–5 (PLWH donor rCD4+ T cell samples, varies by panel) not stated
linear regression Correlation of ATF4 gene expression and SIV RNA in rhesus macaque intestinal tissue (Figure 3D) n=5 macaques not stated
Approaches that could also have been used
  • Multiple t-tests and one-way ANOVAs were performed across doses, time points, and cell types without a stated multiplicity correction
    Could also: A post-hoc correction applied to each ANOVA (e.g., Dunnett's test for all-vs-control comparisons, or Tukey HSD for all pairwise comparisons) would also explicitly control the family-wise error rate within each experiment — Stating and applying a post-hoc correction makes type I error control transparent to readers when multiple conditions are tested within the same experiment, and is standard in journals reporting multi-group comparisons
  • Linear regression was used to assess the ATF4–SIV RNA correlation in n=5 macaques (Figure 3D)
    Could also: Spearman rank correlation could also be used for this small-sample association, as it does not assume bivariate normality or a linear relationship — With only 5 data points, parametric distributional assumptions are difficult to verify; a non-parametric rank-based approach is commonly reported alongside or instead of regression for small n
  • Most p-values are reported as threshold symbols rather than exact values
    Could also: Exact p-values could also be reported throughout, consistent with the one instance where an exact value was given (p=0.0054) — Exact p-values allow readers to gauge the strength of evidence directly, facilitate downstream meta-analyses, and align with current reporting guidelines (e.g., Nature Research reporting standards, APA)
  • Fold-reductions in HIV DNA/RNA across patient samples are summarized as maximum values (e.g., 'up to 19.27-fold', 'up to 2378-fold')
    Could also: Geometric means with 95% confidence intervals, or individual patient-level paired data with a Wilcoxon signed-rank test, could also characterize the central tendency and variability of the reduction across the cohort — For log-scale outcomes such as viral copy numbers, geometric means and CIs summarize the typical response and its precision across donors more completely than maximum-only descriptors
  • The ex vivo comparisons (treated vs. control rCD4+ T cells from the same PLWH donors) were analyzed with two-tailed t-tests without specifying whether a paired or unpaired design was used
    Could also: A paired t-test or Wilcoxon signed-rank test (non-parametric paired equivalent) could also be applied if treated and control cells derive from the same donor isolation — A paired design accounts for inter-donor variability in baseline HIV reservoir size, which can improve sensitivity with small patient cohorts (n=3–5); the signed-rank test would also be appropriate if normality cannot be assumed at this n
  • The dispersion metric for error bars is not explicitly labeled in the available text
    Could also: Standard deviation (SD) or 95% confidence intervals could also be used and clearly labeled in figure legends — For small-n experiments (n=3), SD conveys the spread of the raw replicate data while 95% CIs convey estimation uncertainty; explicit labeling allows readers to distinguish these interpretations, and both are preferred over unlabeled error bars
Software: not stated

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
9
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36590168

Paper: Li et al. 2022, iScience — "Depletion of HIV reservoir by activation of ISR signaling in resting CD4+ T cells." PMCID PMC9800255 · DOI 10.1016/j.isci.2022.105743 Code: https://github.com/jeremymsimon/Li_HIV_HA15 (commit 4108fa2, author's own code — P16 own repo) Data: GEO GSE210824 (10x 3' scRNA-seq, NovaSeq 6000, HTO-multiplexed: NegControl×3 + HA15×3)

Pipeline map

The paper has wet-lab assays (flow, qPCR, latency reversal, viability) out of scope, and ONE bioinformatic pipeline (in scope):

alevin-fry.sh (alevin-fry GEX+HTO quantification, splici hg38+GENCODEv36 + custom HIV pseudo-chromosome) → seurat.R (Seurat SCT-v2 per-sample → integrate 5000 feat → PCA 300 → UMAP/clusters dims1:30, res=0.25 algo=2 → HTODemux CLR q=0.99 → miQC posterior 0.75 + nCount>2000 & nFeature>1000 → presto markers → DESeq2 differential abundance of cluster proportions → Distinct DE → gProfiler2 → ferroptosis heatmap).

In-scope results attempted (the clearly-specified 80%)

Reproduced from the deposited processed outputs (the authors' actual results: GSE210824_ClusterIdents.txt, GSE210824_SampleIdents.txt, GSE210824_SCTnormalized.txt) plus the documented downstream step:

# Reported result Source How reproduced
C1 5,312 cells pass QC seurat.R inline; Methods QC count rows of deposited ClusterIdents
C2 8 distinct clusters Fig 7A text distinct labels in ClusterIdents
C3 clusters 0–3 contain most cells Fig 7A per-cluster cell counts
C4 6-hashtag demux = NegControl×3 + HA15×3 design/Methods distinct labels in SampleIdents
C5 DESeq2 diff. abundance: Cluster5 log2FC=-0.834, padj=0.0276, only sig. cluster seurat.R embedded results table re-run DESeqDataSetFromMatrix(clusters×samples, ~Treatment), DESeq(fitType="mean"), contrast HA15 vs NegControl — deterministic recompute from deposited cell-level idents

Out of scope / NOT attempted (the hard ~20%)

  • Full Seurat re-run from raw alevin-fry counts (GSE210824_RAW.tar → SCT integration → Louvain clustering). Not attempted: cluster assignment from SCT-integration + Louvain is sensitive to exact Seurat/SCT/igraph versions and RNG; it will not be byte-identical and a near-match adds little over the fact that the deposited idents ARE the authors' own pipeline output and the downstream differential-abundance number reproduces them exactly.
  • alevin-fry quantification from FASTQ (SRA PRJNA867681, tens of GB) — same reasoning; the deposited count matrices (RAW.tar) are the documented inputs to the Seurat step.
  • Distinct DE / gProfiler2 / ferroptosis heatmap (Fig 7B): qualitative claims ("ATF4 up in clusters 0,2,3,5"); the paper gives no exact DEG counts/values to grade against.
Figures / tables: Fig 7As table
C0
Reported
9951 GEX∩HTO common cells
Reproduced
9951
exact
C1
Reported
5,312 cells pass QC
Reproduced
5292 (independent full re-run, 99.6%); 5312 (deposited idents)
within tolerance
C2
Reported
8 clusters
Reproduced
8 clusters
exact
C3
Reported
clusters 0-3 contain most cells
Reproduced
clusters 0-3 = 4422/5292 = 83.6%
exact
C4
Reported
6 hashtags = NegControl x3 + HA15 x3
Reproduced
6 samples; per-sample sizes within ~6 cells (HA15_Rep3 exact)
within tolerance
C5
Reported
DESeq2 Cluster5 log2FC=-0.834, padj=0.0276 (only sig)
Reproduced
deterministic: -0.834435/0.0276248 (only sig); independent re-run: Cluster5 -0.808/0.0271 (sig)
exact
C6
Reported
(cross-check, not a paper claim)
Reproduced
ARI=0.892 between re-run clusters and deposited labels
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 89/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

Clean 1:1 reproduction. The authors deposited their actual pipeline outputs (ClusterIdents/SampleIdents/SCTnormalized) on GEO GSE210824, against which C1-C4 (5312 QC cells, 8 clusters, clusters 0-3 = 83.6%, 6 HTO samples) match exactly, and the one precise embedded statistic — Cluster5 DESeq2 log2FC=-0.834, padj=0.0276 — was deterministically recomputed and reproduced to every printed digit. The deviation is none/rounding only, the central conclusion (Cluster5 the sole differentially abundant cluster) holds fully, and there are no fabrication indicators. Caveat (not a defect): the version/RNG-sensitive full Seurat re-run from raw counts and the qualitative Fig 7B analyses were out of scope and not attempted, but no precise reported numbers there were left ungraded.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

353.8 k
tokens (I/O) · 26.9 M incl. cache
63 min
runtime · 0.18 CPU-h
25 GB
peak RAM
4
HPC jobs
hummel
machine