De novo identification of CD4+ T cell epitopes.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> partial 1:1 reproduction of the in-scope computational result (the scRepertoire TCR-clonal analysis), with a strong clean core. SABR-II ships a set of standalone scripts; GSE247410 deposits only the scRNA-seq GEX + cell-hashing + 10x VDJ contigs (the SABR-screen raw FASTQ is NOT deposited, so that sub-pipeline is data_unavailable and out of scope). I reproduced the deterministic clonal analysis on the shipped filtered_contig_annotations.csv (AJ31=8wk, AJ36=6wk, AJ39=10wk) two independent ways: (A) a tool-independent CDR3aa pairing and (B) the authors' actual tool, scRepertoire v1.4.0 combineTCR(cells='T-AB', cloneCall='aa') matching the repo's 'Final version 11/17/2022'. RESULT: all 8 high-confidence validated figure clonotypes (Fig 3 HIP-specific TCRs) are present and clonally expanded in the deposited data (both methods agree within 1-4 cells) -> EXACT, no fabrication signal. The expanded-clonotype landscape reproduces (8W=35, 10W=23, 6W=0 at size>=10), in the same direction and ~1.6x the reported 19/16 BEFORE the paper's interactive CD4 CellSelector gate + HTODemux singlet selection; 6W=0 independently explains why only 8wk+10wk mice were screened. NOT attempted (last-20%/out-of-scope): exact 19/16 selected counts (manual non-deterministic CD4 gate + HTO demux), the 7 CD4 Seurat clusters (integration + manual gate, version/seed-sensitive), SABR-screen per-epitope read counts & enrichment scores (raw FASTQ not in GEO), TCR-similarity amplification, monocle3 pseudotime, DESeq2 expanded-vs-unexpanded DEGs, and all wet-lab validation (non-pipeline). Grades provisional; a human reviewer decides via AUDIT.md.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 83assessed: 2026-06-14 ⛓ 0fb0ec522fee
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a modular, cell-based signaling and antigen-presenting bifunctional receptor platform encoding MHC-II/HLA-II molecules with covalently linked peptides (SABR-IIs) enable de novo, high-throughput identification of CD4+ T cell epitopes, including from scRNA-seq-derived TCRs in type 1 diabetes?
- ★ SABR-IIs encode MHC-II/HLA-II molecules presenting covalently linked peptides and induce NFAT signaling upon cognate TCR recognition, providing a readable output for CD4+ T cell antigen discovery. method
- ★ SABR-II libraries presenting endogenous and non-contiguous epitopes enable de novo identification of CD4+ T cell ligands in the context of type 1 diabetes. finding
- ★ The SABR-II design is modular in signaling (CD28-CD3ζ or B cell CD79A/CD79B domains) and deployable across T cells and B cells (including professional APCs such as Daudi B cells). method
- SABR-II function is retained across host species (human Jurkat and mouse 5KC cells). finding
- ★ Class I (SABR-I) and class II (SABR-II) libraries can be combined at a cellular level and screened simultaneously to increase throughput. method
- ★ Experimental antigen discovery can be amplified post hoc by computational approaches, forming an integrated experimental-computational workflow. method
- ★ Islet-infiltrating CD4+ TCRs from NOD mice recognize physiological hybrid insulin peptides (InsC-ChgA, InsC-Iapp HIPs) and InsB9:23 and analogs. finding
- An enrichment score (ES) metric with two-tiered high/low confidence zones calls putative cognate epitopes from SABR-II library screens. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| NFAT-GFP reporter co-incubation assay (flow cytometry for GFP/CD69) | NFAT-GFP Jurkat cells expressing murine SABR-IIs (I-Ab, I-Ad, I-Ag7) co-cultured with TCR-expressing Jurkat cells (BDC2.5, OT-II, 5-4-E8) | SABR-II overexpression presenting epitopes (Ova, ATEG, 2.5mimo) | GFP and CD69 expression frequency | — |
| NFAT-GFP reporter co-incubation assay | NFAT-GFP Jurkat cells expressing HLA-DQ8 SABR-II co-cultured with T1D patient-derived TCRs GSE.6H9, GSE.20D11 | SABR-II presenting InsB9:23 vs control HEL epitope | GFP+CD69+ cell frequency | — |
| Cross-species co-incubation assay | 5KC mouse thymoma cells expressing SABR-IIs | SABR-II expression; B cell signaling domains (CD79A/CD79B) vs CD28-CD3ζ | NFAT signaling | — |
| Surface FAS upregulation assay | Daudi B cells expressing SABR-IIs (CD28-CD3ζ or CD79A/B) | cognate SABR-II:TCR interaction | surface FAS upregulation | — |
| Pooled SABR-II library screen with FACS sorting and amplicon sequencing | NFAT-GFP Jurkat cells expressing I-Ag7 SABR-II library (4,075 epitopes) screened against BDC2.5 and 4-8Ins TCRs | library presentation; sort top 1-2% GFP+CD69+ cells | epitope read counts / enrichment score | Illumina sequencing; pooled oligonucleotide synthesis, ligation-free cloning |
| Combined class I/II SABR library screen | NFAT-GFP Jurkat cells expressing combined HLA-A*0201 (SABR-I) and HLA-DQ8 (SABR-II) libraries, screened against GSE.20D11 TCR | combined library presentation | enrichment score for DQ8 epitopes | Illumina sequencing |
| scRNA-seq with V(D)J enrichment | Thy1.2+TCRβ+ T cells from individual pancreatic islets of 6-, 8- and 10-week-old NOD mice (11 mice, 3 batches) | none (observational); TotalSeq cell-hashing | transcriptomes, clonal expansion, TCR clonotypes | 10x Genomics; analysis with Seurat, scRepertoire |
| Validation co-incubation / cytokine assays | TCR-expressing 5KC reporter cells and splenic CD4+ T cells; single SABR-II in NFAT-GFP Jurkat cells | stimulation with cognate epitope | mIL-2 secretion; CD25 expression; GFP signal | — |
- ▲ Robust GFP and CD69 expression only in correctly paired SABR-II:TCR assays (murine I-Ab/I-Ad/I-Ag7 with OT-II/5-4-E8/BDC2.5)
- ▲ High frequency of GFP+CD69+ cells only when GSE.6H9/GSE.20D11 TCRs interacted with InsB9:23 in HLA-DQ8 SABR-II, not control HEL
- ▲ Top-scoring epitopes for BDC2.5 TCR were known ligands containing the WXRM(D/E) motif
- ▲ 1-2% sort gate represents >50-fold enrichment of cognate epitopes with minimal signal loss >50-fold
- ▲ Cognate InsB9:23 epitope (SHLVEALYLVCGERG) enriched at high confidence from combined class I and II library for GSE.20D11 TCR
- ▲ High-confidence cognate epitopes obtained for eight islet-derived TCRs, recognizing InsC-ChgA HIP, InsC-Iapp HIP and InsB9:23 analogs
- – Two of ten TCRs (TCR11 and TCR30) with low-confidence hits showed confirmation of reactivity 2 of 10
- – scRNA-seq revealed seven distinct CD4+ T cell clusters with clonal expansion evident in clusters 0 and 3-6, correlating with activation/exhaustion markers
- count 4,075 published epitopes (I-Ag7 SABR-II islet epitope library size)
- mean 708 reads per epitope (mean library coverage after sequencing)
- fold_change >50-fold (enrichment of cognate epitopes at 1-2% sort gate)
- count 35 clonally expanded TCRs (19 from three 8-week-old mice, 16 from two 10-week-old mice) (TCRs selected for screening)
- count 11 mice sequenced in three batches (scRNA-seq of islet-infiltrating T cells from 6-, 8-, 10-week-old NOD mice)
- count eight TCRs (TCRs with high-confidence cognate epitopes identified)
- count two out of the ten TCRs (TCR11 and TCR30) (low-confidence hit TCRs confirmed reactive)
- other ~20 min per replicate, three replicates per TCR (sort time for top 1-2% GFP+CD69+ cells)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This methods-focused paper introduces SABR-II, a cell-based platform for CD4+ T cell antigen discovery. The core analytical framework uses a linear regression model applied to sequencing read counts from sorted vs. unsorted SABR-II libraries to derive a per-epitope enrichment score (ES), with two empirical threshold zones (high- and low-confidence) used for hit calling. Single-cell RNA-seq data from NOD mouse islet-infiltrating T cells were processed with Seurat hierarchical clustering and scRepertoire for clonotype integration, with the Morisita–Horn Index used to quantify TCR overlap across clusters. Results are summarized primarily as means with standard deviations from small numbers of biological or technical replicates, and no formal null-hypothesis significance tests with reported p-values are described for the primary screening readout.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Linear regression (ordinary least squares on read counts) | SABR-II library screens: unsorted library read counts modeled to predict expected epitope abundance; residual used to calculate enrichment score (ES) per epitope per TCR | 3 sorted replicates and 3 unsorted replicates per TCR screen | not stated |
| Empirical two-tiered ES threshold (high-confidence and low-confidence zones derived from positive control TCR–pMHC pairs) | Hit calling across all TCRs screened against I-Ag7 and HLA-DQ8 SABR-II libraries (4,075 and defined smaller sets of epitopes) | 8 biological replicates for BDC2.5 positive control; 3 biological replicates for other TCRs | not stated |
| Morisita–Horn Index | Pairwise comparison of TCR sequence overlap across scRNA-seq-defined CD4+ T cell clusters (Fig. 2d and Extended Data Fig. 5f) | 11 mice across 3 batches (6-, 8-, and 10-week-old NOD mice) | na |
| Hierarchical clustering (Seurat) | Unsupervised clustering of islet-infiltrating CD4+ T cells from scRNA-seq data, yielding 7 clusters (Fig. 2a) | T cells from 11 NOD mice pooled across 3 batches | not stated |
| Bioinformatic gating with re-clustering (Seurat UMAP) | CD4+ T cell subset identification and UMAP visualization (Fig. 2a, Extended Data Fig. 5a,b) | 11 mice | not stated |
-
Read counts from sorted and unsorted SABR-II libraries were modeled with ordinary linear regression to calculate expected epitope abundance and derive the enrichment score↳ Could also: Negative binomial count models implemented in DESeq2 or edgeR could also be applied to model read count data from sorted vs. unsorted conditions, producing per-epitope log2 fold-change estimates with associated standard errors and adjusted p-values — Sequencing read counts are typically overdispersed relative to Poisson and are more accurately modeled by negative binomial distributions; DESeq2/edgeR also provide built-in Benjamini–Hochberg FDR-adjusted p-values across the full epitope library, offering a formal statistical framework alongside the empirical ES threshold approach
-
Hit calling relied on empirical two-tiered ES threshold zones calibrated to positive control TCR–pMHC ES values, without a formal multiplicity adjustment across the ~4,075 epitopes tested↳ Could also: A Benjamini–Hochberg false discovery rate correction applied to enrichment statistics across all epitopes in the library could also be used to control the expected proportion of false positives among called hits — With thousands of epitopes tested per TCR, the probability of false positives accumulates; FDR control provides a principled, widely reported complement to empirical threshold strategies and facilitates comparison of stringency across studies
-
Dispersion was reported as SD from small numbers of biological replicates (typically n=3 or n=8)↳ Could also: 95% confidence intervals or SEM could also be reported alongside or instead of SD, particularly for small-n comparisons — With n as small as 3, SD and 95% CI convey overlapping but distinct information; 95% CIs communicate uncertainty around the mean estimate and are directly interpretable for inference, while SD describes the spread of individual observations — both are standard, and reporting choice can affect how variability is perceived
-
Hierarchical clustering was performed in Seurat on scRNA-seq data to define CD4+ T cell clusters↳ Could also: Leiden community detection (also available in Seurat/scanpy) could also be applied, as it is now widely used in scRNA-seq workflows — Leiden clustering optimizes community structure and has been shown to produce more connected, well-separated partitions than some implementations of Louvain or hierarchical approaches; comparing clustering outputs across methods is a common robustness check in scRNA-seq analysis
-
The Morisita–Horn Index was used to quantify TCR clonotype overlap between scRNA-seq clusters↳ Could also: Other repertoire overlap metrics such as the Jaccard index, Bray–Curtis dissimilarity, or the D50 index could also be applied and are implemented in scRepertoire — Different overlap metrics weight rare vs. abundant clonotypes differently; the Morisita–Horn index is abundance-sensitive, while Jaccard is presence/absence-based — using multiple metrics or reporting which metric is used and why aids reproducibility and cross-study comparison
-
Clonal expansion was categorized into discrete bins (single, low 2–9, medium ≥10) and used as the sole criterion for selecting TCRs for antigen discovery↳ Could also: Continuous expansion metrics (e.g., clone frequency or Shannon entropy per cluster) could also be reported alongside categorical bins, and additional selection criteria such as cluster membership or transcriptional activation state could be combined with expansion — Discrete binning of a continuous variable such as clone size discards within-bin variation; continuous measures allow ranking of TCRs by expansion magnitude and may inform prioritization when screening capacity is limited
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Two of ten islet-derived TCRs (TCR11 and TCR30) with low-confidence SABR-II screen hits show confirmed antigen reactivity on secondary validationflow-cytometry jurkat mixed 2024×1papers★ This paper is the founder (earliest)
-
CD69 and NFAT-GFP reporter expression is upregulated only in correctly matched SABR-II:TCR co-cultures using murine MHC II alleles (I-Ab, I-Ad, I-Ag7), confirming allele-specific antigen presentationflow-cytometry jurkat up 2024×1papers★ This paper is the founder (earliest)
-
HLA-DQ8 SABR-II presenting InsB9:23 but not control HEL epitope activates T1D patient-derived TCRs (GSE.6H9, GSE.20D11) as measured by GFP+CD69+ frequencyflow-cytometry jurkat up 2024×1papers★ This paper is the founder (earliest)
-
Sorting the top 1-2% GFP+CD69+ cells in the SABR-II pooled library screen yields greater than 50-fold enrichment of cognate epitopes with minimal signal lossother jurkat up 2024×1papers★ This paper is the founder (earliest)
-
Top-scoring I-Ag7 SABR-II library epitopes for the BDC2.5 TCR contain the WXRM(D/E) motif, matching known chromogranin A-derived BDC2.5 ligandsother jurkat up 2024×1papers★ This paper is the founder (earliest)
-
SABR-II pooled library screen identifies high-confidence cognate epitopes for eight NOD mouse islet-derived CD4+ TCRs, including InsC-ChgA HIP, InsC-Iapp HIP, and InsB9:23 analogsother jurkat up 2024×1papers★ This paper is the founder (earliest)
-
scRNA-seq of NOD mouse islet Thy1.2+TCRb+ T cells identifies seven CD4+ T cell clusters with clonal expansion in clusters 0 and 3-6 associated with activation and exhaustion markersscRNA-seq nod-mouse pancreatic-islet 2024×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38658646 (SABR-II, Zdinak et al., Nat Methods 2024)
Paper: De novo identification of CD4+ T cell epitopes. PMID 38658646 · PMC11093748 · DOI 10.1038/s41592-024-02255-0 Code: https://github.com/joglekar-lab/SABR-II (GPL-3.0, public, last push 2024-04-15) Data: GEO GSE247410 (scRNA-seq + cell-hashing + TCR-seq of NOD mouse islet T cells; BioProject PRJNA1037506)
The repo (a set of standalone scripts, not one pipeline)
| Script | Function | Needs | Reproducible? |
|---|---|---|---|
Zdinak-et-al-seurat+screpertoire.R |
scRNA-seq + TCR clonal analysis (Seurat + scRepertoire) | GSE247410 GEX matrices, HTO (cell-hashing), VDJ contigs | partly — see below |
TCRgen_mouse.opt_v2.py + imgt_tcr_mouse.nuc.fa |
reconstruct full-length TCR insert from V/J alleles + CDR3 | only shipped FASTA | self-contained utility (no paper number to compare) |
backtranslate_fast_noU_upto25.py |
back-translate epitope AA → nt oligo for library synth | an epitope list | self-contained utility (no paper number) |
demultiplex_dual.py, epitope_extract_fastq_v1.1.py, merge_counts_split_v2.1.py |
SABR screen: raw amplicon FASTQ → per-epitope read counts | raw screen FASTQ — NOT in GSE247410 | OUT — data not deposited |
GSE247410 contents (what is actually public)
Per timepoint AJ36=6-week, AJ31=8-week, AJ39=10-week mice, 10x Genomics output:
<AJ>-barcodes.tsv.gz, -features.tsv.gz, -matrix.mtx.gz (GEX + Antibody-Capture/HTO rows),
<AJ>-filtered_contig_annotations.csv.gz (paired TCR VDJ contigs).
The SABR screen FASTQ / per-epitope count tables are not in this GEO series.
IN SCOPE (attempted) — scRepertoire TCR clonal analysis
The deterministic, low-compute core of the R script: combineTCR(filtered_contig_annotations, cells="T-AB", cloneCall="aa") →
per-cell paired clonotype (CTaa = CDR3α_CDR3β), clone size = #cells sharing it. Needs only the shipped VDJ contig CSVs. Targets:
- C2 — "35 clonally expanded TCRs for screening … 19 TCRs from three 8-week-old mice and 16 TCR from two 10-week-old mice" (Results; Supp. Table 3). Reproduce clone-size distribution per sample (AJ31=8W, AJ39=10W) and count expanded clonotypes.
- C3 — the 8 high-confidence validated TCR clonotypes named verbatim in the R script's
highlightClonotypescalls (Ins-ChgA-HIP ×4, Ins-IAPP-HIP ×2, InsB9:23 ×2). Verify each paired CDR3α_CDR3β is present in the shipped VDJ data and is clonally expanded (fabrication check: are the figure clonotypes real, expanded clones?).
OUT OF SCOPE / last-20% (stated, not attempted)
- C1 — "seven distinct CD4+ T cell clusters" (Fig. 2a). Requires full Seurat SCT/integration and an interactive
CellSelectormanual CD4 gate +HTODemuxsinglet selection. The manual gate is irreducibly non-deterministic; clustering is version/seed-sensitive → not a clean 1:1. Skipped per 80/20. - SABR screen per-epitope counts, enrichment scores, Fig 1/3/5 — raw screen FASTQ not deposited (data_unavailable for that sub-result).
- All wet-lab validation (co-culture GFP/CD69 reporter assays, ES zones) — out of scope (non-pipeline).
- TCR-similarity amplification (TCRdist/GLIPH-type), monocle3 pseudotime, DESeq2 expanded-vs-unexpanded DEGs — version-heavy, not headline numbers; skipped.
Pipeline named per in-scope result
scRepertoire combineTCR (cloneCall="aa") on 10x filtered_contig_annotations.csv; cross-checked with a tool-independent data.table pairing to guard against scRepertoire version drift.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The in-scope deterministic computational result reproduced cleanly: all 8/8 high-confidence validated figure clonotypes (Fig 3) are real, paired, clonally expanded clones derivable from the deposited VDJ data, confirmed by the authors' own tool (scRepertoire v1.4.0) and an independent re-implementation (agreement ±1–4 cells), and 6W=0 expansion independently explains the screening choice. The one deviation — whole-sample 35/23 vs the reported gated 19/16 — sits on the input/preprocessing side and is caused by a non-deterministic interactive CD4 CellSelector + HTODemux singlet gate we did not reproduce, not by an authors' defect. Much of the wider pipeline (SABR-screen enrichment) is out of scope because the raw screen FASTQ is not deposited. Overall: solid reproduction with explainable deviations, no fabrication concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.