Loss of CD4+ T cell-intrinsic arginase 1 accelerates Th1 response kinetics and reduces lung pathology during influenza infection.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values are derivable from the shared data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the one pipeline-derived result (Fig 7P microarray GSEA) from the authors' own shipped inputs. NOTE: brief's accession GSE112244 is wrong (that's PMID 30392958); the paper's real microarray series is GSE229775, but raw IDATs were not needed because the repo (jackbibby1/E-West @ c7c10a1) ships the post-normalization GSEA inputs (stim_gsea.txt, labs.cls). Ran Broad GSEA 4.3.3 standard two-class (healthy vs arg1, REACTOME C2:CP, Diff_of_Classes metric since Signal2Noise errors at n=2/class, gene_set permutation). Result: REACTOME_INTERLEUKIN_10_SIGNALING is the #1 most-enriched pathway overall (strong match to 'among the top pathways perturbed'); REACTOME_METABOLISM_OF_AMINO_ACIDS_AND_DERIVATIVES is enriched in patients and nominally significant (p<0.05, FDR 0.117) but mid-ranked (83/516) rather than top-10 — directionally consistent with ARG1/arginine biology; its exact rank is sensitive to the unrecorded GSEA metric + MSigDB release. Paper prints no NES/FDR for Fig 7P, so this is a qualitative reproduction. NOT attempted: re-deriving normalised_data.csv from raw IDATs (not in repo), and all wet-lab/mouse in-vivo results (not pipeline-derived). No fabrication indicators: the hedged claim is supported by re-running the authors' own data.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 68assessed: 2026-06-15 ⛓ a0b29d4e6440
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates whether CD4+ T cell-intrinsic arginase 1 (Arg1), previously thought to be restricted to M2 macrophages, is expressed in and regulates the kinetics of the Th1 CD4+ T cell response and associated tissue pathology during influenza infection.
- ★ Arg1 is highly and specifically induced in lung CD4+ T cells during in vivo influenza infection finding
- ★ CD4+ T cell-intrinsic Arg1 deletion accelerates Th1 effector expansion and subsequent contraction, reducing lung pathology while preserving viral clearance finding
- ★ Arg1-deficiency causes altered glutamine metabolism and reduced TCA cycle flux, distinct from Arg2-deficiency mechanism
- ★ Rebalancing the perturbed glutamine flux in Arg1-deficient cells normalizes the Th1 response mechanism
- ★ CD4+ T cells from ARG1-deficient patients or CRISPR-Cas9 ARG1-deleted healthy donor cells phenocopy the murine Arg1 CKO Th1 phenotype finding
- Arg2-deficiency in CD4+ T cells restricts Th17 and Th2 cytokine production rather than Th1, distinguishing it functionally from Arg1 finding
- Arg1 CKO CD4+ T cells generate normal levels of ornithine and polyamines despite loss of the arginine-to-ornithine conversion enzyme finding
- ★ CD4+ T cell-intrinsic Arg1 functions as a rheostat pacing the transition of Th1 cells from induction to contraction by balancing glutamine versus arginine usage mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA-seq | flow-sort-purified lung CD4+ T cells (CD11a+CD49d+TCRb+CD4+) vs splenic CD4+ T cells, WT mice | PR8 influenza infection vs uninfected/negative control tissue | differential gene expression | — |
| flow cytometry (immunophenotyping) | lung CD4+ and CD8+ T cells, WT vs CD4cre+ Arg1 fl/fl (Arg1 CKO) mice | PR8 influenza infection; T cell-specific Arg1 conditional knockout | T cell numbers/frequency, Ki67+ proliferation, tetramer (NP311-325) binding, T-bet, cytokine production (IFN-γ, IL-2, TNF-α, IL-10, IL-17A) | — |
| protein expression analysis (flow cytometry-based) | CD4+, CD8+ T cells and CD11b+ macrophages/neutrophils, WT vs Arg1 CKO mice | T cell-specific Arg1 deletion | ARG1 protein expression | — |
| histopathology | lung tissue, WT vs Arg1 CKO mice | PR8 influenza infection | lung inflammation and pathology score, viral titer | — |
| adoptive cell transfer | congenically labeled WT and Arg1 CKO CD4+ T cells + WT CD8+ T cells into Rag KO mice | Arg1 CKO cell transfer; homeostatic proliferation and influenza infection | proliferation, CD11a+CD49d+ frequency | — |
| T cell transfer colitis model | naive CD45RBhi CD25- CD4+ T cells from WT or Arg1 CKO mice into Rag1-deficient recipients | Arg1 CKO donor cells | colitis inflammation and pathology severity score | — |
| in vitro CD3/CD28 activation with cytokine profiling and RNA-seq | splenic CD4+ T cells from WT, Arg1 CKO, and Arg2 KO mice | Arg1 or Arg2 knockout; in vitro TCR stimulation, Th1 polarization | proliferation, cell death, cytokine production (IFN-γ, IL-10, IL-17A, IL-5, IL-13, IL-4), gene expression, FoxP3+ Treg frequency | — |
| metabolomics (mass spectrometry) and Seahorse extracellular flux analysis | splenic/in vitro-activated CD4+ T cells, WT vs Arg1 CKO mice | Arg1 knockout, in vitro activation | ornithine, polyamine, arginine, glutamine levels; oxygen consumption rate (OXPHOS), extracellular acidification rate (glycolysis), glucose uptake, mTOR activation | mass spectrometry; Seahorse |
- ▲ Arg1 was among the most highly induced genes in influenza-activated lung CD4+ T cells versus splenic CD4+ T cells, ranking 3rd among induced enzymes and 11th overall top 1.5 percentile
- ▼ Arg1 CKO mice showed reduced lung inflammation and pathology score at day 9 post-infection compared with WT, with similar viral clearance
- ▲ At day 7 p.i., Arg1 CKO mice had increased numbers/frequency of virus-induced (CD11a+CD49d+) and dividing (Ki67+) lung CD4+ T cells versus WT
- ▼ At day 9 p.i., Arg1 CKO mice showed reduced frequencies and numbers of virus-induced and dividing lung CD4+ T cells versus WT, indicating faster contraction
- ▲ Arg1 CKO CD4+ T cells showed higher frequencies/numbers of IFN-γ, IL-2, TNF-α, and IL-10-producing cells at day 7 p.i. after ex vivo restimulation
- ▼ Arg1 CKO CD4+ T cells failed to reach WT levels of basal and maximal oxygen consumption rate and spare respiratory capacity upon in vitro activation
- ▼ Glutamine was among the most differentially abundant metabolites, strongly reduced in activated Arg1 CKO CD4+ T cells versus WT despite normal glutamine uptake
- – Ornithine, polyamine, and arginine levels were unaltered in Arg1 CKO CD4+ T cells despite loss of ARG1 enzymatic conversion
- count 656 genes increased and 15 genes decreased in expression (differential gene expression in lung CD4+ T cells vs splenic CD4+ T cells after influenza infection)
- other Arg1 ranked 3rd among all induced enzymes and 11th most highly induced gene overall (top 1.5 percentile) (RNA-seq ranking of induced genes in influenza-activated lung CD4+ T cells)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study uses an in vivo mouse influenza (PR8/H1N1) infection model together with in vitro CD4+ T cell cultures, adoptive-transfer (Rag KO) and transfer colitis models, and conditional/global knockouts (Arg1 CKO, Arg2 KO) to compare immune phenotypes against wild-type controls. High-throughput readouts (bulk RNA-seq with differential-expression and pathway-enrichment analysis, and untargeted/targeted mass-spectrometry metabolomics) are paired with flow-cytometric quantification of cell numbers, frequencies, cytokines, and proliferation, and with histopathology scoring. Group comparisons are described qualitatively in the narrative (e.g., 'higher,' 'reduced,' 'trend toward,' 'did not reach statistical significance'), but the excerpt provided does not contain the explicit statistics/methods section naming the specific tests, n, or software.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| differential gene-expression analysis for RNA-seq (specific method/test not stated in the provided text) | lung vs. splenic CD4+ T cells (Figs 1B-1F, Table S1) and WT vs. Arg1 CKO vs. Arg2 KO in vitro CD4+ T cells (Fig 4L, Table S2) | — | not stated |
| pathway/gene-set enrichment analysis (specific method not stated) | genes differentially expressed in Arg1 CKO cells, IFN-γ-associated pathways (Fig 4M) | — | not stated |
| group comparison of cell numbers/frequencies/cytokines (specific test not stated in provided text) | WT vs. Arg1 CKO lung T cell numbers, frequencies, Ki67+, cytokine+ cells at days 7/9 p.i. (Fig 2); in vitro proliferation/cytokines (Fig 4) | — | not stated |
| group comparison for metabolite abundances (specific test not stated) | WT vs. Arg1 CKO ornithine, polyamines, arginine, glutamine (Figs 5B-5D, 6A-6B, Table S3) | — | not stated |
| group comparison for respirometry/glycolysis (OCR/ECAR) (specific test not stated) | basal/maximal OCR, spare respiratory capacity, glycolysis WT vs. Arg1 CKO (Figs 5E-5G) | — | not stated |
| group comparison for histopathology/inflammation scores (specific test not stated) | lung pathology/inflammation (Figs 1J, 1K) and colitis severity (Figs 3I-3R) | — | not stated |
-
Genome-wide differential expression and metabolite comparisons were performed across thousands of features.↳ Could also: Reporting a defined multiplicity-control framework such as Benjamini-Hochberg FDR (for RNA-seq/metabolomics) alongside the per-test results. — An explicit FDR statement makes the family of comparisons and the chosen significance threshold transparent, which readers often find helpful for interpreting omics-scale results.
-
WT and knockout groups were compared across several related readouts (numbers, frequencies, Ki67+, multiple cytokines) at each time point.↳ Could also: A single ANOVA model (e.g., two-way for genotype × time) with a post-hoc adjustment such as Tukey HSD or Sidak, in place of independent pairwise comparisons. — A combined model can control the family-wise error rate across the related comparisons and directly test genotype-by-time interactions, which matches the paper's kinetic ('accelerated') narrative.
-
Several effects are described qualitatively as a 'trend' that 'did not reach statistical significance.'↳ Could also: Accompanying such statements with the exact p-value and an effect size with its 95% confidence interval. — Effect sizes with CIs convey the magnitude and precision of a difference independent of the significance threshold, which is informative for the small-n in vivo comparisons typical of this design.
-
Mouse experiments compared genotypes in an infection/phenotyping setting.↳ Could also: Explicitly reporting randomization of animals to groups and blinding of outcome scoring (e.g., for histopathology/colitis severity). — Documented randomization and blinded scoring are widely recommended for in vivo studies and help readers gauge how subjective endpoints like pathology scores were assessed.
-
Group spread and sample size are summarized in the narrative without the dispersion measure being specified in the provided text.↳ Could also: Stating the dispersion statistic (SD, SEM, or 95% CI) together with the exact n for each comparison, and for small n showing individual data points. — Reporting SD or a 95% CI (rather than SEM alone) and plotting individual values conveys the actual variability and is often preferred when group sizes are small.
-
Ordinal histopathology and colitis severity scores were compared between groups.↳ Could also: Analyzing ordinal scores with a rank-based test (e.g., Mann-Whitney U) or an ordinal-regression model. — Rank-based or ordinal methods do not assume interval-scaled, normally distributed scores, which aligns with the categorical nature of pathology scoring.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
ARG1-deficient CD4+ T cells show accelerated expansion at day 7 (increased numbers and Ki67+) followed by faster contraction by day 9 post influenza infection.flow-cytometry mouse lung mixed 2023×1papers★ This paper is the founder (earliest)
-
ARG1-deficient CD4+ T cells produce higher frequencies of IFN-γ, IL-2, TNF-α, and IL-10 upon ex vivo restimulation during influenza infection.flow-cytometry mouse lung up 2023×1papers★ This paper is the founder (earliest)
-
ARG1-deficient naive CD4+ T cells induce reduced colitis severity and lower inflammation scores in Rag1-deficient recipients compared to WT cells.imaging mouse colon down 2023×1papers★ This paper is the founder (earliest)
-
ARG1 deletion in CD4+ T cells reduces lung inflammation and pathology score at day 9 post influenza infection without impairing viral clearance.imaging mouse lung down 2023×1papers★ This paper is the founder (earliest)
-
Glutamine levels are specifically reduced in ARG1-deficient CD4+ T cells while ornithine, polyamines, and arginine remain unchanged.metabolomics mouse cd4-t-cell down 2023×1papers★ This paper is the founder (earliest)
-
ARG1-deficient CD4+ T cells have reduced basal OCR, maximal OCR, and spare respiratory capacity upon activation, indicating impaired oxidative phosphorylation.other mouse cd4-t-cell down 2023×1papers★ This paper is the founder (earliest)
-
ARG1-deficient CD4+ T cells show enrichment of IFN-γ response gene signatures relative to WT and ARG2-deficient cells in vitro.RNA-seq mouse cd4-t-cell up 2023×1papers★ This paper is the founder (earliest)
-
ARG1 is strongly upregulated in lung CD4+ T cells during influenza infection, ranking as the 11th most induced gene (top 1.5 percentile).RNA-seq mouse lung cd4-t-cell up 2023×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37572656
Paper: West et al., Immunity 2023. "Loss of CD4+ T cell-intrinsic arginase 1
accelerates Th1 response kinetics and reduces lung pathology during influenza infection."
PMID 37572656 · PMC10576612 · DOI 10.1016/j.immuni.2023.07.014
Repo: https://github.com/jackbibby1/E-West @ commit c7c10a11ac2f217f3f1470082fb25df7cbe2cfe0
Accession correction (IMPORTANT)
The room brief lists geo:GSE112244, but that is a different paper
(PMID 30392958, Th1/Th17 ± CB839 glutaminase-inhibitor RNA-seq). The accession for
THIS paper's transcriptomics, per its own Data Availability statement, is
GEO: GSE229775 — Illumina HumanHT-12 V4 microarray of CD4+ T cells from healthy
donors vs arginase-1-deficient (ARG1-def) patients, 6 h post anti-CD3 + anti-CD46
activation. The reproduction targets GSE229775 / the repo, not GSE112244.
In scope (pipeline-derived)
The only bioinformatic-pipeline result in the paper is the microarray GSEA shown in
Figure 7P (human patient arm). Pipeline = limma neqc normalization of Illumina
IDATs (script arg1_analysis.R) → Broad GSEA (standard two-class, REACTOME gene
sets) → pathway plot.
- Claim C1 (Fig 7P, text): "Ranked among the top pathways perturbed in the patients' CD4+ T cells were 'metabolism of amino acids and derivatives' and 'interleukin 10 signaling'." → REACTOME_METABOLISM_OF_AMINO_ACIDS_AND_DERIVATIVES and REACTOME_INTERLEUKIN_10_SIGNALING are top-ranked / significant pathways enriched in the ARG1-deficient patients vs healthy donors.
The repo ships the exact GSEA inputs: stim_gsea.txt (17,053 genes × 4 samples:
healthy_cd3_46 ×2, arg1_def_cd3_46 ×2) and labs.cls (2 classes: healthy, arg1).
We reproduce by running Broad GSEA on these shipped inputs with REACTOME gene sets.
Out of scope (not attempted)
- All wet-lab / mouse in-vivo influenza work (most of the paper — flow, viral titers, histology, metabolomics, Seahorse, in-vivo Arg1-cKO phenotypes). Not pipeline-derived.
- Re-deriving
normalised_data.csvfrom raw IDATs: the raw.idat+.bgxfiles are NOT in the repo (only on GEO GSE229775); the script'sread.idat/neqcstep is not runnable from the repo alone. The repo ships the post-normalization GSEA input, which is exactly what feeds Fig 7P, so we reproduce from there (the authors' own provided data — equally valid per brief P16). - Exact NES/FDR numeric match: the paper text reports no NES/FDR/p values for Fig 7P (only the named pathways), and the authors' precise GSEA version/parameters/MSigDB release are not recorded. We grade the qualitative ranking claim, not exact numbers.
Method note / known deviations
- GSEA tool: Broad
GSEA_Linux_4.3.3(the tool the authors used; their report files are namedgsea_report_for_{healthy,arg1}_*.tsv, i.e. standard two-class GSEA). - Permutation:
gene_set(n=2/class is below the ≥7/class needed for phenotype permutation) — recommended Broad practice for small n; affects FDR stability, not the ranking claim. - Gene sets: MSigDB C2 CP:REACTOME (human symbols); exact MSigDB release unspecified by authors — we record the version actually used.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Reproducing the only pipeline-derived result (Fig 7P human microarray GSEA) from the authors' own shipped inputs confirms IL-10 signaling as the #1 of 771 enriched pathways (NES +2.684, FDR 0.0) — an unambiguous hit. The second named set, amino-acid metabolism, is significant and directionally consistent (NES -1.562, FDR 0.117) but ranks 83/516, so the 'among the top' framing holds only partly. The deviation is on our methodology side (unpinned GSEA metric/MSigDB version, an authors' documentation gap forcing self-chosen parameters), not a derivability or fabrication problem; the qualitative claim is genuinely supported. Overall a solid, explainable yellow reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.