Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Defactinib inhibits PYK2 phosphorylation of IRF5 and reduces intestinal inflammation.

Nat Commun · 2021
L1 67/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
67/100
Reproducibility score
0.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 29% of all assessed papers rank 795 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL 1:1 reproduction of all 5 Figure-5 pipeline results from the deposited GSE141082 featureCounts matrix (DESeq2 1.42.0 -> gsfisher 0.2, the paper's named tool). STRONG: C4 (gsfisher GO) recovers all three Fig5c-named terms with p.adj<1e-7 (cellular response to interferon-beta, cytokine activity, regulation of inflammatory response), and C3 (PCA) reproduces genotype+treatment separation. C5 recovers 6/8 named inflammatory genes. APPROXIMATE: DEG counts are the right order of magnitude and robustly confirm the headline biological conclusion (defactinib's transcriptional effect is largely IRF5-dependent: WT response ~10-50x larger than KO) but are not exact -- C1 is robustly ~15% high (4567-4776 vs 4026) and C2 is highly design-sensitive with the reported 217 BRACKETED by our pairwise(96) and full-model(518) runs. The paper under-specifies the DESeq2 design formula/version/pre-filtering, which fully accounts for the count differences; no sign of fabrication -- all reported values are derivable from the shipped matrix within method ambiguity. NOT ATTEMPTED (out of scope, wet-lab): phospho-proteomics, Western blots, kinase assays, mouse colitis models, IHC, ELISAs.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 67
    assessed: 2026-06-21 ⛓ d36b4da01598
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The molecular mechanisms activating the transcription factor IRF5 (a key driver of gut inflammation) are unclear, so the authors set out to identify novel upstream kinases—specifically testing whether PTK2B/PYK2 phosphorylates and activates IRF5 to drive macrophage inflammatory responses relevant to inflammatory bowel disease.

Core claims
  • PYK2 was identified as a putative IRF5 kinase via a kinase inhibitor library screen in macrophages finding
  • PYK2 binds IRF5 through IRF5's C-terminal serine-rich region (SRR) and PYK2's kinase domain mechanism
  • PYK2-deficient macrophages show impaired endogenous IRF5 activation and reduced inflammatory gene expression finding
  • PYK2 phosphorylates IRF5 at tyrosine 172 (human), contributing to LPS-induced IRF5 activation mechanism
  • The PYK2 inhibitor defactinib suppresses IRF5 activation and induces a transcriptomic signature similar to IRF5 deficiency finding
  • Defactinib prevents LPS-induced nuclear translocation of IRF5 but not of p65/RelA, indicating IRF5-specific, NF-κB-independent action mechanism
  • Defactinib reduces pro-inflammatory cytokine expression in human ulcerative colitis colon biopsies and in a mouse colitis model finding
  • Defactinib is a selective PYK2/FAK1 inhibitor already used in cancer models and clinical trials resource
Experimental setups
Assay System Perturbation Readout Platform
Kinase inhibitor library screen (PKIS, 221 kinases) with TNF-promoter luciferase reporter RAW264.7 macrophages; 293 TLR4/CD14/MD-2 cells small molecule kinase inhibitors IRF5-dependent TNF-luciferase reporter activity PKIS library
Co-immunoprecipitation and Western blot 293 ET cells (overexpression) overexpression of Myc-tagged kinases and HA-tagged IRF5 IRF5-kinase binding
In vitro kinase assay 293 ET cells co-transfection of HA-IRF5 with myc/FLAG-tagged kinases IRF5 phosphorylation
Co-IP domain mapping with truncation mutants HEKTLR4 cells IRF5 and PYK2 truncation mutants interaction domains required for IRF5-PYK2 binding
Endogenous co-IP and phospho-Western blot RAW264.7 macrophages LPS stimulation endogenous IRF5-PYK2 binding; PYK2 Y402 phosphorylation
CRISPR-Cas9 knockout, luciferase reporter, ChIP, qPCR RAW264.7 macrophages (PYK2 KO, IRF5 KO) CRISPR knockout + LPS TNF-luc reporter activity; IRF5/Pol II promoter recruitment (Il6, Il1a, Tnf); cytokine/chemokine mRNA (Il6, Il1a, Ccl4, Ccl5, Il10)
qPCR gene expression HoxB8-derived macrophages (PYK2 KO, IRF5 KO) CRISPR knockout + LPS inflammatory cytokine/chemokine mRNA levels
Phospho-proteomics (nUPLC-MS/MS) RAW264.7 macrophages (WT and PYK2 KO) LPS stimulation, PYK2 knockout IRF5 phosphorylation sites nano ultra-high-pressure liquid chromatography-MS/MS
In vitro kinase assay with site-specific IRF5 tyrosine mutants HEKTLR4 cells IRF5 Y104F/Y172F/Y312F/Y334F mutants + PYK2 co-transfection IRF5 phosphorylation and reporter activity
Luciferase reporter, cell fractionation, ChIP, qPCR, Western blot RAW264.7 macrophages and primary mouse BMDMs defactinib (PYK2/FAK inhibitor) ± LPS reporter activity, IRF5 nuclear translocation, IRF5/Pol II promoter recruitment, cytokine expression
Ex vivo cytokine expression assay human colon biopsies from ulcerative colitis patients defactinib treatment pro-inflammatory cytokine expression
In vivo colitis pathology and cytokine assessment mouse colitis model defactinib treatment colitis pathology and cytokine expression
Key results
  • LPS-induced TNF-reporter activity markedly reduced in IRF5-expressing PYK2 KO RAW264.7 cells
  • IRF5 and RNA Pol II recruitment to Il6, Il1a, Tnf promoters impaired in PYK2 KO cells
  • IRF5 Y171 (murine)/Y172 (human) phosphorylation detected only in WT cells, absent in PYK2-deficient cells
  • Defactinib inhibited TNF-reporter activity in WT and IRF5-restored cells but had no additional effect in PYK2 KO cells
  • Defactinib blocked LPS-induced nuclear translocation of IRF5 but not p65/RelA
  • Defactinib reduced pro-inflammatory cytokine expression in UC patient colon biopsies and in mouse colitis model
  • Cytokine/chemokine mRNA induction reduced in PYK2 KO and IRF5 KO HoxB8 macrophages, comparable to or stronger than IRF5 KO
  • LPS-induced IL-10 expression increased in PYK2 knockout cells
Key statistics
  • pvalue P < 0.05, ** P < 0.01, *** P < 0.001, **** P < 0.0001 (two-way ANOVA with Sidak's correction) (Il6 and Il1a mRNA/ChIP in WT vs PYK2 KO vs IRF5 KO RAW264.7 cells)
  • pvalue P < 0.05, ** P < 0.01, *** P < 0.001, **** P < 0.0001 (two-way ANOVA) (Gene expression in WT vs PYK2 KO vs IRF5 KO HoxB8 macrophages)
  • pvalue * P < 0.05, ** P < 0.01, *** P < 0.001 (two-way ANOVA with Tukey's correction) (TNF-luc reporter activity with defactinib in WT/IRF5 KO/PYK2 KO RAW264.7 cells)
  • pvalue * P < 0.05, ** P < 0.01 (one-way ANOVA with Tukey's correction) (IRF5/Pol II ChIP binding to Il6 and Il1b promoters in defactinib-treated BMDMs)
  • pvalue * P < 0.05, *** P < 0.001 (one-way ANOVA with Tukey's correction) (Il6 and Il1b expression in defactinib-treated GM-BMDMs from individual mice)
  • count 34 candidate IRF5 kinases shortlisted from 221 screened (Kinase inhibitor library screening workflow)
  • other defactinib effective concentration range 0.3-1 μM (RAW264.7) and 3.5 μM (BMDMs) (Dose-dependent inhibition of IRF5-PYK2 pathway)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper combines a kinase inhibitor library screen with molecular and cellular validation experiments in macrophage cell lines (RAW264.7, HoxB8) and primary mouse bone marrow-derived macrophages (BMDMs). Group comparisons across multiple genotypes (WT, PYK2 KO, IRF5 KO) and treatment conditions (LPS ± defactinib) were analysed by one-way or two-way ANOVA with post-hoc corrections (Sidak or Tukey). Results throughout are expressed as mean ± SEM from n = 3 independent experiments (cell lines) or n = 4 individual mice (BMDMs), with significance indicated by asterisk thresholds rather than exact p-values.

Replicationbiological Sample sizen = 3 independent experiments for cell line assays; n = 4 individual mice for primary BMDM assays; no formal power calculation reported GroupsWT vs. PYK2 KO vs. IRF5 KO macrophages, with or without LPS stimulation and/or defactinib treatment Pairingunpaired Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionSidak's correction (Fig. 2c–e) and Tukey's HSD (Figs. 4a, 4c, 4d) as post-hoc tests within individual ANOVAs
Statistical tests used
Test Applied to n Assumptions
Two-way ANOVA with Sidak's post-hoc correction TNF-luc reporter activity and qPCR mRNA induction (Il6, Il1a, Ccl4, Ccl5) in WT, PYK2 KO, IRF5 KO RAW264.7 macrophages ± LPS (Figs. 2c–e) n = 3 independent experiments not stated
Two-way ANOVA Gene expression (qPCR) in WT, PYK2 KO, IRF5 KO HoxB8 macrophages ± LPS (Fig. 2f) n = 3 experiments not stated
Two-way ANOVA with Tukey's post-hoc correction TNF-luc reporter activity in WT, IRF5 KO, PYK2 KO RAW264.7 cells ± defactinib ± LPS (Fig. 4a) n = 3 independent experiments not stated
One-way ANOVA with Tukey's post-hoc correction IRF5 and RNA Pol II ChIP at Il6 and Il1b promoters in BMDMs ± defactinib ± LPS (Fig. 4c) n = 3 independent experiments not stated
One-way ANOVA with Tukey's post-hoc correction Il6 and Il1b expression (qPCR) in GM-BMDMs ± defactinib ± LPS (Fig. 4d) n = 4 individual mice not stated
Approaches that could also have been used
  • Dispersion is reported as SEM throughout, with n = 3–4 per group
    Could also: SD or 95% confidence intervals could also be used to summarise spread — With small n (3–4), SEM can make variability appear smaller than it is; SD directly reflects the spread of the observed data, and 95% CIs convey both precision and effect magnitude, which can aid interpretation and comparison across studies
  • Significance is indicated only by asterisk threshold bands (* < 0.05 through **** < 0.0001)
    Could also: Exact p-values (e.g., p = 0.0032) could also be reported for each comparison — Exact p-values allow readers to judge effect evidence on a continuous scale rather than discrete bins, support meta-analyses, and are increasingly recommended by journals and reporting guidelines (e.g., APA, Nature guidelines)
  • The two-way ANOVAs in Figs. 2c–e use Sidak's correction, while those in Fig. 4a and the one-way ANOVAs in Figs. 4c–d use Tukey's HSD
    Could also: A consistent single post-hoc strategy (e.g., Tukey HSD throughout, or Dunnett's test when all groups are compared to a single control) could also be applied uniformly — Using a single post-hoc method throughout simplifies interpretation; Dunnett's test would be slightly more powerful than Tukey when comparisons are only made against one control group (e.g., WT + LPS), which matches several figures here
  • Multiple separate ANOVAs are performed across figures without a study-wide multiplicity adjustment
    Could also: A global false discovery rate (FDR) correction (e.g., Benjamini–Hochberg) across all reported tests could also be applied — When many hypothesis tests are conducted across a paper, the experiment-wise type I error accumulates; an FDR procedure applied across the full test family would quantify and control that accumulation, which is standard practice in high-throughput contexts and increasingly adopted in focused mechanistic studies as well
  • Western blot and co-immunoprecipitation experiments are described qualitatively from representative blots (n = 3) without quantitative statistical analysis
    Could also: Densitometric quantification of band intensities followed by a statistical test (e.g., one-way ANOVA or paired t-test across n = 3 replicates) could also be reported — Quantitative band analysis with a formal test makes the conclusions from Western blot experiments more reproducible and allows readers to assess effect sizes and variability, complementing the representative image
  • No sample size justification or power calculation is reported for any experiment
    Could also: A brief power or sample size rationale (e.g., based on pilot data or published effect sizes for similar LPS-stimulation assays) could also be included — Reporting the basis for choosing n = 3 or n = 4 allows readers to assess whether the study was adequately powered to detect biologically meaningful differences, and is encouraged by reproducibility initiatives such as ARRIVE guidelines for in vivo work
Software: Not stated for statistical analyses · nUPLC-MS/MS (nano ultra-high-pressure liquid chromatography coupled mass spectrometry, for phosphoproteomics)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34795257

Paper: Ryzhakov et al. 2021, Defactinib inhibits PYK2 phosphorylation of IRF5 and reduces intestinal inflammation. Nat Commun 12:6702. DOI 10.1038/s41467-021-27038-5.

Tool (code link): https://github.com/sansomlab/gsfisher — third-party R package for Fisher's-exact gene-set/GO over-enrichment testing (P16: applying an existing tool to the paper's own data is equally valid).

Data: GEO GSE141082 — bulk RNA-seq, Mus musculus BMDMs, Illumina NovaSeq 6000. 18 samples = 6 conditions × 3 biological replicates:

  • WT + DMSO 0h, WT + DMSO 2h, WT + Defactinib 2h
  • IRF5ko + DMSO 0h, IRF5ko + DMSO 2h, IRF5ko + Defactinib 2h
  • (Defactinib 3.5 µM pretreat 1h; LPS stimulation 0 or 2h)
  • Processed count matrix deposited: GSE141082_featureCounts.txt.gz (4.1 MB). Raw fastq via SRA SRP233379.

In scope (pipeline-derived; reproducible from deposited count matrix)

The relevant computational results are all in Figure 5 (the only sequencing figure; rest of paper is wet-lab: Western blots, mass-spec phospho-IRF5, mouse colitis models, IHC — OUT of scope).

Pipeline as described in Methods: STAR (mm10) → featureCounts → DESeq2 (DEGs: |FC|>2 & FDR<0.05) → gsfisher (one-sided Fisher's exact GO enrichment). Since the featureCounts matrix is deposited, we start the pipeline from DESeq2 (STAR/featureCounts upstream not re-run; their output IS the deposited matrix).

ID Result Reported value Source
C1 DEGs WT defactinib-vs-DMSO @2h LPS 4,026 DEGs Fig 5b / text
C2 DEGs IRF5ko defactinib-vs-DMSO @2h LPS 217 DEGs Fig 5b / text
C3 PCA separates genotype & treatment qualitative separation Fig 5a
C4 GO enrichment of defactinib-downregulated genes "cellular response to interferon-beta", "regulation of inflammatory response", "cytokine activity" among top Fig 5c
C5 Overlap (IRF5-up ∩ defactinib-down) inflammatory genes Il6, Il1a, Il1b, Il12a, Il12b, Il23a, Ccl3, Ccl4 Fig 5d/e

Out of scope (wet-lab / not pipeline-derived)

Phospho-proteomics (mass-spec IRF5 phosphosites), Western blots, in-vitro kinase assays, mouse DSS/oxazolone colitis, histology/IHC, cytokine ELISAs — manual/experimental, not reproducible from deposited sequencing data.

Quick-minimum (~80% floor) vs stretch

  • Floor: C1, C2 (the two headline DEG counts — directly checkable from the matrix).
  • Stretch: C4 (gsfisher GO enrichment — the actual named tool), C5 (overlap gene membership), C3 (PCA).
Figures / tables: Fig 5bFig 5aFig 5cFig 5d
C1
Reported
4026 DEGs WT defactinib-vs-DMSO @2h LPS (Fig5b)
Reproduced
4654 full-model / 4664 pairwise (range 4567-4776 across designs)
partial
C2
Reported
217 DEGs IRF5ko defactinib-vs-DMSO @2h LPS (Fig5b)
Reproduced
518 full-model / 96 pairwise (reported 217 bracketed by 96-518)
partial
C3
Reported
PCA separates genotype & treatment (Fig5a)
Reproduced
PC1=49% PC2=33%; clear genotype + treatment/time separation
within tolerance
C4
Reported
GO: IFN-beta response, inflammatory response, cytokine activity (Fig5c)
Reproduced
all 3 significant (p.adj 1.4e-12 / 8.3e-8 / 1.2e-11) via gsfisher
exact
C5
Reported
Overlap inflammatory genes Il6/Il1a/Il1b/Il12a/Il12b/Il23a/Ccl3/Ccl4 (Fig5d/e)
Reproduced
6/8 present (Il1a,Il1b,Il6,Il12b,Ccl3,Ccl4)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 67/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

297.2 k
tokens (I/O) · 23.7 M incl. cache
94 min
runtime · 0.03 CPU-h
2.5 GB
peak RAM
2
HPC jobs
hummel
machine