Competitive binding of STATs to receptor phospho-Tyr motifs accounts for altered cytokine responses.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce; mostly 1:1. The repo (PollyJeffrey/Cytokine_modelling, pinned commit 9c3e0dd) is a self-contained Python ABC-SMC pipeline; GEO GSE164479 is mass-spec data and is NOT the modelling input (the pSTAT time courses are shipped in-repo). Ran on «our HPC» (SLURM 2176945). RESULTS: (1) Table 1's 16 reported posterior mean+median rate constants are reproduced EXACTLY from the shipped RPE1_posteriors.txt (no fabrication signal; fully derivable). (2) Re-running the authors' ODE model verbatim on 400 posterior draws reproduces the Fig 2c model fit (median dist 0.42 to data) and confirms every shipped draw has distance <= 0.6 = the stated final threshold, i.e. the posteriors file genuinely is the accepted set. (3) Qualitative claims (k3a+>k3b+, >=1 order-of-magnitude differences) hold. (4) Model-selection direction reproduced (H1 strongly favoured, rising as delta falls). PARTIAL on the single headline number: the stated 99% preference for H1 comes out as 92.55% from the shipped final-iteration file -- same conclusion, but the exact 99% is not derivable from the repo's shipped artifact (likely a different ABC-SMC seed/run or the limiting probability). NOT ATTEMPTED (optional hard ~20%): re-running the full ABC-SMC from scratch (N=10^4 x 15 iterations, ~10^6-10^7 ODE solves, stochastic); Th-1 cell inference; and all wet-lab/biophysics results (SPR, smFRET, mass-spec GSE164479), which are out of scope.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 89assessed: 2026-06-15 ⛓ c6f302fe905f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusHow do IL-6 and IL-27, two cytokines that activate the same JAK1/STAT1/STAT3 signaling pathway and share the GP130 receptor subunit, elicit non-redundant biological responses? The paper tests whether differential competitive binding of STATs to receptor phospho-tyrosine motifs determines distinct signaling kinetics and gene programs.
- ★ IL-27 induces more sustained and higher-amplitude STAT1 phosphorylation than HypIL-6, while both induce comparable STAT3 phosphorylation finding
- ★ Differential binding of STAT1 to IL-27Rα and STAT3 to GP130 (competitive binding to receptor phospho-Tyr motifs) is the main dynamical process driving sustained pSTAT1 by IL-27 mechanism
- ★ Tyr613 on IL-27Rα is required for IL-27-induced STAT1 phosphorylation but not STAT3 phosphorylation mechanism
- ★ Sustained STAT1 phosphorylation and IRF1 expression drive a unique IL-27 gene program enriched in classical Interferon Stimulated Genes that shapes the T-cell proteome finding
- ★ Receptor and STAT concentrations critically shape cytokine responses and generate functional pleiotropy; predicted by mathematical/statistical modeling mechanism
- ★ SLE patients show higher STAT1 expression and biased/more potent STAT1 activation by IL-6/IL-27 than healthy controls finding
- ★ Sub-saturating doses of the JAK inhibitor Tofacitinib specifically lower STAT1 activation by IL-6, offering a strategy to selectively target individual STATs finding
- Single-molecule mathematical modeling of IL-6 and IL-27 STAT signaling kinetics as a framework to identify molecular determinants of functional selectivity method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| High-throughput multiparameter flow cytometry (phospho-flow, fluorescent barcoding) | Human Th-1 cells (CD4+ from buffy coat), activated PBMCs (CD4+, CD8+) | IL-27 (mIL-27sc) and HypIL-6 stimulation (dose/response and kinetics) | pSTAT1 and pSTAT3 levels (MFI) | — |
| Dual-color single-molecule TIRF imaging with co-localization/co-tracking | RPE1 cells (GP130 KO reconstituted with mXFPe-IL-27Rα and mXFPm-GP130) | IL-27 / HypIL-6 stimulation | Receptor heterodimerization (IL-27Rα/GP130) and GP130 homodimerization via co-trajectories | Dye-conjugated anti-GFP nanobodies (RHO11, DY649) |
| Single-molecule bleaching analysis and single-molecule FRET | RPE1 cells expressing mXFPe-IL-27Rα and GP130 | IL-27 stimulation | Receptor complex stoichiometry (1:1) and molecular proximity | — |
| Flow cytometry receptor surface quantification | RPE1 GP130 KO, wt RPE1, RPE1 GP130KO + mXFPm-GP130 (10x GP130) | GP130 overexpression; HypIL-6 stimulation | Cell surface GP130 levels and pSTAT1/pSTAT3 kinetics | — |
| Phospho-proteomics | Human Th-1 cells | Early IL-27 / HypIL-6 stimulation | Phosphorylation events | — |
| Transcriptomics (gene expression kinetics) | Human Th-1 cells | IL-27 / HypIL-6 stimulation over time | Kinetics of transcriptomic changes / gene expression program | — |
| Proteomics | Human T-cells | Prolonged IL-27 / HypIL-6 exposure | T-cell proteome alterations | — |
| Receptor mutagenesis with phospho-flow readout | Cells expressing IL-27Rα Tyr613 mutant | Mutation of Tyr613 on IL-27Rα; IL-27 stimulation | IL-27-induced pSTAT1 and pSTAT3 levels | — |
- ▼ Mutation of Tyr613 on IL-27Rα decreased IL-27-induced STAT1 phosphorylation with limited effect on STAT3 phosphorylation 80% decrease
- ▲ IL-27/HypIL-6 induced more potent STAT1 activation in SLE patients than healthy controls, correlating with higher STAT1 expression
- – STAT1/3 phosphorylation more sensitive to IL-27 (EC50 ~20 pM) than HypIL-6 (EC50 ~400 pM) EC50 ~20 pM vs ~400 pM
- – Both cytokines yielded same maximal pSTAT3 amplitude, but HypIL-6 gave significantly reduced maximal pSTAT1 amplitude relative to IL-27
- – Both cytokines exhibited nearly identical sustained pSTAT3 profile with ~20% activation remaining after 3 hr continuous stimulation ~20% after 3 hr
- – 10x GP130 overexpression in RPE1 made STAT3 activation more sustained with HypIL-6 but had very little effect on STAT1 kinetics 10x GP130
- ▲ After IL-27 stimulation, IL-27Rα and GP130 form substantial heterodimers with 1:1 stoichiometry; no pre-assembly in resting state 1:1
- – Purchased IL-27 EC50 pSTAT1 ~70 pM, pSTAT3 ~80 pM; mIL-27sc EC50 pSTAT1 ~20 pM, pSTAT3 ~30 pM in activated CD4+ cells ~70/80 pM vs ~20/30 pM
- fold_change 80% decrease in IL-27-induced STAT1 phosphorylation upon Tyr613 mutation (IL-27Rα Tyr613 mutation effect on pSTAT1)
- other EC50 ~20 pM (IL-27) vs ~400 pM (HypIL-6) (STAT1/3 phosphorylation dose-response in Th-1 cells)
- other ~20% pSTAT3 activation remaining after 3 hr (Sustained STAT3 kinetics for both cytokines)
- other 10x higher GP130 surface levels (RPE1 clone overexpressing GP130)
- count minimum of 23 cells measured per condition (Single-molecule co-trajectory analysis)
- other EC50 purchased IL-27 pSTAT1 ~70 pM, pSTAT3 ~80 pM; mIL-27sc pSTAT1 ~20 pM, pSTAT3 ~30 pM (Activated CD4+ cells dose-response)
- count five biological replicates with two technical replicates each (STAT1/3 phosphorylation kinetics in Th-1 cells)
- count three biological replicates with two technical replicates each (Dose-dependent STAT phosphorylation in Th-1 cells)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combines quantitative cell-biology and -omics measurements with mathematical/statistical modeling of STAT1/STAT3 signaling kinetics. Group comparisons of single-molecule co-tracking data were assessed with two-tailed Student's t-tests, and dose-response/kinetic measurements were summarized as mean ± standard deviation from biological replicates each containing technical replicates. Significance was reported using thresholded p-value tiers (e.g., *p<0.05, **p≤0.01, ***p≤0.001) rather than exact values for the displayed comparisons.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-tailed Student's t-test | relative number of co-trajectories for IL-27Rα/GP130 heterodimerization and GP130 homodimerization across stimulation conditions (Figure 1g) | each data point = one cell; minimum 23 cells per condition | not stated |
| two-tailed Student's t-test | diffusion coefficients of receptors with/without cytokine stimulation (Figure 1—figure supplement 3c) | each data point = one cell; minimum 23 cells per condition | not stated |
-
Pairwise group comparisons of single-molecule co-tracking and diffusion data were made with two-tailed Student's t-tests.↳ Could also: A non-parametric test such as Mann-Whitney U, or a permutation/bootstrap test, could also be applied. — Per-cell distributions can be skewed or have unequal variance; a rank-based or resampling approach makes fewer distributional assumptions and is often chosen when normality is not formally checked.
-
Several conditions (cytokines, stimulation states) were compared using multiple two-tailed t-tests.↳ Could also: A single one-way (or two-way) ANOVA followed by a post-hoc test such as Tukey HSD could also be used. — An omnibus model with post-hoc correction would also control the family-wise error rate across the set of related comparisons in one analysis.
-
Significance was reported as thresholded tiers (*, **, ***).↳ Could also: Exact p-values alongside an effect-size estimate (e.g., difference in means with a 95% confidence interval) could also be reported. — Exact values and effect sizes convey both the strength and magnitude of a difference, which complements categorical significance markers.
-
Spread of replicate measurements was summarized with standard deviation.↳ Could also: A 95% confidence interval (or showing all individual data points) could also be displayed. — A confidence interval conveys precision of the estimate and is often preferred for small replicate numbers, while plotting individual points shows the full data distribution.
-
Sample sizes were stated descriptively per panel without a formal power analysis.↳ Could also: An a priori power/sample-size justification could also accompany the design. — A stated power calculation makes explicit the sensitivity of each comparison and aids interpretation and reproducibility.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Engineered mIL-27sc is ~3-4-fold more potent than commercial IL-27 for STAT1/STAT3 activation (EC50 ~20-30 pM vs ~70-80 pM) in activated human CD4+ cells.flow-cytometry human-cd4 2021×1papers★ This paper is the founder (earliest)
-
IL-27- and HypIL-6-induced STAT1 phosphorylation is more potent in SLE patient PBMCs than healthy donor PBMCs, correlating with elevated basal STAT1 protein expression.flow-cytometry human-pbmc-sle up 2021×1papers★ This paper is the founder (earliest)
-
HypIL-6 produces significantly lower maximal STAT1 phosphorylation amplitude than IL-27 despite eliciting identical maximal STAT3 activation in human Th-1 cells.flow-cytometry human-th1 down 2021×1papers★ This paper is the founder (earliest)
-
IL-27 activates STAT1 and STAT3 with ~20-fold higher potency (EC50 ~20 pM) than HypIL-6 (EC50 ~400 pM) in human Th-1 cells.flow-cytometry human-th1 2021×1papers★ This paper is the founder (earliest)
-
IL-27 and HypIL-6 produce identical sustained STAT3 phosphorylation profiles (~20% activation remaining at 3 hr continuous stimulation) in human Th-1 cells.flow-cytometry human-th1 none 2021×1papers★ This paper is the founder (earliest)
-
Mutation of IL-27Rα Tyr613 reduces IL-27-induced STAT1 phosphorylation by ~80% with minimal effect on STAT3 phosphorylation, identifying Tyr613 as the principal STAT1-recruiting phospho-Tyr motif.flow-cytometry rpe1 down 2021×1papers★ This paper is the founder (earliest)
-
10-fold GP130 overexpression prolongs HypIL-6-induced STAT3 activation without altering STAT1 phosphorylation kinetics in RPE1 cells.flow-cytometry rpe1 up 2021×1papers★ This paper is the founder (earliest)
-
IL-27 stimulation induces ligand-dependent 1:1 IL-27Rα:GP130 receptor heterodimerization in RPE1 cells; no pre-formed dimers exist in the resting state.imaging rpe1 up 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- Topological signatures in regulatory network e... L1 94/100
- Generic injuries are sufficient to induce ecto... L1 100/100
- Co-regulation and function of <i>FOXM1</i>/<i>... L1 76/100
- Tbx5 drives <i>Aldh1a2</i> expression to regul... L1 84/100
- Disrupted PGR-B and ESR1 signaling underlies d... L1 78/100
- Firefly genomes illuminate parallel origins of... L1 93/100
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33871355
Paper: Wilmes et al. (2021) eLife. "Competitive binding of STATs to receptor phospho-Tyr motifs accounts for altered cytokine responses." DOI 10.7554/eLife.66014.
Code: https://github.com/PollyJeffrey/Cytokine_modelling
(redirects to PollyJeffrey/Cytokine-modelling-paper), Python, MIT, pinned commit
9c3e0ddc7a96eac941baad560d1541d660b0515d (the exact SHA cited in the paper's Data
Availability via Software Heritage swh:1:rev:9c3e0dd…). Zenodo DOI 10.5281/zenodo.4609852.
Data: GEO GSE164479 = phospho-proteomic / proteomic (mass-spec) datasets.
NOTE: the GEO accession is not the input to the modelling pipeline. The model is
calibrated against pSTAT1/pSTAT3 phospho-flow time courses, which are shipped
inside the repo as 8 small normalised .txt files (Eq. 5 normalisation). So the
in-scope computational pipeline is fully self-contained in the repo.
In scope (pipeline-derived, attempted)
The repo is an ABC-SMC (Approximate Bayesian Computation – Sequential Monte
Carlo) model-selection + Bayesian parameter-inference pipeline (numpy+scipy.odeint).
Three scripts: ABC_SMC_model_selection.py, ABC_SMC_RPE1.py, ABC_SMC_TH1.py.
Shipped pipeline outputs in the repo: Accepted_models_iteration_0..14.txt,
RPE1_posteriors.txt, TH1_posteriors.txt.
| Result | Paper location | Pipeline | In scope |
|---|---|---|---|
| Model selection: hypothesis 1 favoured over hypothesis 2 with probability 99 % (RPE1) | Results §"…model selection…"; Figure 2b | ABC-SMC model selection → Accepted_models_iteration_14.txt |
YES |
| Fig 2b trajectory: rel. prob. of H1 rises as distance threshold δ falls | Figure 2b | same, all 15 iteration files | YES |
| Table 1: posterior mean & median of 8 STAT binding/dissociation rates k1a±,k1b±,k3a±,k3b± (RPE1) | Table 1 | ABC-SMC inference → RPE1_posteriors.txt |
YES |
| Qualitative ordering k3a+ > k3b+ ; "≥1 order-of-magnitude" rate differences | Results, Fig 2d | RPE1_posteriors.txt |
YES |
| Fig 2c: pointwise median + 95% CI model fit to pSTAT1/3 vs time, calibrated by posterior | Figure 2c | forward ODE sim of authors' model with posterior draws | YES (deterministic forward sim) |
Out of scope (not attempted, why)
- Re-running the full ABC-SMC from scratch (N=10⁴ particles × 15 iterations, acceptance rate falls steeply at small δ → ~10⁶–10⁷ ODE solves). Stochastic; the hard last ~20%. We instead (a) verify the shipped pipeline outputs reproduce the reported numbers and (b) re-run the authors' ODE model forward. We do NOT claim a from-scratch re-inference.
- Wet-lab / biophysics: surface plasmon resonance, single-molecule FRET, mass-spec phospho-proteomics (GSE164479), flow cytometry, microscopy — all experimental, not pipeline-derived.
- Th-1 cell inference (
TH1_posteriors.txt): same machinery as RPE1; we focus the gradeable comparison on RPE1 (where the paper states the 99 % headline and Table 1). Th-1 left as optional.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a mathematical-modelling paper whose ABC-SMC pipeline and pSTAT time-course inputs are shipped in-repo (GEO GSE164479 is unrelated mass-spec data). The reproduction is essentially 1:1: all 16 Table-1 rate constants match exactly from RPE1_posteriors.txt, the Fig-2c ODE fit reproduces (median dist 0.42; all 400 draws ≤0.6=δ_final), and the qualitative claims (H1 favoured, k3a+>k3b+, ≥1 order of magnitude) hold. The only deviation is the headline 99% → 92.55% P(H1), which is explainable by ABC-SMC stochasticity/seed (a technical, not authors', issue) and does not change any conclusion. Overall yellow: solid reproduction with one explainable, beyond-rounding headline discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.