Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Competitive binding of STATs to receptor phospho-Tyr motifs accounts for altered cytokine responses.

Elife · 2021
L1 89/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
89/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 77% of all assessed papers rank 247 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce; mostly 1:1. The repo (PollyJeffrey/Cytokine_modelling, pinned commit 9c3e0dd) is a self-contained Python ABC-SMC pipeline; GEO GSE164479 is mass-spec data and is NOT the modelling input (the pSTAT time courses are shipped in-repo). Ran on «our HPC» (SLURM 2176945). RESULTS: (1) Table 1's 16 reported posterior mean+median rate constants are reproduced EXACTLY from the shipped RPE1_posteriors.txt (no fabrication signal; fully derivable). (2) Re-running the authors' ODE model verbatim on 400 posterior draws reproduces the Fig 2c model fit (median dist 0.42 to data) and confirms every shipped draw has distance <= 0.6 = the stated final threshold, i.e. the posteriors file genuinely is the accepted set. (3) Qualitative claims (k3a+>k3b+, >=1 order-of-magnitude differences) hold. (4) Model-selection direction reproduced (H1 strongly favoured, rising as delta falls). PARTIAL on the single headline number: the stated 99% preference for H1 comes out as 92.55% from the shipped final-iteration file -- same conclusion, but the exact 99% is not derivable from the repo's shipped artifact (likely a different ABC-SMC seed/run or the limiting probability). NOT ATTEMPTED (optional hard ~20%): re-running the full ABC-SMC from scratch (N=10^4 x 15 iterations, ~10^6-10^7 ODE solves, stochastic); Th-1 cell inference; and all wet-lab/biophysics results (SPR, smFRET, mass-spec GSE164479), which are out of scope.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 89
    assessed: 2026-06-15 ⛓ c6f302fe905f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

IL-6 and IL-27 activate the same JAK1/STAT1/STAT3 signaling pathway yet produce opposite immunomodulatory outcomes (pro- vs anti-inflammatory); the paper tests whether differential/competitive binding of STAT1 and STAT3 to phospho-tyrosine motifs on the shared and unique receptor chains (GP130, IL-27Rα) accounts for this functional divergence.

Core claims
  • IL-27 induces more sustained STAT1 phosphorylation than HypIL-6, while both cytokines induce comparable levels and kinetics of STAT3 phosphorylation finding
  • Mathematical and statistical modeling identified STAT3 binding to GP130 and STAT1 binding to IL-27Rα as the main dynamical processes contributing to sustained pSTAT1 levels induced by IL-27 mechanism
  • Mutation of Tyr613 on IL-27Rα decreased IL-27-induced STAT1 phosphorylation by 80% but had limited effect on STAT3 phosphorylation finding
  • Strong receptor/STAT coupling by IL-27 initiates a unique gene expression program that requires sustained STAT1 phosphorylation and IRF1 expression and is enriched in classical Interferon Stimulated Genes finding
  • STAT/receptor coupling for IL-6/IL-27 is altered in SLE patients, who show more potent STAT1 activation correlating with higher STAT1 expression than healthy controls finding
  • Sub-saturating doses of the JAK inhibitor Tofacitinib specifically lower STAT1 activation levels induced by IL-6 finding
  • IL-27 stimulation drives formation of a 1:1 stoichiometric IL-27Rα/GP130 heterodimer at the cell surface, with no pre-assembly detected in the unstimulated state finding
  • A recombinant murine single-chain IL-27 (mIL-27sc) cross-reacts with human receptors and triggers signaling comparable to commercial human IL-27, and a single-chain HyperIL-6 (HypIL-6) fusion reduces IL-6 signaling variability from IL-6Rα expression changes resource
Experimental setups
Assay System Perturbation Readout Platform
high-throughput barcoded flow cytometry (phospho-flow, dose-response and kinetics) primary human Th-1-polarized CD4+ T cells IL-27 or HypIL-6 stimulation (dose and time course) pSTAT1 and pSTAT3 levels (MFI)
flow cytometry (phospho-flow) activated human PBMCs (CD4+, CD8+ T cells) IL-27 (2 nM) or HypIL-6 (20 nM) stimulation pSTAT1 and pSTAT3 kinetics
flow cytometry (phospho-flow kinetics) wt RPE1 cells and RPE1 GP130KO reconstituted with 10x mXFPm-GP130 HypIL-6 (20 nM) stimulation vs unstimulated pSTAT1 and pSTAT3 kinetics; cell surface GP130 levels flow cytometry
flow cytometry (phospho-flow dose-response and kinetics) RPE1 GP130KO, wt RPE1, and RPE1 stably expressing mXFPe-IL-27Rα IL-27 or HypIL-6 stimulation pSTAT1/pSTAT3 dose-response and kinetics
single-molecule dual-color TIRF imaging with nanobody labeling and co-tracking RPE1 GP130KO cells transfected with mXFPe-IL-27Rα and GP130 IL-27 (20 nM) or HypIL-6 (20 nM) stimulation vs unstimulated receptor heterodimer/homodimer co-trajectories, diffusion coefficients TIRF microscopy with dye-conjugated anti-GFP nanobodies (RHO11/DY649)
single-molecule photobleaching step analysis and single-molecule FRET RPE1 cells expressing mXFPe-IL-27Rα and GP130 IL-27 stimulation receptor complex stoichiometry and molecular proximity TIRF microscopy
Key results
  • IL-27 phosphorylates STAT1/STAT3 with EC50 ~20 pM versus ~400 pM for HypIL-6 in Th-1 cells ~20x lower EC50 for IL-27
  • Both cytokines yield the same maximal amplitude for pSTAT3, but HypIL-6 shows significantly reduced maximal pSTAT1 amplitude relative to IL-27
  • pSTAT3 peaks at ~15-30 min then gradually declines for both cytokines, with ~20% activation remaining after 3 hr of continuous stimulation ~20% remaining at 3h
  • 10-fold overexpression of GP130 in RPE1 cells prolongs HypIL-6-induced STAT3 activation but has little effect on STAT1 kinetics 10x GP130
  • IL-27 stimulation induces substantial IL-27Rα/GP130 heterodimerization on the cell surface, not observed at baseline
  • Bleaching analysis confirms a 1:1 stoichiometry of the IL-27Rα/GP130 complex 1:1
  • Purchased human IL-27 and mIL-27sc show comparable potency (EC50 ~70-80 pM vs ~20-30 pM respectively) for pSTAT1/pSTAT3 induction EC50 20-80 pM range
Key statistics
  • other EC50 ~20 pM (IL-27) vs ~400 pM (HypIL-6) (pSTAT1/3 dose-response in Th-1 cells)
  • other EC50 pSTAT1 ~70 pM, pSTAT3 ~80 pM (purchased IL-27) (dose-response in activated CD4+ cells)
  • other EC50 pSTAT1 ~20 pM, pSTAT3 ~30 pM (mIL-27sc) (dose-response in activated CD4+ cells)
  • fold_change 80% decrease (IL-27-induced STAT1 phosphorylation after Y613 mutation on IL-27Rα)
  • other ~20% of maximal activation remaining (pSTAT3 levels after 3 hr continuous IL-27/HypIL-6 stimulation)
  • count minimum of 23 cells measured per condition (single-molecule co-tracking/dimerization analysis)
  • pvalue *p<0.05, **p≤0.01, ***p≤0.001 (two-tailed Student's T-test significance thresholds for co-trajectory/dimerization comparisons)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combines quantitative cell-biology and -omics measurements with mathematical/statistical modeling of STAT1/STAT3 signaling kinetics. Group comparisons of single-molecule co-tracking data were assessed with two-tailed Student's t-tests, and dose-response/kinetic measurements were summarized as mean ± standard deviation from biological replicates each containing technical replicates. Significance was reported using thresholded p-value tiers (e.g., *p<0.05, **p≤0.01, ***p≤0.001) rather than exact values for the displayed comparisons.

Replicationmixed Sample sizedescribed qualitatively per panel (e.g., three or five biological replicates each with two technical replicates; single-molecule panels report a minimum number of cells per condition), with no formal power/sample-size calculation stated GroupsIL-27 vs HypIL-6 (and unstimulated) across cell types; SLE patients vs healthy controls Pairingunclear Randomization/blindingnot stated DispersionSD Exact p-valuesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
two-tailed Student's t-test relative number of co-trajectories for IL-27Rα/GP130 heterodimerization and GP130 homodimerization across stimulation conditions (Figure 1g) each data point = one cell; minimum 23 cells per condition not stated
two-tailed Student's t-test diffusion coefficients of receptors with/without cytokine stimulation (Figure 1—figure supplement 3c) each data point = one cell; minimum 23 cells per condition not stated
Approaches that could also have been used
  • Pairwise group comparisons of single-molecule co-tracking and diffusion data were made with two-tailed Student's t-tests.
    Could also: A non-parametric test such as Mann-Whitney U, or a permutation/bootstrap test, could also be applied. — Per-cell distributions can be skewed or have unequal variance; a rank-based or resampling approach makes fewer distributional assumptions and is often chosen when normality is not formally checked.
  • Several conditions (cytokines, stimulation states) were compared using multiple two-tailed t-tests.
    Could also: A single one-way (or two-way) ANOVA followed by a post-hoc test such as Tukey HSD could also be used. — An omnibus model with post-hoc correction would also control the family-wise error rate across the set of related comparisons in one analysis.
  • Significance was reported as thresholded tiers (*, **, ***).
    Could also: Exact p-values alongside an effect-size estimate (e.g., difference in means with a 95% confidence interval) could also be reported. — Exact values and effect sizes convey both the strength and magnitude of a difference, which complements categorical significance markers.
  • Spread of replicate measurements was summarized with standard deviation.
    Could also: A 95% confidence interval (or showing all individual data points) could also be displayed. — A confidence interval conveys precision of the estimate and is often preferred for small replicate numbers, while plotting individual points shows the full data distribution.
  • Sample sizes were stated descriptively per panel without a formal power analysis.
    Could also: An a priori power/sample-size justification could also accompany the design. — A stated power calculation makes explicit the sensitivity of each comparison and aids interpretation and reproducibility.

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
32
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE164479 GEO in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet
PXD024188 PRIDE in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet
PXD024657 PRIDE in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33871355

Paper: Wilmes et al. (2021) eLife. "Competitive binding of STATs to receptor phospho-Tyr motifs accounts for altered cytokine responses." DOI 10.7554/eLife.66014.

Code: https://github.com/PollyJeffrey/Cytokine_modelling (redirects to PollyJeffrey/Cytokine-modelling-paper), Python, MIT, pinned commit 9c3e0ddc7a96eac941baad560d1541d660b0515d (the exact SHA cited in the paper's Data Availability via Software Heritage swh:1:rev:9c3e0dd…). Zenodo DOI 10.5281/zenodo.4609852.

Data: GEO GSE164479 = phospho-proteomic / proteomic (mass-spec) datasets. NOTE: the GEO accession is not the input to the modelling pipeline. The model is calibrated against pSTAT1/pSTAT3 phospho-flow time courses, which are shipped inside the repo as 8 small normalised .txt files (Eq. 5 normalisation). So the in-scope computational pipeline is fully self-contained in the repo.

In scope (pipeline-derived, attempted)

The repo is an ABC-SMC (Approximate Bayesian Computation – Sequential Monte Carlo) model-selection + Bayesian parameter-inference pipeline (numpy+scipy.odeint). Three scripts: ABC_SMC_model_selection.py, ABC_SMC_RPE1.py, ABC_SMC_TH1.py. Shipped pipeline outputs in the repo: Accepted_models_iteration_0..14.txt, RPE1_posteriors.txt, TH1_posteriors.txt.

Result Paper location Pipeline In scope
Model selection: hypothesis 1 favoured over hypothesis 2 with probability 99 % (RPE1) Results §"…model selection…"; Figure 2b ABC-SMC model selection → Accepted_models_iteration_14.txt YES
Fig 2b trajectory: rel. prob. of H1 rises as distance threshold δ falls Figure 2b same, all 15 iteration files YES
Table 1: posterior mean & median of 8 STAT binding/dissociation rates k1a±,k1b±,k3a±,k3b± (RPE1) Table 1 ABC-SMC inference → RPE1_posteriors.txt YES
Qualitative ordering k3a+ > k3b+ ; "≥1 order-of-magnitude" rate differences Results, Fig 2d RPE1_posteriors.txt YES
Fig 2c: pointwise median + 95% CI model fit to pSTAT1/3 vs time, calibrated by posterior Figure 2c forward ODE sim of authors' model with posterior draws YES (deterministic forward sim)

Out of scope (not attempted, why)

  • Re-running the full ABC-SMC from scratch (N=10⁴ particles × 15 iterations, acceptance rate falls steeply at small δ → ~10⁶–10⁷ ODE solves). Stochastic; the hard last ~20%. We instead (a) verify the shipped pipeline outputs reproduce the reported numbers and (b) re-run the authors' ODE model forward. We do NOT claim a from-scratch re-inference.
  • Wet-lab / biophysics: surface plasmon resonance, single-molecule FRET, mass-spec phospho-proteomics (GSE164479), flow cytometry, microscopy — all experimental, not pipeline-derived.
  • Th-1 cell inference (TH1_posteriors.txt): same machinery as RPE1; we focus the gradeable comparison on RPE1 (where the paper states the 99 % headline and Table 1). Th-1 left as optional.
Figures / tables: Figure 2bTableFigure 2dFigure 2c
c1_modelsel
Reported
P(H1)=99% (Fig 2b, RPE1)
Reproduced
P(H1)=92.55% from shipped Accepted_models_iteration_14.txt
partial
c2_fig2b_traj
Reported
rel. prob. of H1 rises as distance threshold delta falls (Fig 2b)
Reproduced
P(H1) rises 0.579->0.926 over iterations 8->14
exact
c3-c10_table1_medians
Reported
Table 1 medians: k1a+ 2.2e-2, k1a- 1.8e-1, k1b+ 3.4e-1, k1b- 7.2e-1, k3a+ 3.7e-1, k3a- 8.0e-1, k3b+ 7.5e-4, k3b- 4.3e-1
Reproduced
2.19e-2, 1.83e-1, 3.41e-1, 7.19e-1, 3.71e-1, 7.96e-1, 7.49e-4, 4.35e-1
exact
c11-c18_table1_means
Reported
Table 1 means: 1.1e-1, 4.2e-1, 1.2, 1.2, 1.1, 1.4, 6.8e-2, 1.2
Reproduced
1.10e-1, 4.22e-1, 1.20, 1.17, 1.14, 1.35, 6.76e-2, 1.16
within tolerance
c19_order_mag
Reported
STAT binding-rate differences >= 1 order of magnitude; k3a+ > k3b+
Reproduced
k3a+/k3b+ median ratio = 495 (~2.7 orders); k3a+ > k3b+ TRUE
exact
c20_fig2c_fit
Reported
Fig 2c: posterior-calibrated ODE model reproduces pSTAT1/3 kinetics vs data
Reproduced
median fit dist to data = 0.42; all 400 posterior draws have distance <= 0.6 = the stated final threshold
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 89/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

This is a mathematical-modelling paper whose ABC-SMC pipeline and pSTAT time-course inputs are shipped in-repo (GEO GSE164479 is unrelated mass-spec data). The reproduction is essentially 1:1: all 16 Table-1 rate constants match exactly from RPE1_posteriors.txt, the Fig-2c ODE fit reproduces (median dist 0.42; all 400 draws ≤0.6=δ_final), and the qualitative claims (H1 favoured, k3a+>k3b+, ≥1 order of magnitude) hold. The only deviation is the headline 99% → 92.55% P(H1), which is explainable by ABC-SMC stochasticity/seed (a technical, not authors', issue) and does not change any conclusion. Overall yellow: solid reproduction with one explainable, beyond-rounding headline discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

192.3 k
tokens (I/O) · 15.1 M incl. cache
20 min
runtime · 0.06 CPU-h
2.5 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine