IL-6 trans-Signaling Regulates Neutrophilic Inflammation in Alcohol-Associated Hepatitis.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
EXECUTED on «our HPC» (the prior session was a compute-gated 'error' that staged but ran nothing, and was correctly quarantined as bogus). The brief's code pointer is the STAR aligner (third-party, P16-valid) and the data is GSE143318 - a public HUMAN liver RNA-seq set the paper REUSES as a validation cohort (Fig 3D), not the authors' own mouse data. The paper itself only reused GSE143318's shipped processed counts, so the faithful reproduction is the count->DESeq2 stage, which I ran: GSE143318_Rawcount.txt.gz (586,782 ENSEMBL features x 25 samples) -> DESeq2 severe AH (n=13) vs donor control (n=7) -> 7324 DEGs (padj<0.05; 4468 up/2856 down). RESULTS: (M1) the cohort is exactly n=13 AH / n=7 control as the paper states - this OVERTURNS the prior session's 'possible fabrication' flag, which had wrongly claimed GEO=10 AH/5 controls. (R1) 6/8 Fig-3D genes reproduce as up in AH (CXCL1/3/5/8 strongly significant; SOCS3/CCL20 up by mean, ns), BUT SAA1 is ~8x DOWN and CRP ~2x DOWN in AH - directly contradicting the figure's claim that these acute-phase genes are upregulated. The contradiction is driven by very high SAA1/CRP expression in the deceased-donor 'healthy control' livers (a known donor-liver inflammation / hepatocyte-mass confound); flagged for a human to check against the actual Fig 3D panel, provisional, not asserted as fabrication. (R2) my DESeq2 recovers 1293/2255 (57%) of the original GSE143318 submitters' shipped DE genes, with direction+magnitude matching closely on shared key genes (their list was 5v5 with a Cuffdiff-like FPKM method vs my DESeq2 13v7). (R3) IL6R is significantly down in severe AH (padj 9e-8), reproducing the paper's IL6R-reduction claim, with a coherent IL-6 trans-signaling pattern. NOT ATTEMPTED: STAR 2.6.1d realignment from SRA fastq (SRP240640) - the heavy last-20% the paper did not do for this dataset; the authors' own mouse RNA-seq (no accession); HepG2/GSE255379, scRNA-seq/GSE255772, InTeam cohort, IPA, all wet-lab assays (out of scope). CAVEAT: env used DESeq2 1.38.0 (paper: 1.20.0) - a version difference; directional results are robust. All small results + the reproduction figure are under datasets/pmid-40562277/reproduction/; the 57 MB full DESeq2 table stays on «infra» (path + sha256 recorded).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-16 ⛓ 760ccddbbd38
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tested whether IL-6 trans-signaling (via soluble IL-6 receptor), rather than classical membrane-bound IL-6R signaling, drives STAT3 activation in hepatocytes and the resulting elaboration of neutrophilic activators that promote neutrophilic inflammation in alcohol-associated hepatitis (AH).
- ★ Hepatic IL-6R expression progressively declines with increasing severity of alcohol-related liver disease, from normal to early ASH to nonsevere and severe AH finding
- ★ TGF-β1 is the most potent negative regulator of IL-6R expression among factors examined finding
- ★ STAT3-dependent gene expression is increased in severe AH despite reduced IL-6R finding
- ★ TGF-β1 treatment suppresses IL-6R expression in HepG2 cells in vitro finding
- ★ Hyper-IL-6 (trans-signaling agonist), but not classical IL-6, restores STAT3 activation when IL-6R is suppressed finding
- ★ A hyper-IL-6-induced gene signature stratifies a subset of AH patients with enhanced IL-6 trans-signaling activity, increased intrahepatic neutrophilic infiltration, and enrichment of leukocyte migration pathways finding
- ★ In a chronic-plus-binge ethanol mouse model, female mice show enhanced STAT3 activation despite reduced hepatic IL-6R, increased neutrophilic activators, and colocalization of Ly6G+ leukocytes with STAT3+ hepatocytes finding
- ★ IL-6 trans-signaling preserves hepatocyte STAT3-dependent gene expression and promotes neutrophilic inflammation in AH mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA sequencing | human liver biopsies (InTeam cohort: normal, early ASH, nonsevere AH, severe AH) | none (observational, ALD severity spectrum) | IL-6R, IL-6, IL-6ST mRNA expression (tpm) and gene signatures | Illumina sequencing via Novogene, NEBNext Ultra RNA library kit |
| immunohistochemistry/quantitative staining | human liver biopsy sections (healthy control vs AH) | none | IL-6R protein signal intensity | — |
| single-cell RNA sequencing | human hepatocytes (healthy controls vs severe AH; public datasets GSE255772, GSE136103) | none | IL-6R mRNA raw counts | — |
| cell culture treatment + Western blot/qPCR | HepG2 cells | TGF-β1 treatment | IL-6R protein and mRNA expression | — |
| cell culture treatment + Western blot | HepG2 cells (TGF-β1 pretreated) | IL-6 (classical) vs hyper-IL-6 (trans-signaling agonist), low/high dose | STAT3/pSTAT3 activation | — |
| RNA sequencing | HepG2 cells | hyper-IL-6 stimulation | gene expression signature | Illumina sequencing via Novogene |
| chronic-plus-binge ethanol feeding model | female Ptpn11 fl/fl mice, liver | 10-day ethanol diet + single ethanol binge vs pair-fed control | hepatic IL-6R expression, STAT3 activation, neutrophilic activator expression | — |
| immunofluorescence | mouse liver sections | ethanol feeding (chronic-plus-binge) | Ly6G+ leukocyte and pSTAT3+/STAT3+ hepatocyte colocalization, MPO staining | Zeiss LSM 700 confocal microscope |
- ▼ IL-6R protein is significantly reduced in AH liver versus healthy control liver by IHC
- ▼ Hepatocyte IL-6R mRNA is significantly lower in severe AH than healthy controls by single-cell RNA-seq
- ▼ Whole-liver IL-6R mRNA declines progressively across normal, early ASH, nonsevere AH, and severe AH
- ▼ TGF-β1 identified as the most potent negative regulator of IL-6R expression
- ▲ STAT3-dependent gene expression is increased in severe AH
- ▼ TGF-β1 treatment suppresses IL-6R expression in HepG2 cells
- ▲ Hyper-IL-6 restores STAT3 activation despite suppressed IL-6R, while IL-6 alone does not
- – Ethanol-fed mice show enhanced STAT3 activation despite reduced hepatic IL-6R, with increased neutrophilic activators and Ly6G+/STAT3+ colocalization
- pvalue P < 0.05, P < 0.001, P < 0.0001 (denoted *, ***, ****) (IL-6R IHC staining intensity comparison between healthy control and AH liver (Figure 1A))
- pvalue P = 1.906 × 10^-241 (Hepatocyte single-cell IL-6R mRNA expression, healthy controls vs severe AH (Figure 1D))
- count n = 51 total (normal n=10, early ASH n=12, nonsevere AH n=11, severe AH n=18) (InTeam RNA-seq patient cohort)
- count severe AH n=13, healthy controls n=7 (Validation whole-liver RNA-seq dataset GSE143318)
- count severe AH n=5, healthy controls n=5 (Validation hepatocyte single-cell RNA-seq datasets GSE255772/GSE136103)
- count severe AH n=57, nonsevere AH n=17, no liver disease n=16 (Serum proteomics cohort)
- count AH n=6, healthy control n=5 (Genomic DNA methylome (Infinium MethylationEPIC) cohort)
- count AH n=6, healthy control n=4 (ChIP-seq (H3K27ac, H3K4me1, H3K4me3, H3K27me3) cohort)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This multi-component study combined bulk RNA-seq from human liver biopsies (n=51 across four ALD severity groups) with publicly available validation RNA-seq and single-cell RNA-seq datasets, in vitro HepG2 cell experiments, and a 10-day chronic-plus-binge female murine ethanol model. IL-6R expression trajectories across disease severity and STAT3 signaling responses to IL-6 versus hyper-IL-6 were primary outcomes. Significance was reported primarily via asterisk-based P-value thresholds (P < 0.05, < 0.001, < 0.0001), with at least one exact P value (P = 1.906 × 10⁻²⁴¹ for single-cell comparison); the full statistical methods section is not present in the provided text excerpt.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| not stated in available text (significance thresholds *, ***, **** imply pairwise or multi-group inferential test) | IHC optical density quantification of IL-6R: AH patients vs healthy controls (Figure 1A) | n=3 subjects per group | not stated |
| not stated in available text; P = 1.906 × 10⁻²⁴¹ is consistent with Wilcoxon rank-sum or similar non-parametric test used in scRNA-seq pipelines | Single-cell IL-6R mRNA expression: severe AH vs healthy controls (Figure 1D) | n=5 severe AH subjects, n=5 healthy controls (public datasets GSE255772 and GSE136103) | not stated |
| differential expression analysis — specific method not stated in available text; expression summarized as transcripts per million (tpm) | Whole-liver bulk RNA-seq across normal, early ASH, nonsevere AH, severe AH (InTeam cohort, Figure 1C) | n=10 normal, n=12 early ASH, n=11 nonsevere AH, n=18 severe AH | not stated |
| ΔCT method for relative quantification; inferential test not stated in available text | qPCR of target genes in HepG2 cells and mouse liver (Tables 1 and 2) | all samples run in triplicate; biological n not stated in available text | not stated |
| densitometry with normalization to β-actin; inferential test not stated in available text | Western blot quantification of pSTAT3, STAT3, IL-6R, gp130 in HepG2 cells and mouse liver | not stated in available text | not stated |
| Infinium MethylationEPIC BeadChip array; statistical testing method not stated in available text | Genomic DNA methylome analysis: 6 AH explant livers vs 5 healthy controls | n=6 AH, n=5 healthy controls | not stated |
-
Four ALD severity groups were compared, with asterisk-based significance levels implying multiple pairwise tests↳ Could also: One-way ANOVA (or Kruskal-Wallis for non-normally distributed data) with a post-hoc correction such as Tukey HSD or Dunn's test with Benjamini-Hochberg adjustment — An omnibus test followed by a controlled post-hoc procedure is a standard way to formally account for the family-wise error rate when comparing four ordered groups, and would complement the pairwise threshold approach used
-
IHC signal intensity comparisons were made with n=3 subjects per group, with results reported as P < 0.05 / < 0.001 / < 0.0001↳ Could also: Report effect size (e.g., Cohen's d or rank-biserial r) and 95% confidence interval alongside the P value — With n=3 per group, P-value thresholds alone have limited precision; an effect size and CI communicate magnitude and uncertainty, which is especially informative for small-sample comparisons
-
Bulk RNA-seq expression was summarized as transcripts per million (tpm) for visualization across groups↳ Could also: Count-based differential expression with DESeq2 or edgeR (negative binomial model, size-factor normalization, Wald or likelihood-ratio test, Benjamini-Hochberg FDR) — TPM is appropriate for within-sample comparisons and visualization, while DESeq2/edgeR are generally recommended for between-group statistical testing because they model count overdispersion and provide FDR-controlled results
-
Relative gene expression by qPCR was calculated using the ΔCT method, with amplification efficiency measured from a standard curve↳ Could also: Efficiency-corrected ΔΔCT (Pfaffl method) or REST software for statistical comparison of relative expression ratios — Incorporating the measured per-assay efficiency (already collected via standard curve) into the fold-change calculation can improve accuracy when efficiencies deviate from the ideal 100%, and REST provides bootstrapped significance testing of expression ratios
-
A hyper-IL-6 RNA-seq gene signature was used to stratify a subset of AH patients↳ Could also: Gene set enrichment analysis (GSEA) or single-sample GSEA (ssGSEA) scored per patient — GSEA-based approaches yield a continuous enrichment score per sample using the full ranked gene list rather than discrete gene sets, potentially offering finer-grained patient stratification and integration with existing pathway databases
-
Dispersion in figures is described as median bars for Figure 1D; dispersion measure is not stated for most other figures↳ Could also: Consistently report SD or IQR (for skewed/small-n data) alongside means or medians, and follow ARRIVE 2.0 guidelines for the murine experiment reporting — Explicit dispersion measures allow readers to assess data spread and evaluate biological versus statistical significance; ARRIVE 2.0 guidelines specifically recommend reporting variability, sample-size justification, and blinding for animal studies
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40562277
Paper: IL-6 trans-Signaling Regulates Neutrophilic Inflammation in Alcohol-Associated Hepatitis. Am J Pathol 2025. PMID 40562277 / PMC13168973 / DOI 10.1016/j.ajpath.2025.05.023.
Brief's code pointer: https://github.com/alexdobin/STAR (the STAR aligner — a third-party tool, P16-valid). The paper has NO authors' own analysis repo; it names a standard RNA-seq pipeline.
Brief's data pointer: GSE143318.
What the paper's RNA-seq pipeline is
Methods (Bioinformatics): paired-end clean reads → STAR v2.6.1d against mouse GRCm38 → featureCounts v1.5.0-p3 (RPKM) → DESeq2 v1.20.0 (adjusted P < 0.05). Pathway analysis with Ingenuity (IPA, proprietary — out of scope).
What GSE143318 actually is
GSE143318 = "RNAseq analysis ... within the livers of patients with alcoholic hepatitis." Human (Homo sapiens), Illumina NextSeq 500 (GPL18573). 25 samples: 10 alcoholic hepatitis (AH), 5 alcoholic cirrhosis, 5 donor controls. This is an external/public dataset reused by the paper as a validation cohort — NOT the authors' own mouse RNA-seq. GEO ships processed GSE143318_Rawcount.txt.gz (raw counts) and GSE143318_SAH_Diff_genes.xls.gz (the original authors' severe-AH DE list). Raw fastq: SRA SRP240640 / PRJNA600011.
✅ CORRECTION (2026-06-16, verified on «our HPC»): the earlier "n=13/n=7 mismatch / possible fabrication" note was WRONG (it misread the GEO summary as 10 AH). The shipped GSE143318_Rawcount.txt.gz has exactly 7 donor-control columns (N1–N6,N8) + 13 AH columns (AH1–AH64) (+5 cirrhosis AC1–AC5, unused), and the series-matrix titles classify identically to AH=13, control=7, cirrhosis=5. The paper's "severe AH n=13; healthy controls n=7" is CORRECT (M1 = exact); the fabrication flag is overturned.
In scope (pipeline-derived, low-hanging — what we reproduce)
- R1 (Fig 3D): In GSE143318, STAT3-dependent / neutrophil-chemokine / acute-phase genes (SAA1, CRP, SOCS3, CXCL1, CXCL3, CXCL5, CXCL8, CCL20) are upregulated in severe AH vs controls. Reproduce by: GEO raw counts → DESeq2 (AH vs donor control) → check these genes' direction + significance. Deterministic.
- R2 (clean 1:1 on shipped data): reproduce the shipped
SAH_Diff_genesDE list size from the shipped raw counts via DESeq2 (count→DESeq2 stage). Compare DEG count + overlap. - R3 (paper claim, Fig 1C): IL6R reduced in severe AH — check IL6R direction in GSE143318 (note: Fig 1C is the InTeam cohort, not GSE143318; we check IL6R here as supporting/contextual).
Out of scope (not attempted, with reason)
- Full STAR realignment from fastq (SRP240640, ~20 human 150bp PE libraries): the paper's STAR stage. This is the heavy last-20% — large compute + genome index — and the count matrix is already shipped, so the count→DE stage is the clear, low-cost reproduction. Skipped by 80/20.
- The mouse RNA-seq (the pipeline the Methods actually describe, GRCm38): no mouse accession given in the brief; brief's data pointer is the human validation set GSE143318. Not attempted.
- HepG2/IL-6 stimulation DE (2211/398 genes, GSE255379), hepatocyte scRNA-seq (GSE255772), InTeam human cohort, IPA pathway analysis, all wet-lab (flow, IHC, mouse models): out of scope (different accessions / proprietary / non-computational).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
On the same public GSE143318 data the paper reused (cohort confirmed exact at n=13 AH / n=7 control, overturning a prior bogus fabrication flag), the central IL-6 trans-signaling / neutrophilic-chemokine conclusion reproduces cleanly: IL6R down (padj 9e-8), IL6 up, CXCL1/3/5/8 strongly up, corroborated by the submitters' own DE list. The one substantive discrepancy is on the authors'/figure side: Fig 3D claims SAA1 and CRP are upregulated in severe AH, but the cited raw counts show them ~8x and ~2x DOWN — not derivable from the deposited data, most plausibly a deceased-donor control acute-phase confound. Severity is moderate (2/8 peripheral acute-phase genes flip; core story intact), so overall a solid reproduction with one explainable, human-checkable deviation rather than a critical failure.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.