Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Exploring the prognostic and diagnostic value of lactylation-related genes in sepsis.

Sci Rep · 2024
L1 73/100 PQI 91
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Reported values are derivable from the shared data
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
73/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 41% of all assessed papers rank 664 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for the PUBLIC-DATA validation claims, which reproduce ~1:1. The paper's two signature genes (S100A11, CCNA2) were validated on public GEO sets. Diagnostic ROC AUC on GSE69528 reproduces essentially exactly: S100A11 0.9607 vs reported 0.961 (exact), CCNA2 0.8865 vs reported 0.890 (within-tol); sample split 83 sepsis/55 control matches. 28-day survival on GSE65682 (n=479): S100A11 low-expression-worse reproduces with log-rank p=0.0056 (matches paper, significant); CCNA2 direction matches (high expr = higher mortality) but a plain median split is not significant (p=0.18) vs the paper's p<0.05 from a multi-cohort meta-analysis -> partial. NOT attempted: the discovery DEG chain (4890 DEGs, 55 DE-LRGs, 9 PPI hub genes) runs on the authors' in-house RNA-seq (CNGBdb CNP0002611) plus an unpublished 332-gene lactylation list, neither derivable from public data -> outside the 80% public-data scope, flagged for human auditor, no fabrication claim. No authors' analysis repo exists (cited 'code' is BGI SOAPnuke, a generic read-cleaner); reproduced via standard tools (GEOquery/pROC/survival) per P16.

💻 Code ↗ 🗄 Data: GSE65682

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 73
    assessed: 2026-06-14 ⛓ 9e0921a37e5e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Lactate-derived histone/protein lactylation may play a role in sepsis pathogenesis, and lactylation-related genes identified from differential expression analysis could serve as prognostic and diagnostic markers for sepsis.

Core claims
  • Intersecting sepsis-associated differentially expressed genes with a curated list of 332 lactylation genes yields 55 sepsis-related lactylation genes. finding
  • PPI network analysis identifies GAPDH, S100A11, H2BC14, PARP1, TP53, CCNA2, NCL, S100A4, and H2BC13 as core lactylation-related genes in sepsis. finding
  • Low S100A11 expression is associated with decreased 28-day survival in sepsis patients. finding
  • Low CCNA2 expression is associated with increased 28-day survival in sepsis patients. finding
  • S100A11 and CCNA2 have high diagnostic accuracy for sepsis based on ROC/AUC analysis. finding
  • Meta-analysis across independent GEO cohorts confirms differential expression patterns of S100A11 and CCNA2 between sepsis survivors and non-survivors. finding
  • Single-cell RNA sequencing shows monocyte macrophages, T cells, and B cells have high expression of the sepsis-associated lactylation hub genes. finding
  • S100A11 and CCNA2 are proposed as lactylation-related biomarkers for sepsis diagnosis, prognosis, and treatment guidance. finding
Experimental setups
Assay System Perturbation Readout Platform
bulk mRNA-seq peripheral blood, human (20 sepsis patients, 10 healthy controls) none (disease state comparison) differentially expressed genes (DEGs) BGISEQ-500/MGISEQ-2000
GO and KEGG enrichment analysis 55 overlapping sepsis-lactylation genes (bioinformatic dataset) none functional/pathway enrichment clusterProfiler (R v4.2.1)
protein-protein interaction (PPI) network analysis 55 overlapping sepsis-lactylation genes none hub gene identification STRING database
survival analysis (Kaplan-Meier, log-rank test) GEO dataset GSE65682, whole blood, human (478 sepsis patients, 365 survivors) none (gene expression stratification) 28-day survival rate by gene expression level GraphPad Prism 8
ROC curve analysis GEO dataset GSE69528, whole blood, human (83 sepsis, 55 normal controls) none diagnostic sensitivity/specificity (AUC) MedCalc
meta-analysis of gene expression (forest plot) GEO datasets GSE54514, GSE63042, GSE95233, whole blood, human (sepsis survivors vs non-survivors) none cross-cohort gene expression levels
single-cell RNA sequencing (10x Genomics) PBMCs from 5 peripheral blood samples (2 healthy, 1 SIRS, 2 sepsis) none (disease state comparison) cell-type-specific expression of hub genes 10x Genomics, Cell Ranger, Seurat
Key results
  • 4890 DEmRNAs identified between sepsis and normal groups (2498 upregulated, 2392 downregulated) |FC| ≥ 2, FDR < 0.05
  • 55 sepsis-related lactylation genes obtained by intersecting DEGs with lactylation gene list
  • 9 hub genes identified at core of PPI network
  • Low S100A11 expression linked to lower 28-day survival P < 0.05
  • Low CCNA2 expression linked to higher 28-day survival P < 0.05
  • S100A11 and CCNA2 show high sensitivity/specificity in ROC analysis AUC = 0.961 (S100A11), AUC = 0.890 (CCNA2)
  • Meta-analysis: S100A11 high in survivors/low in non-survivors; CCNA2 low in survivors/high in non-survivors P < 0.05
  • Monocyte macrophages, T cells, and B cells show high expression of hub genes in scRNA-seq
Key statistics
  • count 4890 DEmRNAs (2498 up, 2392 down) (DEGs between sepsis (n=20) and healthy (n=10) blood samples)
  • fold_change |FC| ≥ 2, FDR < 0.05 (DEG selection criteria)
  • count 55 overlapping lactylation genes (intersection of 4890 DEGs and 332 lactylation genes)
  • count 332 lactylation genes (background lactylation gene list from prior studies)
  • pvalue P < 0.05 (log-rank test, survival difference by S100A11/CCNA2 expression (GSE65682))
  • other AUC = 0.961 (ROC diagnostic accuracy of S100A11 (GSE69528: 83 sepsis, 55 normal))
  • other AUC = 0.890 (ROC diagnostic accuracy of CCNA2 (GSE69528))
  • pvalue P < 0.05 (meta-analysis expression differences across GSE54514, GSE63042, GSE95233)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study performed bulk RNA sequencing on peripheral blood from 20 sepsis patients and 10 healthy controls, applying fold-change and FDR thresholds to identify differentially expressed genes (DEGs), which were intersected with a curated 332-gene lactylation set to yield 55 candidate genes. Hub genes were prioritized by visual PPI network centrality in STRING, yielding S100A11 and CCNA2, which were then validated in independent GEO cohorts via log-rank survival analysis and ROC curve analysis; a meta-analysis across three additional GEO datasets further assessed expression directionality in survivors vs. non-survivors. Single-cell RNA sequencing on five peripheral blood samples characterized cell-type expression patterns of the hub genes, with results reported as threshold p-values (P < 0.05) and AUC point estimates.

Replicationbiological Sample size20 sepsis patients and 10 healthy controls for primary RNA-seq; external validation in GEO datasets (GSE65682 n=802, GSE69528 n=138, GSE54514 n=163, GSE63042 n=106, GSE95233 n=124); no formal power calculation described GroupsSepsis vs. healthy controls (primary RNA-seq); survivors vs. non-survivors (survival and meta-analysis); high vs. low gene expression groups (log-rank) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFDR < 0.05 for DEG discovery (exact correction method not named; iDEP 2.1 default); P < 0.05 without stated correction for survival and ROC analyses across two hub genes; clusterProfiler applies BH FDR correction by default for GO/KEGG enrichment
Statistical tests used
Test Applied to n Assumptions
Differential expression analysis with |FC| ≥ 2 and FDR < 0.05 thresholding (via iDEP 2.1; underlying statistical model not explicitly named) Bulk RNA-seq: sepsis (n=20) vs. healthy controls (n=10) 20 sepsis, 10 healthy controls not stated
Hypergeometric test (GO functional enrichment, via R/clusterProfiler) GO enrichment of 55 lactylation–DEG overlapping genes 55 overlapping genes not stated
Hypergeometric test (KEGG pathway enrichment, via R/clusterProfiler) KEGG pathway enrichment of 55 overlapping genes 55 overlapping genes not stated
Log-rank test 28-day survival analysis for S100A11 and CCNA2 (high vs. low expression groups) in GSE65682 802 total samples in GSE65682; paper states 478 patients with sepsis and 365 survivors not stated
ROC curve analysis (AUC) Diagnostic discrimination of S100A11 and CCNA2 in GSE69528 (sepsis vs. normal control) 138 (83 sepsis, 55 normal) na
Meta-analysis (specific pooling model not stated; forest plot reported) Expression of S100A11 and CCNA2 in survivors vs. non-survivors across GSE54514, GSE63042, GSE95233 163 + 106 + 124 = 393 combined across three datasets not stated
Approaches that could also have been used
  • Survival analysis dichotomized continuous gene expression into high vs. low groups for log-rank testing
    Could also: Cox proportional hazards regression using continuous gene expression as a covariate — Cox regression retains the full continuous expression spectrum, avoids an arbitrary split point, and directly yields hazard ratios with confidence intervals as quantitative effect sizes; it is the conventional approach when expression is treated as a prognostic variable
  • Meta-analysis pooled expression data across three GEO cohorts, but the pooling model (fixed vs. random effects) and between-study heterogeneity were not stated
    Could also: Explicitly specified random-effects meta-analysis (e.g., DerSimonian–Laird) with the heterogeneity statistic I² and Cochran's Q reported alongside the forest plot — Reporting the model choice, I², and individual study weights allows readers to evaluate cross-cohort consistency and is standard practice for meta-analyses of gene expression data; random-effects models are generally preferred when cohort-level differences in platform or population are expected
  • ROC curves were reported as point-estimate AUC values only (0.961 and 0.890)
    Could also: Report 95% confidence intervals around each AUC (e.g., via DeLong's method) and specify the optimal operating threshold with corresponding sensitivity and specificity — CIs around AUC quantify estimation uncertainty — particularly relevant at n=138 — and an explicit threshold with paired sensitivity/specificity directly supports potential clinical translation of a diagnostic marker
  • Hub genes were identified by visual inspection of centrality in a STRING PPI network
    Could also: Compute formal network centrality metrics (degree, betweenness, eigenvector centrality) or apply a penalized regression approach (e.g., LASSO) on the 55-gene candidate set — Quantitative centrality metrics or data-driven variable selection provide a reproducible, algorithmically defined ranking that can be reported and re-run, rather than relying on graphical or qualitative prioritization
  • The specific differential expression model used by iDEP 2.1 was not named, though |FC| ≥ 2 and FDR < 0.05 thresholds were stated
    Could also: Explicitly name and parameterize the DE model (e.g., DESeq2 with negative binomial Wald test and shrinkage estimation, or limma-voom with TMM normalization) — Naming the exact model, normalization strategy, and dispersion estimation approach enables full reproducibility, allows readers to evaluate suitability for the small-n design (n=30 total), and follows community norms for RNA-seq publications
  • Survival and ROC results for two genes were each reported at a nominal P < 0.05 without a stated multiplicity correction
    Could also: Apply a two-test Bonferroni correction (adjusted α = 0.025) or Benjamini-Hochberg FDR across all hub-gene-level tests — Even a minimal correction for two simultaneous gene-level tests keeps the effective significance threshold interpretable and is straightforward to implement; this is particularly relevant if additional hub genes beyond S100A11 and CCNA2 were also examined at this stage
Software: iDEP 2.1 · R/clusterProfiler 4.2.1 · R/ggplot2 · GraphPad Prism 8 · MedCalc · 10x Genomics Cell Ranger · Seurat · BGI SOAPnuke

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
14
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE54514 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE63042 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE95233 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE65682 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE69528 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

C-AUC-S100A11
Reported
0.961
Reproduced
0.9607 (95% CI 0.930-0.991)
exact
C-AUC-CCNA2
Reported
0.890
Reproduced
0.8865 (95% CI 0.830-0.943)
within tolerance
C-SURV-S100A11
Reported
low expr -> decreased 28-day survival (p<0.05)
Reproduced
low_expr_worse; log-rank p=0.0056; death low=0.292 vs high=0.184 (GSE65682, n=479)
exact
C-SURV-CCNA2
Reported
low expr -> improved 28-day survival; high in non-survivors (p<0.05)
Reproduced
high_expr_worse (direction matches); log-rank p=0.18 NS; death high=0.264 vs low=0.212 (GSE65682, n=479)
partial
C-DEG-COUNT
Reported
4890 DEGs (2498 up / 2392 down)
Reproduced
not attempted (in-house CNGBdb CNP0002611 data)
partial
C-DELRG
Reported
55 DE-LRGs
Reproduced
not attempted (depends on in-house DEGs + unpublished 332-LRG list)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 73/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The public-data validation core reproduces ~1:1: S100A11 diagnostic AUC 0.961→0.9607 (exact) and CCNA2 0.890→0.8865 (within-tol) on GSE69528 with matching 83/55 split, and S100A11 28-day survival reproduces with correct direction and significance (p=0.0056). The one substantive deviation — CCNA2 prognostic significance (paper p<0.05 vs reproduced p=0.18) — keeps the correct direction and is explained by our single-cohort median split versus the paper's multi-cohort meta-analysis, i.e. a method choice, not an authors' defect. The discovery DEG chain (4890 DEGs, 55 DE-LRGs) could not be re-derived because it depends on unavailable in-house data and an unpublished gene list, a data-availability limitation rather than fabrication. Overall: solid, with explainable deviations — yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

88.1 k
tokens (I/O) · 4.2 M incl. cache
13 min
runtime · 0.03 CPU-h
1.7 GB
peak RAM
3 (2 failed)
HPC jobs
hummel
machine