The molecular landscape of sepsis severity in infants: enhanced coagulation, innate immunity, and T cell repression.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Secondary re-analysis paper. Described well enough ONLY for its deconvolution step, and only as a third-party tool: the paper ships NO analysis code -- the sole resolvable code artifact is the cited ABIS deconvolution Shiny app (github.com/giannimonaco/ABIS, P16). The paper's numbers derive from a custom merge of FIVE GEO datasets (we hold GSE25504, 1 of 5) via COCONUT, with no per-sample group assignment provided, so DEG counts / GO / pseudotime / classifier results (C5-C10) are NOT reproducible from shipped artifacts and were not attempted (documented, not fabricated). What we DID reproduce, faithfully and from public data: ran the cited ABIS microarray signature (rlm robust regression, server.R method) on GSE25504 (GPL6947 Illumina neonatal, 26 infected vs 37 control). All four in-scope directional deconvolution claims that anchor the paper title reproduce with matching direction and high significance: Neutrophils UP (p_BH=1e-07), Monocytes UP (1e-02) = 'enhanced innate immunity'; T Naive/Memory DOWN (2e-08) = 'T cell repression'; naive B DOWN (1e-04). Result = PARTIAL: directional reproduction of the deconvolution claim on the paper's own data with the paper's own cited tool; absolute values not comparable (single dataset vs merged matrix); the preprocessCore quantile step was reimplemented in base R due to a pthread bug on the nodes (core rlm unchanged). NOT attempted: the full 5-dataset COCONUT pipeline, DE, GO, Monocle3, RF, and the 'enhanced coagulation' DE/GO finding -- all lacking shipped code.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-15 ⛓ d24f45716913
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusBecause infant immune responses differ from adults and existing adult-derived sepsis gene signatures may not apply, the study asks what transcriptomic processes characterize sepsis severity and progression to septic shock in infants under 6 months, using a merged multi-study microarray dataset.
- ★ Most published adult/other-cohort sepsis gene signatures have limited utility for infant sepsis; only 2 of 7 achieved >80% accuracy in infants finding
- ★ Progression to septic shock in infants involves late-stage induction of clotting/coagulation factors, heightened innate immunity, and suppression of adaptive (T cell) functionality finding
- ★ Pseudotime analysis of individual gene expression profiles reveals a continuum of molecular change forming tight clusters between healthy controls and septic shock concurrent with disease progression finding
- ★ Disease progression is marked by a transition from adaptive immune cues (e.g., IFN-gamma production) toward greater activation of innate cells and pathways mechanism
- ★ COCONUT co-normalization removes platform, age, and sex biases while preserving case-control gene expression contrast across merged datasets method
- A merged multi-transcriptomic infant sepsis dataset (5 studies, n=335) was assembled as a resource for comparative analysis resource
- Random forest-derived gene modules can classify individuals as Healthy Control or Case method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-blood microarray transcriptomics (re-analysis of public datasets) | Infant whole blood (<6 months), bacteremia/septic shock/healthy controls, n=335 | none (observational disease vs control) | Genome-wide gene expression | Illumina HumanHT-12 V4.0/V3.0, Codelink 55K, Affymetrix HG-U219/U133 Plus 2.0/Custom HTA (6 platforms) |
| COCONUT co-normalization and concordance assessment | Merged multi-study infant dataset (8466 shared genes) | none | Normalized log2 expression; housekeeping (ATP6V1B1, GAPDH) and infection genes (CEACAM1, DYSF) | COCONUT algorithm (R) |
| Published gene-signature classification (AUROC/confusion matrix) | Merged infant dataset, Case vs Healthy Control | none | Accuracy, sensitivity, specificity of 7 signatures (SMS, NS, PD25, PD3, SLS, RG, GD) | Caret v.3.45 (R) |
| Pseudotime trajectory analysis | Individual subject expression profiles, all groups | none | Pseudotime clusters/continuum and group-specific marker genes | Monocle3 (UMAP, preprocess_cd num_dim=50) |
| Differential expression and GO over-representation enrichment | Pairwise disease groups / pseudotime clusters | none | DEGs (log2FC>1, BH p.adjust<0.05) and enriched biological processes | Wilcoxon rank-sum; clusterProfiler / org.Hs.eg.db |
| Immune cell type deconvolution | Merged infant blood samples (288 samples after filtering) | none | Relative (22 cell types) and absolute (29 cell types) immune cell proportions | CIBERSORTx (LM22) and ABIS Shiny app |
| Random forest classification of gene modules | Merged dataset, Healthy Control vs Case | none | Feature importance (mean decrease accuracy >0.5%), classification scores, sensitivity/specificity/accuracy | MetaboAnalyst 5.0; Caret v.3.45 |
- – Only 2 of 7 published sepsis gene signatures reached >80% accuracy in the infant cohort 2 of 7; accuracy >80%
- – COCONUT-normalized data correlated strongly with pre-normalized distribution while reducing platform bias r=0.982
- – Housekeeping genes (ATP6V1B1, GAPDH) showed much smaller variance after normalization while infection genes (CEACAM1, DYSF) remained elevated
- – Sweeney et al. neonatal signature achieved high accuracy for sepsis classification across three cohorts (cited prior work) accuracy=0.9
- – Pseudotime clustering showed bacteremia subjects spread along a continuum fitting Healthy Control-like, Septic Shock-like, or transitory clusters with a shift from adaptive to innate immunity
- – After deconvolution filtering (p<0.05), 288 samples retained and all but one Septic Shock sample dropped, excluding that group from CIBERSORTx analysis 288 samples
- count 335 (Total subjects in merged dataset (Bacteremia 151, Septic Shock 30, Healthy Controls 154))
- correlation cor = 0.982, p-value < 2.2e-16 (Pearson correlation of pre vs post COCONUT normalized distributions)
- mean 18 days (95% CI 15-22) (Overall mean age of subjects)
- count 59% male / 41% female (Estimated sex distribution (available for 3 of 5 datasets))
- count 8466 genes (Genes present across all platforms retained for analysis)
- other median 21,107 (18,947 to 22,296) (Median gene number per platform after probe summarization)
- other accuracy > 80% (Threshold met by only 2 of 7 published adult/other signatures)
- pvalue p.adjust < 0.05; log2FC >1 (DEG significance thresholds (Wilcoxon, BH-adjusted))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study re-analyzed five publicly available whole-blood microarray datasets from infants with bacterial sepsis (total n=335), co-normalizing them with the COCONUT empirical-Bayes algorithm before downstream analyses. Differential gene expression between disease groups (Bacteremia, Septic Shock, Healthy Controls) and pseudotime-derived clusters was assessed with Wilcoxon rank-sum tests, with Benjamini-Hochberg FDR correction applied throughout. Classification performance of published sepsis gene signatures and newly derived gene modules was evaluated via AUROC and a random forest approach, with accuracy, sensitivity, and specificity reported.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Two-sided Wilcoxon rank-sum test | Pairwise DEG analysis between disease groups (Bacteremia, Septic Shock, Healthy Controls) and between pseudotime clusters | Subsets of the merged dataset (Bacteremia n=151, Septic Shock n=30, Healthy Controls n=154; total n=335) | not stated |
| Kruskal-Wallis one-way analysis of variance by ranks | Multiple-group comparisons throughout (stated as the default for >2 groups in the Statistical Analysis section) | Up to n=335 | not stated |
| Pearson's correlation | Assessment of pre- and post-COCONUT normalization distributions (cor=0.982, p<2.2e-16) | n=335 | not stated |
| Area under the ROC curve (AUROC) | Classification performance of seven published sepsis gene predictor sets applied to the merged infant dataset | n=335 | na |
| Random forest classification | Classification performance of gene modules and the Garnett marker signature (MetaboAnalyst 5.0) | n=335 | not stated |
| Monocle3 topmarker() test (q-value < 0.01, specificity ≥ 0.50) | Identification of genes most specifically expressed in each group along the pseudotime trajectory | n=335 | not stated |
-
Differential gene expression was identified using a two-sided Wilcoxon rank-sum test applied pairwise across disease groups↳ Could also: Linear model-based methods such as limma (with voom or lmFit for microarray data) could also be applied, modeling all group contrasts simultaneously within a single framework — A single linear model framework would handle all pairwise contrasts jointly, naturally propagating variance estimates across groups, and is widely used for microarray meta-analyses; it also provides moderated t-statistics that stabilize estimates for genes with small within-group variance
-
Batch effects across five microarray platforms and studies were addressed with COCONUT co-normalization using control samples↳ Could also: ComBat (without requiring matched controls) or surrogate variable analysis (SVA) could also model and remove latent batch structure — ComBat and SVA do not require explicitly matched control samples in every batch, making them applicable when control representation is uneven; SVA additionally estimates unknown sources of variation that may not correspond to known platform labels
-
Gene module and signature classification performance was evaluated using a random forest model with mean decrease accuracy for feature selection↳ Could also: Regularized regression approaches such as LASSO or elastic net could also be used to select a sparse gene set while simultaneously fitting a classification model — Regularized regression provides an explicit penalization framework that directly controls model complexity and yields calibrated coefficient estimates, which can facilitate interpretation of gene contributions and generalization to external cohorts
-
The study aggregated five datasets from different platforms and populations into a single merged dataset for joint analysis↳ Could also: A formal random-effects meta-analysis of effect sizes computed within each dataset independently could also integrate evidence across studies — A meta-analytic framework explicitly models between-study heterogeneity and yields study-weighted summary effect estimates with associated confidence intervals, which directly quantifies how consistently a finding replicates across the contributing cohorts
-
Age (mean 18 days, 95% CI 15-22) was reported as the primary demographic summary, and sex was available for only three of five datasets↳ Could also: Sensitivity analyses stratifying by postnatal age or including age and sex as covariates in the expression model could also be performed — Postnatal age is a well-documented driver of neonatal immune gene expression; explicit covariate adjustment or stratified analyses would help distinguish age-driven expression variation from disease-driven variation, particularly given the wide age range (0-140 days) across contributing datasets
-
Classification performance of seven published gene signatures was compared using AUROC and threshold-based accuracy metrics↳ Could also: Confidence intervals for AUROC (e.g., via DeLong's method) and pairwise statistical comparisons of AUROCs across signatures could also be reported — Point-estimate AUROCs alone do not convey uncertainty; CIs and formal pairwise comparisons would clarify whether observed differences in classification performance across the seven signatures are statistically distinguishable given the sample sizes available
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-38817614
Title: The molecular landscape of sepsis severity in infants: enhanced coagulation, innate immunity, and T cell repression. Huang SSY, Toufiq M, Eghtesady P, Van Panhuys N, Garand M. Front Immunol 2024. PMID 38817614 · PMCID PMC11137207 · DOI 10.3389/fimmu.2024.1281111
Nature of the study
This is a secondary re-analysis (no new wet-lab data). The authors merge five public microarray series across six platforms and re-analyze them:
| Accession | Role in paper |
|---|---|
| GSE25504 | neonatal sepsis (whole blood) — our BRIEF accession |
| GSE64456 | pediatric febrile infants |
| GSE69686 | neonatal |
| GSE26378 | pediatric septic shock |
| GSE26440 | pediatric septic shock |
Combined n = 335 → Bacteremia n=151, Septic Shock n=30, Healthy Controls n=154 (per Methods, verbatim sample counts).
Pipeline (as described in Methods):
- Per-dataset normalization (RMA / Illumina neqc), log2, probe→gene by mean.
- COCONUT cross-platform co-normalization using controls → 8466 common genes.
- Differential expression: two-sided Wilcoxon rank-sum, log2FC>1, BH p.adj<0.05.
- Immune deconvolution: CIBERSORT (LM22, 22 types) and ABIS (29 types, via the Shiny app https://github.com/giannimonaco/ABIS).
- GO enrichment: clusterProfiler (minGSSize=3, FDR<0.05, dispensability 0.4).
- Trajectory: Monocle3 (num_dim=50, UMAP) → 3 pseudotime clusters; topmarker().
- Classification: random forest via MetaboAnalyst 5.0; published sepsis signatures.
Code availability — CRITICAL CONSTRAINT
The paper provides no repository for its own analysis scripts. The only code artifact cited is the third-party ABIS deconvolution Shiny app (github.com/giannimonaco/ABIS, last push 2020-04-28, no license). All other steps (COCONUT merge, sample→group assignment, Wilcoxon DE, GO, Monocle3, RF) have no shipped code and no per-sample group table. This is the dominant reproducibility gap and is recorded honestly below.
IN SCOPE (will attempt — repo-anchored, P16)
The one step backed by a resolvable code artifact applied to the paper's own data:
- R1 — ABIS immune-cell deconvolution of GSE25504 (neonatal whole-blood,
Illumina HumanHT-12 / GPL6947), infected (bacteremia) vs control, using the cited
ABIS microarray signature (
sigmatrixMicro.txt) and the exactserver.Rmethod (rlmrobust regression, quantile-normalize to shipped target, ×100). Pipeline: ABIS (cited repo).- Tests the directional claims that anchor the paper's title:
- neutrophils ↑ in infected vs control ("enhanced innate immunity")
- T cells (CD4/CD8, naïve) ↓ ("T cell repression")
- naïve B cells ↓
- NOTE: the paper reports deconvolution on the merged 5-dataset matrix grouped Bacteremia/Shock/HC, not on GSE25504 alone — so this is a directional / partial reproduction of the deconvolution claim on the paper's own data with the paper's own cited tool, not an exact value match. Graded accordingly.
- Tests the directional claims that anchor the paper's title:
OUT OF SCOPE (not attempted — reason recorded, per 80/20 + no-completeness rule)
- Full 5-dataset COCONUT-merged matrix — no shipped code; per-sample group
assignment across the 5 series is not provided → cannot reconstruct the exact
335-sample / 3-group matrix the paper's numbers derive from. (
docs_insufficientfor the merge step.) - Wilcoxon DE / DEG counts (Fig 4), GO enrichment (Fig 5/6, Table 2), Monocle3 pseudotime clusters (Fig 3), RF classification & published-signature scoring (Fig 2) — all computed on the merged matrix with no shipped code; the reported values (e.g. "12 DEGs Bacteremia-vs-Shock", "106 DEGs cluster2-vs-3", GO-term counts, classifier accuracy 0.81) are not pinnable to a runnable artifact. Not attempted.
- CIBERSORT/CIBERSORTx (LM22) — the paper's primary deconvolution, but it is not the cited code and CIBERSORTx is access-gated (registration/license). We run the cited
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The four deconvolution claims that anchor the paper's title — innate-immunity activation (neutrophils/monocytes up) and T cell repression (T cells down) — reproduce with matching direction and high significance using the paper's own cited ABIS tool on public GSE25504, with no fabrication signal. The principal limitation is on the authors' side: no analysis code and no per-sample grouping for the merged 5-dataset COCONUT matrix were released, so C5–C10 (DEGs, GO, pseudotime, RF acc=0.813, cor=0.982) are not derivable from shipped artifacts. A secondary, explainable gap is our methodology (single dataset vs merged matrix → only directional, not absolute, comparison; base-R preprocessCore substitution). Overall a solid partial reproduction of the testable core with deviations attributable to missing released code rather than incorrect numbers.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.