Reverse Engineering of the Pediatric Sepsis Regulatory Network and Identification of Master Regulators.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL but thorough reproduction of Oliveira et al. 2021 (pediatric sepsis RTN/ARACNe master regulators). Re-ran the authors' OWN RTN pipeline (dalmolingroup/gene-regulatory-networks) on the paper's OWN GEO microarray data (GSE13904 + GSE4607 networks; GSE26378 signature) end-to-end on «our HPC», BOTH cohorts. EXACT: cohort sizes 99/108 (C5), signature groups 82/21 (C4). PARTIAL on the headline MRA candidate counts: net1=261 vs reported 233 (ratio 1.12), net2=265 vs reported 215 (ratio 1.23) on the repo-code signature path - same order of magnitude, same method; and 12/15 named conserved master regulators recovered as candidates in BOTH networks (paper-text signature). Counts are not bit-identical because the repo sets NO random seed (stochastic MI network inference), RTN is 2.26.0 (2026) vs the authors' ~2021 build, and TF mapping used hgu133plus2.db vs the authors' live biomaRt - the last directly explains 2 of the 3 missing conserved MRs (ZKSCAN8, ZNF134 not in our TF set). All documented, no fabrication. Two genuine engineering fixes, both saved to the RTN kartei card: (1) tni2tna.preprocess must be called WITHOUT phenoIDs in RTN 2.26.0 (its probe->symbol remap crashes on the probe-keyed network with 'subscript out of bounds'); (2) a snow-parallel override of RTN's serial tni.dpi, VALIDATED bit-identical to serial on net1 (max_abs_diff=0), to bring net2's dpi.filter (which exceeds the 12h walltime cap serially) down to 84 min. NOT attempted (out of scope, downstream/manual): RA/MS cross-filter, cluster A/B assignment, RedeR figures.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 67assessed: 2026-06-21 ⛓ 895b0ea4970f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether a discrete set of transcription factors act as master regulators driving the transcriptional state of pediatric sepsis, distinguishable from transcriptional changes shared with other unrelated inflammatory conditions.
- ★ A set of 15 TFs was identified as sepsis-specific master regulators of pediatric sepsis, dividing into two non-overlapping clusters. finding
- ★ Cluster A TFs (e.g., GATA3, RORA) show decreased activity/expression in pediatric sepsis, while Cluster B TFs (MEF2A, RFX2, TRIM25) show increased activity/expression. finding
- ★ Regulatory networks were reconstructed from whole-blood gene expression using mutual information (RTN/ARACNe) to build TF-centered regulons. method
- ★ Filtering candidate master regulators against unrelated inflammatory disease signatures (rheumatoid arthritis, multiple sclerosis) increases specificity of true sepsis master regulators. method
- ★ TRIM25, RFX2, and MEF2A form a coordinated regulatory cluster not previously described together in pediatric sepsis. finding
- GATA3 regulon shows the highest similarity between the two independently reconstructed networks, while ZNF134 shows the least. finding
- MEF2A is the only one of the 15 regulons enriched for apoptosis-related biological processes. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| microarray gene expression / co-expression network reconstruction (RTN/ARACNe) | whole blood, pediatric sepsis/SIRS/septic shock patients (GSE13904) | none (observational, disease vs disease-state comparison) | TF-centered regulatory network (regulons) | Affymetrix Human Genome U133 Plus 2.0 |
| microarray gene expression / co-expression network reconstruction (RTN/ARACNe) | whole blood, pediatric sepsis/SIRS/septic shock patients (GSE4607) | none (observational) | TF-centered regulatory network (regulons) | Affymetrix Human Genome U133 Plus 2.0 |
| differential gene expression (limma) | whole blood, pediatric septic shock vs healthy children (GSE26378) | none (case-control) | sepsis gene signature (differentially expressed genes) | Affymetrix Human Genome U133 Plus 2.0 |
| differential gene expression (limma) | CD4+ T-cells, rheumatoid arthritis patients vs healthy donors (GSE56649) | none (case-control) | RA gene signature (differentially expressed genes) | Affymetrix Human Genome U133 Plus 2.0 |
| differential gene expression (limma) | PBMC, multiple sclerosis patients vs healthy donors (GSE21942) | none (case-control) | MS gene signature (differentially expressed genes) | Affymetrix Human Genome U133 Plus 2.0 |
| Master Regulator Analysis (hypergeometric enrichment test) | regulons from Pediatric Sepsis 1 and 2 networks vs sepsis/RA/MS signatures | none (computational) | enrichment significance of regulon for gene signature | RTN R/Bioconductor package |
| regulon activity analysis (GSEA2, Normalized Enrichment Score) | whole blood, pediatric sepsis patients (GSE13904, GSE4607) | none | TF/regulon activity level per patient | RTN R/Bioconductor package |
| Gene Ontology functional enrichment | regulon gene sets (merged from both networks) | none | enriched Biological Process terms per regulon | clusterProfiler R/Bioconductor package |
- – MRA identified 233 MRs in Pediatric Sepsis 1 network and 215 in Pediatric Sepsis 2 network, with 179 overlapping between both
- – 15 MRs were specifically enriched only in the sepsis signature (not RA or MS), forming the final prioritized master regulator list
- ▼ 10 of 15 MRs (Cluster A) showed lower expression in septic samples compared to controls
- ▲ MEF2A, RFX2, and TRIM25 (Cluster B) showed higher expression in septic samples compared to controls
- – ZNF235 and ZKSCAN8 showed no significant difference in expression between sepsis and control samples
- – GATA3 regulon was the most similar (largest shared gene overlap) between Pediatric Sepsis 1 and 2 networks; ZNF134 was the least similar
- – MEF2A regulon uniquely enriched for apoptosis-related GO Biological Process terms among the 15 MRs
- – TRIM25 regulon enriched for neutrophil activation and erythrocyte maturation processes
- count 233 MRs (Pediatric Sepsis 1), 215 MRs (Pediatric Sepsis 2), 179 overlapping (initial MRA candidate master regulators before specificity filtering)
- count 50 MRs retained within cluster A/B of both networks (MRs contained in clusters A and B present in both networks)
- count 15 final sepsis-specific master regulators (MRs enriched only in sepsis signature, not RA or MS)
- pvalue BH adjusted p < 0.05 (significance threshold for TNI permutation testing (10,000 permutations) and MRA hypergeometric enrichment)
- pvalue Bonferroni-adjusted p < 0.01 (pairwise Wilcoxon's test for MR gene expression comparison between case and control samples)
- count case = 99, control = 18 (Pediatric Sepsis 1 dataset (GSE13904) samples used for MR expression comparison)
- count case = 108, control = 15 (Pediatric Sepsis 2 dataset (GSE4607) samples used for MR expression comparison)
- count sepsis n=82, control n=21 (GSE26378); RA n=13, control n=9 (GSE56649); MS n=12, control n=15 (GSE21942) (sample sizes for gene signature datasets used in Master Regulator Analysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This retrospective bioinformatics study reconstructed gene regulatory networks from two independent public pediatric sepsis microarray datasets (GSE13904, n=99; GSE4607, n=108) using mutual information-based ARACNe via the RTN package, then identified master regulators (MRs) by hypergeometric enrichment testing of limma-derived differential expression signatures. Specificity of candidate MRs was assessed by cross-testing against rheumatoid arthritis and multiple sclerosis signatures from unrelated datasets. Regulon activity was quantified with a two-tailed GSEA (GSEA2) normalized enrichment score, and MR gene expression in septic versus control samples was compared with pairwise Wilcoxon tests corrected by Bonferroni.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Permutation-based mutual information (ARACNe) with Benjamini-Hochberg correction (10,000 permutations) | TF–gene regulatory network reconstruction in Pediatric Sepsis 1 and Pediatric Sepsis 2 networks | 99 samples (GSE13904) and 108 samples (GSE4607) | not stated |
| limma moderated t-test (linear model for microarrays), Benjamini-Hochberg adjusted p < 0.05 | Differential expression to derive gene signatures for sepsis (GSE26378), rheumatoid arthritis (GSE56649), and multiple sclerosis (GSE21942) | Sepsis: 82 case + 21 control; RA: 13 case + 9 control; MS: 12 case + 15 control | not stated |
| Hypergeometric test (Master Regulator Analysis), Benjamini-Hochberg adjusted p < 0.05 | Enrichment of disease gene signatures within each TF regulon to identify master regulator candidates | null | not stated |
| Two-tailed Gene Set Enrichment Analysis (GSEA2) — Normalized Enrichment Score | Per-sample regulon activity estimation across both pediatric sepsis cohorts | null | not stated |
| Pairwise Wilcoxon test with Bonferroni correction, p < 0.01 | MR gene expression comparison between septic cases and healthy controls in both datasets; also used for best-probe selection per gene | Pediatric Sepsis 1: case n=99, control n=18; Pediatric Sepsis 2: case n=108, control n=15 | not stated |
| GO Biological Process over-representation test (clusterProfiler) with Benjamini-Hochberg correction, adjusted p < 0.05 | Functional enrichment of gene sets regulated by each of the 15 prioritized master regulator regulons | null | not stated |
-
Master Regulator Analysis applied a hypergeometric test to a binary list of genes passing a BH-adjusted p < 0.05 threshold from limma↳ Could also: Rank-based enrichment methods such as GSEA or fgsea applied to the full continuously ranked gene list (e.g., ranked by limma t-statistic or log-fold-change) could also have been used for MRA — Using a continuous ranking avoids the information loss from binary thresholding and reduces sensitivity to the chosen significance cutoff; the RTN package itself supports GSEA2-based MRA as an alternative to the hypergeometric approach
-
Pairwise Wilcoxon tests with Bonferroni correction were used to compare MR gene expression between septic and control groups↳ Could also: The limma moderated t-test framework already used for differential expression upstream could also have been applied for these same group comparisons — Limma's empirical Bayes variance shrinkage is particularly suited to small control groups (n=15–18) and would provide consistent methodology across the pipeline; it also produces log-fold-changes and standard errors alongside p-values
-
Regulon activity was summarized by the GSEA2 Normalized Enrichment Score, a cohort-relative measure↳ Could also: Per-sample, absolute regulon activity scoring methods such as VIPER (also implemented in RTN/Bioconductor) or ssGSEA could also have been applied — Per-sample scores enable direct group-level statistical comparisons with standard tests and facilitate integration with individual patient clinical metadata
-
Two independent networks were inferred from two separate datasets and compared qualitatively by overlap counting to assess reproducibility↳ Could also: A quantitative network meta-analysis — such as batch-correcting and pooling expression matrices before inference, or using ensemble network averaging — could also have been applied — Quantitative integration would yield a single consensus network with combined statistical power; the current overlap-based approach provides a valid robustness check but does not formally quantify agreement between networks
-
Bonferroni correction was applied to Wilcoxon MR expression comparisons while Benjamini-Hochberg FDR was used for all other multiple-testing steps↳ Could also: BH-FDR could also have been applied consistently to the MR expression comparisons, as it was for all upstream analysis steps — Bonferroni is more conservative than BH-FDR at the same alpha level; consistent use of BH-FDR throughout the pipeline would make error-rate control uniform and potentially improve sensitivity for the MR expression step with small control groups
-
Differential expression results and MR expression comparisons were reported solely in terms of significance thresholds, without accompanying effect sizes or fold changes↳ Could also: Log2 fold changes (for limma DE) and rank-biserial correlations or common language effect sizes (for Wilcoxon comparisons) could also have been reported alongside adjusted p-values — Effect sizes allow readers to assess biological magnitude independently of sample size and are recommended alongside p-values by many biomedical reporting guidelines to support downstream prioritization and replication
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34680414
Paper: Oliveira RAC, Imparato DO, Fernandes VGS, Cavalcante JVF, Albanus RD, Dalmolin RJS (2021). Reverse Engineering of the Pediatric Sepsis Regulatory Network and Identification of Master Regulators. Biomedicines 9(10):1297. PMID 34680414 · PMCID PMC8533457 · DOI 10.3390/biomedicines9101297.
Code: https://github.com/dalmolingroup/gene-regulatory-networks
— pinned commit 2d594ff00024212f7417831105f4739c9414145a (HEAD, 2019-10-22, not
archived). The authors' own step-by-step RTN pipeline (P16: authors' code).
Data accessions (all GEO, all Affymetrix HG-U133 Plus 2.0):
- GSE13904 — Pediatric Sepsis 1 network reconstruction cohort (whole blood).
- GSE4607 — Pediatric Sepsis 2 network reconstruction cohort (whole blood).
- GSE26378 — sepsis (septic-shock vs control) gene signature (82 shock, 21 control).
- GSE56649 (RA, CD4+) & GSE21942 (MS, PBMC) — used only to filter out MRs whose regulons enrich in non-sepsis inflammatory signatures.
Method (what the paper / repo does)
A transcriptional regulatory network is reverse-engineered with the RTN (Reconstruction of Transcriptional Networks) Bioconductor package, which implements ARACNe (mutual-information inference + DPI pruning):
- RMA normalize raw CELs (
affy::ReadAffy→rma). Probe→gene via GPL570. - TF list: 1,388 human transcription factors from the Fletcher2013b package, mapped to U133 Plus 2.0 probes.
- TNI (
tni.constructor→tni.permutation→tni.bootstrap→tni.dpi.filter) builds the regulon network from the cohort's expression. - Signature: GSE26378 septic-shock-vs-control differential expression
(
limma, BH, p<0.01, |logFC|>1) →pheno+hits. - MRA (
tni2tna.preprocess→tna.mra): hypergeometric overlap of each regulon with the sepsis signature (BH-adjusted) → master-regulator candidates. - The same is done independently on both cohorts; MRs significant in both networks, after removing RA/MS-shared regulons and clustering, give the 15 conserved master regulators.
Reported headline numbers (targets — see original/claims.tsv)
- Pediatric Sepsis 1 (GSE13904): 233 MR candidates.
- Pediatric Sepsis 2 (GSE4607): 215 MR candidates.
- 179 MRs shared between the two networks.
- 15 conserved master regulators (final): GATA3, HOXB2, KLF12, MEF2A, NR3C2, RFX2, RORA, TRIM25, ZKSCAN8, ZNF134, ZNF234, ZNF235, ZNF329, ZNF331, ZNF529 (12 down-regulated cluster A; 3 up-regulated cluster B = MEF2A, RFX2, TRIM25).
In scope (pipeline-derived, attempted)
| id | result | pipeline step | reproducibility from public data |
|---|---|---|---|
| C1 | Pediatric Sepsis 1 MRA candidate count (paper: 233) | RTN TNI+TNA MRA on GSE13904 vs GSE26378 signature | YES — all inputs are public GEO; re-run the authors' RTN pipeline with default params |
| C2 | Recovery of the 15 conserved master regulators as significant MRs in the GSE13904 network | TNA MRA output regulon list | YES (partial by nature) — the 15 must appear among net-1 candidates; exact p-values are stochastic (no seed in repo) |
Out of scope (not attempted — and why)
- Second network (GSE4607) + the 179-overlap + the final 15-intersection. Reproducing both networks doubles the (heavy) MI compute; per the 80/20 rule we reproduce network 1 (GSE13904, this RU's named dataset) and check the 15 named MRs are recovered there (a necessary condition for the conserved set). Skipped to avoid the hard/expensive last 20%, not because it is infeasible.
- RA/MS regulon-filtering step (GSE56649, GSE21942) and the clustering / cluster-A-vs-B up/down assignment — secondary curation on top of the MRA.
- GSEA-2 / regulon-activity heatmaps, RedeR tree-and-leaf figures — visualization.
- Exact bit-identical p-values — the repo calls
tni.permutation/tni.bootstrapwith **no `set
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.