Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

An OMICs-based meta-analysis to support infection state stratification.

Bioinformatics · 2021
L1 44/100 3/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
44/100
Reproducibility score
1.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 5% of all assessed papers rank 1109 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to RUN the authors' own pipeline 1:1 (P16: own R code @20742a6d on its own shipped example data), with a partial 1:1 match on the headline classifier. The repo ships two demo files used by the two scripts: example_raw.RDS (3 Affymetrix platform slices) for preprocess.R and example_bc.RDS (a ready 60/20/20 train/test/val split, 982 samples, 384 genes) for backwardsElimination.R. I rebuilt the env on «our HPC» (r-base 4.3.3 + sva/ranger/caret/varSelRF) and ran both steps on «infra». CLASSIFIER (the Table-4 result): the backward-elimination RF reproduces the Control and Viral classes within tolerance on the held-out validation set (Control sensitivity 0.556 vs reported 0.57, Control specificity 0.942 vs 0.96, Control balanced-acc 0.749 vs 0.78; Viral sensitivity 0.952 vs 0.97), but Bacterial collapses (sensitivity 0.413 vs 0.90) because the shipped demo carries only ~9% bacterial (53 train / 15 val) vs the paper's full 18.7% (314); gene-set size is smaller (best 21 vs 33) because the example config is light (trees=100, runs=30, 384 pre-filtered genes) vs the paper's 240 procedures over 13383 genes. BATCH CORRECTION: the two-step ComBat pipeline runs end-to-end and is deterministic, but example_raw's 3 platform slices intersect in only 21 genes (0 overlap with the 384-gene example_bc), so the shipped raw and bc files are independent demo artifacts, not a matched pair -> the corrected values can't be 1:1 validated against a shipped target. NO FABRICATION SIGNAL: the Control-class numbers reproduce almost exactly (incl. n=241 exact) on the authors' own code; every gap is explained by demo down-sizing + lighter config, and Table 4 is plausibly derivable from the full GSE162329/GSE162330 cohort (not shipped as the demo). NOT ATTEMPTED (the hard 20%): the GALGO arm (Table 4 GA columns) because the GALGO package is 'currently unsupported by R 3.6' per the authors' own README (env_unresolvable); the full 240-procedure production run on the complete GSE162329/GSE162330 matrices; the Illumina arm (example data is Affy only); and the downstream network reverse-engineering / pathway-convergence analysis (external, not in the repo scripts).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 44
    assessed: 2026-06-15 ⛓ 540ad978c238
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a generalizable, multi-class predictive biomarker panel be developed from a meta-analysis of human blood gene expression data to distinguish between bacterial, viral, and no-infection disease states using machine learning?

Core claims
  • Multi-class machine learning models built from cross-platform microarray meta-analysis can distinguish bacterial, viral and no-infection states with high accuracy (best model: 93% bacterial, 89% viral correct). finding
  • Although gene-level feature sets differ between Affymetrix and Illumina models, selected features converge to the same functional network regions (Type I interferon signalling, chemotaxis, apoptosis, inflammatory/innate response). finding
  • A two-step sequential ComBat batch-correction pipeline integrates multi-study, multi-platform gene expression data for meta-analysis. method
  • Reverse-engineering an ARACNe gene regulatory network and overlaying model-selected genes reveals shared predictive functional space across technologies. method
  • Type-I interferon-inducible genes (LY6E, IFI27, IFI44) are among the most consistently selected discriminative features across all search procedures. finding
  • Out-of-sample cross-platform validation (training/testing models on the other technology's dataset) demonstrates model generalizability and lack of overfitting. method
  • Both backward elimination and a genetic algorithm (GALGO) feature selection with a Random Forest classifier yield predictive infection-status models. method
Experimental setups
Assay System Perturbation Readout Platform
gene expression microarray (Affymetrix) human whole blood, infection studies (10 studies, GPL570/GPL571/GPL9188) none (observational: bacterial/viral/control infection states) gene expression intensity for class prediction Affymetrix microarray (GPL570, GPL571, GPL9188)
gene expression microarray (Illumina) human whole blood, infection studies (GPL10558) none (observational: bacterial/viral/control infection states) gene expression intensity for class prediction Illumina microarray (GPL10558)
feature selection / classification (Backward Elimination + Random Forest) merged Affy_I (1676 samples) and Illumina_I (1892 samples) datasets none OOB error, balanced accuracy, sensitivity, specificity, selected gene panel R VarSelRF, Ranger packages
feature selection / classification (Genetic Algorithm, GALGO + Random Forest) merged Affy_I and Illumina_I datasets none selected gene chromosomes, model accuracy R GALGO, Ranger packages
gene regulatory network inference and clustering merged batch-corrected gene expression dataset none interaction network, functional sub-network clusters ARACNe, Cytoscape (GLay/Girvan–Newman), DAVID
out-of-sample cross-platform validation (Random Forest) Affymetrix-optimized genes tested on Illumina and vice versa none classification error on non-discovery data R Ranger, 60/40 train/test split
Key results
  • Best model correctly predicted bacterial samples 93%
  • Best model correctly predicted viral samples 89%
  • LY6E was present among the nine most frequently selected genes in all four search procedures (Affy BW/GA, Illumina BW/GA)
  • Affymetrix backward elimination converged to a stable gene set 14 genes at relative frequency 1.0
  • Illumina backward elimination converged to a stable gene set 12 genes at relative frequency 1.0
  • Functional enrichment of 88 intersecting genes between manufacturer models highlighted immune terms (Antiviral defense, type I interferon signalling, Immunity) 12 / 10 / 17 genes
  • Merged datasets retained distinct genes after integration (Illumina vs Affymetrix) 19,947 vs 13,383 genes
Key statistics
  • count 3868 samples from 21 studies (total samples included across Affymetrix and Illumina platforms)
  • count Affy_I: 1676 samples (314 bacterial 18.74%, 1121 viral 66.89%, 241 control 14.38%) (merged batch-corrected Affymetrix modelling dataset)
  • count Illumina_I: 1892 samples (356 bacterial 18.82%, 1069 viral 56.50%, 467 control 24.68%) (merged batch-corrected Illumina modelling dataset)
  • count 88 genes with >5% aggregated inclusion across all search procedures (intersecting genes compared between manufacturers)
  • count 240 BW search procedures per dataset; 250 GA models per dataset (number of feature selection runs)
  • count type I interferon signalling pathway included 10 genes (functional enrichment of 88 intersecting genes)
  • other P-value threshold < 0.05 (significance threshold for ARACNe network interactions)
  • other GA fitness goal 0.95, chromosome size 15 genes (GALGO genetic algorithm parameters)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper performed a meta-analysis of publicly available human blood microarray gene expression data (Affymetrix and Illumina platforms; 21 studies; 3568 samples post-processing) to build multi-class Random Forest classifiers distinguishing bacterial infection, viral infection, and no-infection states. Batch effects across studies and platforms were corrected using a two-step sequential ComBat pipeline validated by PCA and overlap of differential expression results. Feature selection employed Backward Elimination (240 runs, OOB-minimization, 60/20/20 split) and a Genetic Algorithm (GALGO, 250 models, k-fold cross-validation), with model performance summarised as balanced accuracy, sensitivity, specificity, and McNemar's test p-value on a held-out split; cross-platform out-of-sample validation served as the primary generalisability check.

Replicationbiological Sample sizeSample counts per platform and class reported in tables (1676 Affymetrix, 1892 Illumina); no a priori power calculation described GroupsBacterial infection vs. viral infection vs. no infection (control); secondary cross-platform validation (Affymetrix model tested on Illumina data and vice versa) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Random Forest classifier (R Ranger) with out-of-bag (OOB) error minimisation Backbone classifier for both Backward Elimination and Genetic Algorithm feature selection across Affymetrix and Illumina datasets 1676 (Affy_I) and 1892 (Illumina_I) not stated
McNemar's test Testing consistency in model responses and potential bias toward classifying a certain class for each representative model (Affy_BW, Affy_GA, Illumina_BW, Illumina_GA) held-out evaluation split size not stated in main text not stated
ARACNe mutual information with unadjusted P-value threshold < 0.05 Selecting significant edges in reverse-engineered gene regulatory networks from batch-corrected expression data not stated per network build not stated
k-fold cross-validation (within GALGO genetic algorithm) Counter-overfitting guard during genetic algorithm feature selection not stated (k not specified in main text) not stated
Principal component analysis (PCA) Visual assessment of batch correction success pre- and post-ComBat 1676 (Affy_I) and 1892 (Illumina_I) na
DAVID functional enrichment analysis Mapping immune-response ontology terms onto network sub-clusters and onto the 88 genes with >5% aggregated inclusion frequency 88 genes (aggregated across four feature-selection runs) not stated
Approaches that could also have been used
  • Batch effects across 21 studies and two platforms were corrected using a two-step sequential ComBat approach with study ID and class as covariates
    Could also: Surrogate Variable Analysis (SVA) or limma's removeBatchEffect, which estimate latent confounders without requiring fully known batch structure — When batch labels are partially uncertain or confounded with biology, SVA can capture residual variation that a label-based approach may miss; comparing multiple correction strategies also allows sensitivity analysis of downstream results
  • Multi-class classification relied on Random Forest as the sole underlying classifier for both feature selection strategies
    Could also: L1-regularised (LASSO) or elastic-net multinomial logistic regression, or gradient-boosted trees (e.g., XGBoost) — Regularised regression yields explicit coefficient estimates and built-in sparsity directly interpretable as feature weights; evaluating more than one classifier family can clarify whether selected biomarkers are robust across model assumptions
  • Held-out model performance was estimated from a single 60/20/20 training/test/evaluation split for Backward Elimination
    Could also: Repeated stratified k-fold cross-validation (e.g., 10 × 5-fold) or nested cross-validation that wraps feature selection inside the outer loop — A single partition can yield higher-variance performance estimates, particularly with class imbalance; nested cross-validation provides less optimistically biased generalisation estimates and is widely recommended when the same data inform both selection and evaluation
  • ARACNe network edge significance was assessed using an unadjusted P-value threshold of 0.05 across all gene-pair mutual information scores
    Could also: False Discovery Rate correction (e.g., Benjamini-Hochberg) applied to the full set of pairwise tests before thresholding — With tens of thousands of gene pairs tested simultaneously, an unadjusted threshold is expected to admit many false-positive edges; FDR correction quantifies and bounds the expected proportion of spurious interactions retained in the network
  • Model performance metrics (balanced accuracy, sensitivity, specificity) were reported as point estimates without measures of uncertainty
    Could also: Bootstrap 95% confidence intervals around each performance metric, or reporting the full per-class confusion matrix with confidence bounds — Point estimates alone do not convey sampling uncertainty; given pronounced class imbalance (>50% viral, ~18% bacterial), confidence intervals allow readers to judge whether differences across models or platforms exceed chance variation
  • Class imbalance (bacterial samples most underrepresented at ~18%) was addressed with a class penalty weight within the Random Forest
    Could also: Synthetic minority oversampling (SMOTE) applied before training, or primary evaluation using macro-averaged F1 or AUROC rather than accuracy — Class penalties adjust the learning objective but not the data distribution; SMOTE augments the minority class directly; AUROC and macro-F1 are threshold-independent and weight classes more evenly, facilitating comparisons across imbalanced multi-class settings
Software: R/Ranger · VarSelRF (R package) · GALGO (R library) · ComBat (R/sva package) · ARACNe · Cytoscape · DAVID

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
5
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE162329 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE162330 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
P40305 UniProt in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
Q16553 UniProt in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
Q8TCB0 UniProt in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33560295

Paper: Myall AC et al. (2021) An OMICs-based meta-analysis to support infection state stratification. Bioinformatics 37. DOI 10.1093/bioinformatics/btab089. PMCID PMC8388022.

Repo (authors' own code): https://github.com/PGB-LIV/Classifying-disease-state-in-high-dimensional-data @ commit 20742a6d3a5650f7bb35d8e6b93a7752e39ea9a4 (2020-11-22). No license file.

Data: GEO GSE162329 (Affymetrix combined) / GSE162330 (Illumina combined). The repo also ships example data: example data/example_raw.RDS (raw, 29 MB, 3 Affy platforms GPL570/571/9188) and example data/example_bc.RDS (batch-corrected, 2.7 MB).

Pipeline (from Methods + repo scripts)

  1. preprocess.R — two-step ComBat batch correction: (a) within-platform over study GSE (mod = ~1 + sample.type, par.prior=T), merge over common genes, then (b) over platform GPL. Deterministic (parametric ComBat, no RNG).
  2. backwardsElimination.R — feature selection by backward elimination using varSelRF2 (a ranger-based reimplementation of varSelRF) + final ranger RF with class weights. 60/20/20 train/test/val split. OOB error as minimization criterion. Repo example config: trees=100, dropFrac=0.1, runs=30 (paper used 240 BW procedures). Stochastic (RF, no fixed seed in repo).
  3. geneticSearchAlgo.R — GALGO genetic-algorithm feature selection. README states the GALGO package is "currently unsupported by R 3.6" → known environment blocker.

In scope (pipeline-derived, attempted)

  • C-data: dataset composition of the shipped Affy example (sample count, class distribution, gene count) vs Table 2 Affy_I (1676 samples; 314 B / 1121 V / 241 C; 13,383 genes). Deterministic, auditable.
  • C-batch: reproduce ComBat batch correction (preprocess.R on example_raw.RDS) and compare to the shipped batch-corrected matrix. ComBat parametric → deterministic anchor.
  • C-clf: run the repo's backward-elimination RF on the example data; compare gene-set size + per-class balanced accuracy / sensitivity / specificity to Table 4 (Affymetrix BW column: gene-set 33; bal.acc B/C/V 0.94/0.78/0.86; sens 0.90/0.57/0.97; spec 0.93/0.96/0.76). Stochastic → provisional / partial expected.

Out of scope / not attempted (the hard 20%)

  • GALGO feature selection (geneticSearchAlgo.R): package unsupported by R 3.6 per the authors' own README → env_unresolvable for that sub-result. Skipped, reason recorded.
  • Full Table 4 with 240 BW search procedures at production tree counts on the FULL GSE162329/GSE162330 matrices (example ships a reduced demo config).
  • Network reverse-engineering / clustering results (Type I IFN pathway convergence, FR clusters, 16-gene intersection) — downstream, external tooling, not in the repo scripts.
  • Illumina arm beyond the shipped Affy example (example data is Affy only).

Possible-fabrication watch

Headline abstract numbers ("93% of bacterial and 89% viral correct") vs Table 4 per-class values to be checked for internal consistency; flag for human review, not asserted.

Figures / tables: Table
C-data
Reported
Affy_I 1676 samples (B314/V1121/C241), 13383 genes (Table 1/2)
Reproduced
shipped classifier demo (example_bc) = 982 samples (B87/C241/V654), 384 genes; Control n=241 matches Table 2 exactly
partial
C-batch
Reported
two-step ComBat batch correction (deterministic pipeline)
Reproduced
preprocess.R ran end-to-end & deterministically (ComBat par.prior=T); but example_raw (3-platform 21-gene intersection) and example_bc (384 genes) have 0 overlap -> not a lineage-matched pair, no value validation possible
partial
C-clf-size
Reported
33-gene optimal model (Table 4 Affy backward-elim)
Reproduced
mean 19.2 / median 17 / best 21 genes
partial
C-clf-sens
Reported
sensitivity B0.90 / C0.57 / V0.97 (Table 4 Affy BW)
Reproduced
B0.413 / C0.556 / V0.952 (mean over 30 runs, validation set)
partial
C-clf-spec
Reported
specificity B0.93 / C0.96 / V0.76 (Table 4 Affy BW)
Reproduced
B0.979 / C0.942 / V0.616
partial
C-clf-bacc
Reported
balanced accuracy B0.94 / C0.78 / V0.86 (Table 4 Affy BW)
Reproduced
B0.696 / C0.749 / V0.784
partial
C-galgo
Reported
GALGO genetic-algorithm models (Table 4 GA columns, gene-set 36/37)
Reproduced
NOT ATTEMPTED
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 44/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🔴6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

Running the authors' own R pipeline on the shipped demo (example_bc: 982/1676 samples, 384/13383 genes, bacterial down-sampled to ~9%) reproduces the Control class almost exactly (sens 0.556≈0.57, spec 0.942≈0.96, bal.acc 0.749≈0.78, n=241 exact) and Viral sensitivity (0.952≈0.97), with no fabrication signal. The large deviations — Bacterial sensitivity 0.90→0.413, bal.acc 0.94→0.696, and gene-set 33→21 — sit on the input/data-definition side and are explained by the demo down-sizing plus a lighter config, i.e. our method choice, not an authors' defect; Table 4 is plausibly derivable from the full public GSE cohort that was simply not shipped. The GALGO arm is unreproduced (env-unresolvable). Net: solid, explainable partial reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

182.6 k
tokens (I/O) · 13.9 M incl. cache
18 min
runtime · 0.12 CPU-h
3.6 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine