Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Detecting tipping points of complex diseases by network information entropy.

Brief Bioinform · 2024
L1 78/100 PQI 87
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce; faithful 1:1 on the flagship influenza result. Ran the authors' own code (NIEE.py @ 9a5acfb) unmodified on the paper's own data (GEO GSE30550) at the documented StringDB threshold 0.85, all compute on «our HPC» SLURM. C3 (study design + parameter) reproduced exactly. C2 (key local network) is the strongest data point: all nine reported biomarker genes (CCL8, CCR1, CCR3, CCR5, CXCL10, DDX58, RNF125, USP25, XCR1) are recovered, CXCL10 is the dominant hub as emphasized in the paper, and the top key genes are the expected interferon-stimulated/chemokine antiviral set; only the literal '13 key edges' count is threshold-dependent (21 at >=7/9 subjects). C1 (early-warning, Fig 2) reproduced qualitatively: every symptomatic subject shows a sharp NIEE early-warning peak, stronger than asymptomatic subjects (5.79x vs 4.43x). NOT attempted (80/20): the other datasets (TCGA cancers, single-cell), Supplementary Notes S6-S8 method comparisons, GO enrichment, and the exact per-subject clinical onset hours needed to strictly verify '8 of 9 before symptoms'. No fabrication signal: the reproduction independently supports the reported results.

💻 Code ↗ 🗄 Data: GSE30550

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 78
    assessed: 2026-06-14 ⛓ ab314aeb24d4
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether an edge-based network information entropy measure (NIEE), built on dynamic network biomarker theory and sample-specific networks, can detect critical states (tipping points) in complex disease progression from bulk or even single-sample expression data, enabling early-warning prediction before disease onset.

Core claims
  • NIEE can detect critical states or tipping points in diverse data types, including bulk and single-sample expression data method
  • Applying NIEE to real disease datasets successfully identified critical predisease stages and tipping points before disease onset finding
  • NIEE effectively predicted upcoming diseases and showed early-warning signals when assessed on influenza, acute lung injury, cancer, and pancreatic disease datasets finding
  • Enrichment analysis of the key local network identified by NIEE revealed significant associations between biological functions and diseases in the experimental samples finding
  • NIEE is a model-free algorithm requiring only a suitable background network, without any prior training on extensive data method
  • NIEE calculates network information entropy of local edges by considering the interdependent regulation (first- and second-order neighbors) among genes mechanism
  • Sample-specific networks (SSN) construct a unique correlation network for each individual sample based on single-sample differential association information resource
  • The background network is constructed from the StringDB database using a confidence score threshold of ≥0.85, excluding isolated nodes method
Experimental setups
Assay System Perturbation Readout Platform
network information entropy analysis (NIEE) of gene expression data real disease datasets (bulk and single-sample expression data) disease progression / disease state (influenza, acute lung injury, cancer, pancreatic disease) local and global NIEE scores indicating critical transitions/tipping points and early-warning signals
background/reference network construction via protein-protein interaction database filtering StringDB database none background network edges retained above confidence score ≥0.85, isolated nodes excluded StringDB (https://www.string-db.org)
sample-specific network (SSN) construction from correlation of gene expression individual patient/sample expression data comparison of perturbed sample vs. reference/healthy control samples sample-specific correlation network used to derive local/global NIEE scores
enrichment analysis on key local network genes/edges identified via NIEE landscape analysis and threshold screening none biological functions and disease associations enriched among filtered edges/genes
Key results
  • NIEE successfully identified critical predisease stages and tipping points before disease onset in real disease datasets
  • NIEE effectively predicted upcoming diseases and showed early-warning signals across influenza, acute lung injury, cancer, and pancreatic disease datasets
  • Enrichment analysis of the NIEE-identified key local network showed significant associations between biological functions and diseases in the experimental samples
Key statistics
  • other confidence score ≥ 0.85 (threshold used to filter StringDB background network edges)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational methods paper proposing NIEE (Network Information Entropy of Edges), a model-free algorithm that combines Dynamic Network Biomarkers (DNB), Sample-Specific Networks (SSN), and Shannon-type information entropy to detect tipping points in complex disease progression. The core computation derives per-edge entropy and standard deviation scores by calculating Pearson Correlation Coefficients between each gene and its first-order network neighbors (weighted by node degree), then compares individual perturbed samples against a reference set of healthy samples. Results are interpreted via landscape analysis and threshold-based filtering to yield key local gene networks, with downstream enrichment analysis applied to those modules. The available text describes a score-based, algorithmic output rather than classical inferential hypothesis testing with p-values.

Replicationunclear Sample sizeNot stated in available text; multiple real-world disease datasets used; reference group defined as healthy/normal-tissue samples; perturbed samples are individual disease-progression samples evaluated one at a time GroupsReference (healthy/normal) samples vs. single perturbed (disease-stage) samples, evaluated in a single-sample framework Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Pearson Correlation Coefficient (PCC) — used as a correlation measure for network construction, not as a significance test Calculation of gene-gene expression correlations across reference samples to build sample-specific networks (Equations 4–5) s reference samples per edge; exact n varies by dataset and is not stated in the available text not stated
Shannon-type normalized network information entropy (H_E) Per-node entropy computed over first-order neighbor correlations; combined into per-edge NIEE scores (Equations 1–3); applied across all edges in the StringDB background network s reference samples per calculation not stated
Standard deviation (SD_E) — used as a fluctuation/variability measure Per-node and per-edge SD score computation (Equations 2 and 6); combined with entropy to characterize pre-disease state volatility s reference samples not stated
Gene enrichment analysis (specific method not named in available text) Downstream characterization of key local networks identified by NIEE across multiple disease datasets (influenza, acute lung injury, cancers, pancreatic diseases) not stated
Approaches that could also have been used
  • Gene-gene correlations were computed using Pearson Correlation Coefficient (PCC) to construct sample-specific networks
    Could also: Spearman rank correlation could also be used for the same network construction step — Spearman correlation is less sensitive to outliers and does not assume linearity, which may be relevant for RNA-seq or microarray data with skewed expression distributions or non-linear co-regulation patterns
  • Node degree was used as the weighting factor when combining per-node entropy scores into a per-edge NIEE score (Equations 1–2)
    Could also: Other centrality measures — such as betweenness centrality or eigenvector centrality — could also serve as node weights — Betweenness or eigenvector centrality captures a node's global network influence rather than local connectivity alone, which may up-weight biologically important hub genes that are not necessarily high-degree in the local neighborhood
  • The background PPI network was sourced from StringDB with a single fixed confidence threshold of ≥ 0.85
    Could also: Alternative databases (e.g., BioGRID, HPRD, or a consensus multi-database network) or a sensitivity analysis across a range of confidence thresholds could also be used — Different databases reflect different evidence types and species coverage; exploring threshold sensitivity would help characterize how network density and edge selection affect NIEE score distributions and tipping-point detection stability
  • Critical-state detection is framed as a score-based landscape analysis with empirical threshold setting, without a formal null-hypothesis test for the significance of an observed NIEE spike
    Could also: Permutation testing or bootstrapping over sample labels could also be applied to establish a null distribution for NIEE scores and derive empirical p-values or confidence bands — A permutation-derived null would allow researchers to quantify the probability of observing a given NIEE elevation by chance, enabling more formal statistical claims about tipping-point timing, especially in small-n datasets
  • NIEE performance was demonstrated on multiple real disease datasets without a stated cross-validation scheme or held-out test set
    Could also: Leave-one-out cross-validation or evaluation on prospective independent cohorts could also be applied to assess generalizability of tipping-point detection — Held-out or cross-validated evaluation helps distinguish whether the identified tipping-point signals reflect dataset-specific patterns or genuinely generalizable disease dynamics transferable to new patient populations
  • The enrichment analysis method for identified key local networks is mentioned in the abstract but not specified in the available text
    Could also: Gene Set Enrichment Analysis (GSEA) could also complement over-representation analysis (ORA) for pathway-level characterization of identified gene modules — GSEA operates on ranked gene lists rather than a binary membership threshold, reducing sensitivity to the specific edge-filtering cutoff used in the NIEE landscape step and potentially surfacing weaker but coordinated pathway-level signals
Software: StringDB (protein-protein interaction background network database)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
12
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

1ESR PDBe in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
1O80 PDBe in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
8DVU PDBe in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38960408 (NIEE: Detecting tipping points by Network Information Entropy)

  • Paper: Lyu, Chen, Liu. Detecting Tipping Points of Complex Diseases by Network Information Entropy. Brief Bioinform 2024; bbae311. PMID 38960408 / PMC11221888.
  • Code: https://github.com/lllvcs/NIEE (authors' own code — P16 own-repo).
  • Data: GEO GSE30550 — H3N2 influenza challenge, 17 subjects, 16 timepoints (−24,0,5,12,21,29,36,45,53,60,69,77,84,93,101,108 h), 9 symptomatic / 8 asymptomatic.

Method (from NIEE.py + notebooks)

NIEE = Network Information Entropy of Edges. Single-sample / SSN style:

  1. Background network = StringDB edges with combined_score ≥ sl (paper: 0.85 → 850).
  2. Reference (Baseline) sample group → gene SD + |Pearson| correlation, masked by network. Per network edge (i,j): edge entropy = degree-weighted mean of node neighbor-prob Shannon entropies; edge sd = degree-weighted mean of node SDs.
  3. For each perturbed (non-baseline) sample: append to reference, recompute, take |Δentropy|·|Δsd| per edge → "NIEE landscape" value per edge per sample (NIEE.py).
  4. Per-sample NIEE score = sum of the top edges (top-100 / top-1%); the timepoint where it peaks = tipping point / early-warning signal (Result_analysis.ipynb).

IN SCOPE (pipeline-derived, attempted)

  • C1 — early-warning before symptoms (Fig 2). Run NIEE on GSE30550 symptomatic subjects; per-subject NIEE score curve over the 16 timepoints should show a sharp early-warning peak. Paper: "eight of [the nine symptomatic subjects] had NIEE early-warning signals before clinical symptoms appeared." Grade by reproducing the curves + counting subjects whose peak precedes symptom onset.
  • C2 — key local network / biomarker genes. Top-1% edges across symptomatic subjects → key local network of 13 key edges; reported biomarker genes CCL8, CXCL10, XCR1, CCR1, CCR3, CCR5, DDX58(RIG-I), RNF125, USP25. Grade by overlap of reproduced top-edge gene set with this set.
  • Parameter check: StringDB combined_score ≥ 850.

OUT OF SCOPE (not attempted, with reason)

  • Other datasets in the paper (TCGA cancers, single-cell) — 80/20, focus on the flagship influenza result (Fig 2).
  • Supplementary Notes S6–S8 method comparisons (KEGG background net, MIC, protein- length Weight) — auxiliary robustness checks, not the headline result.
  • GO/pathway functional enrichment of the key network — downstream annotation, not the NIEE pipeline output itself.
  • Exact per-subject symptom-onset hours from the original Huang 2011 clinical scores are only partially recoverable; C1 onset comparison is therefore provisional.

Compute plan

All heavy compute on «our HPC» (SLURM, «infra»). Download GSE30550 series matrix + StringDB v12 + GPL platform annotation inside the compute job; preprocess to NIEE inputs (expr by symbol, ENSP→symbol dict, anno); run NIEE.py at sl=850; aggregate to per-subject score curves + top-edge gene set. Small results → «host» dataset folder.

Figures / tables: Fig 2
C1
Reported
All 9 symptomatic subjects had NIEE early-warning signals; 8 of them before clinical symptoms (Fig 2, GSE30550)
Reproduced
9/9 symptomatic subjects show a clear NIEE early-warning peak (peak hours 29-77h); symptomatic mean peak = 5.79x median vs asymptomatic 4.43x; Fig 2 structure reproduced
partial
C2
Reported
Key local network (~13 key edges); biomarker genes CCL8, CXCL10, XCR1, CCR1, CCR3, CCR5, DDX58, RNF125, USP25
Reproduced
All 9 reported biomarker genes recovered in the key network; CXCL10 dominant hub; top genes STAT1/CXCL10/CCL2/RSAD2/IFIT1 (ISG+chemokine antiviral set); 21 recurrent key edges at >=7/9 subjects
within tolerance
C3
Reported
17 subjects, 16 time points, 9 symptomatic / 8 asymptomatic; StringDB confidence threshold 0.85
Reproduced
17 subjects, 268 samples (16 tp), 9 symptomatic / 8 asymptomatic; NIEE run at combined_score>=850
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

This is a strong, faithful reproduction: the authors' own NIEE code run on the paper's own public GEO data (GSE30550) at the documented StringDB 0.85 threshold recovers all nine reported biomarker genes with CXCL10 as the dominant hub, reproduces the 17-subject/9–8 design exactly, and shows the early-warning peak in 9/9 symptomatic subjects (5.79× vs 4.43×). The deviations are minor and lie on our side / paper under-specification — the '13 key edges' count is threshold-dependent (21 at ≥7/9) and the precise 'eight of nine before clinical onset' was not strictly verified. No fabrication signal; the reproduction independently supports the central claim, so overall quality is solid with only explainable, definitional deviations.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

206.9 k
tokens (I/O) · 14.6 M incl. cache
36 min
runtime · 0.37 CPU-h
5.3 GB
peak RAM
2
HPC jobs
hummel
machine