Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Minimal metabolic pathway structure is consistent with associated biomolecular interactions.

Mol Syst Biol · 2014
L1 30/100 3/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • No authors-side cause for any deviation
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
30/100
Reproducibility score
2.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 2% of all assessed papers rank 1148 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Exact reproduction of paper's headline E. coli null-space dim (750) and toy-model checks. Full MILP not achieved: E. coli is compute-duration-bound (~43.7h serial vs 12h ceiling), yeast blocked by ATPM + model drift (431 vs 332). GSE48324 fully profiled (14/14, 24/24).

💻 Code ↗ 🗄 Data: GSE48324

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-01
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-01
no human curator yet
Last updated
2026-08-01

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Do currently used, human-defined pathway structures correctly account for the observed biomolecular interactions between the macromolecules that carry out metabolic function? The authors test whether an unbiased, parsimony-based (minimal, linearly independent) pathway decomposition of genome-scale metabolic networks describes independent biomolecular interaction data better than canonical textbook/database pathways.

Core claims
  • MinSpan, a mixed-integer linear optimization algorithm, computes the shortest, linearly independent pathways (sparsest basis of the null space of the stoichiometric matrix S) for genome-scale metabolic networks, which convex approaches (extreme pathways, elementary flux modes) cannot do at genome scale. method
  • MinSpan pathways are biologically supported by independent genome-scale datasets: protein-protein interactions, positive genetic interactions, and transcriptional regulation. finding
  • MinSpan pathways are statistically more consistent with positive genetic interactions and with local and intermediate transcriptional regulation than human-defined pathway databases (KEGG, BioCyc/EcoCyc/YeastCyc, Gene Ontology). finding
  • A minimal (sparsest) null space basis is more biologically relevant than alternative bases (least-sparse MaxSpan and random RandSpan bases), confirming parsimony as the underlying organizing principle. finding
  • Local regulation (TFs regulating at most 30 metabolic genes, ~90% of E. coli TFs regulating metabolism) is pathway based, acting directly on minimal, linearly independent pathways, whereas global regulation (>30 regulated metabolic genes) does not mimic the metabolic scaffold. mechanism
  • MinSpan pathways predicted and guided experimental discovery of novel transcriptional regulatory interactions in E. coli metabolism for three transcription factors, effectively doubling the known regulatory roles for Nac and MntR. finding
  • MinSpan pathways are significantly different from (depleted in similarity to) human-defined pathway databases; differences arise because MinSpan enforces steady-state flux balance, so pathways include all required cofactors and precursors (notably for fatty acid metabolism), and some traditional pathways are split into smaller MinSpan pathways. finding
  • Genome-scale MinSpan pathway sets for E. coli and S. cerevisiae are provided as a resource (Supplementary Dataset S1). resource
Experimental setups
Assay System Perturbation Readout Platform
MinSpan mixed-integer linear optimization (sparsest null-space basis of stoichiometric matrix S) — constraint-based modeling Simplified illustrative glycolysis/TCA model (14 metabolites, 18 reactions, rank 14, 4-dimensional null space) and E. coli core metabolism none Set of shortest linearly independent pathway vectors (pathway matrix P)
MinSpan pathway computation at genome scale Genome-scale metabolic network reconstructions of Escherichia coli (Orth et al, 2011) and Saccharomyces cerevisiae (Mo et al, 2009) none Number and composition of minimal pathways (null space dimension); gene/protein sets via gene-protein-reaction (GPR) associations
Spearman correlation analysis of protein co-occurrence/co-absence across pathway protein sets, scored by ROC AUC vs. yeast two-hybrid protein-protein interaction network S. cerevisiae metabolism; PPI data from BioGRID (Stark et al, 2006) none Accuracy (AUC of ROC) and coverage (number of interactions predicted); fraction of known Y2H metabolic PPIs recovered
Spearman correlation analysis of gene co-occurrence/co-absence vs. positive genetic interaction network S. cerevisiae metabolism; genetic interaction data from Costanzo et al, 2010 gene deletion-based genetic interaction data (P < 0.05, ε > 0.16) ROC AUC accuracy and coverage for predicting positive genetic interactions
Spearman correlation analysis of gene sets vs. transcriptional regulatory network (co-regulation by the same transcription factor) E. coli metabolism; TRN from RegulonDB (Gama-Castro et al, 2011) none ROC AUC accuracy/coverage for co-regulated gene pairs; stratified by local (≤30 metabolic genes), intermediate, and global (>30) regulators
Comparative benchmarking against alternative null-space bases and human-defined pathway databases (two-tailed t-test on accuracy) E. coli and S. cerevisiae metabolic networks; MaxSpan (least sparse basis), RandSpan (n = 100 random bases), KEGG Modules, EcoCyc/YeastCyc, Gene Ontology Biological Process none ROC AUC accuracy and coverage; P-values of pairwise comparisons
Pairwise connection specificity index (CSI) computation with hierarchical clustering All pathway definitions (MinSpan, KEGG, BioCyc, Gene Ontology) for E. coli and S. cerevisiae, based on pathway gene products none Pathway-pathway similarity/specificity matrix; percentage of pathways sharing high CSI (top 15% of interactions) across databases
K-nearest neighbor classification of pathways across databases (significance by binomial distribution with Bonferroni correction) E. coli and S. cerevisiae pathway sets from MinSpan, KEGG, BioCyc, Gene Ontology none Number of pathways assigned to each other database; enrichment/depletion of similarity
Key results
  • Genome-scale MinSpan computation yielded 750 pathways for E. coli and 332 for S. cerevisiae, matching the dimensions of the two null spaces. 750 and 332 pathways
  • 80% of known yeast two-hybrid protein-protein interactions in S. cerevisiae metabolism were found within MinSpan pathways by significant correlation of protein co-occurrence/co-absence. 80%
  • MinSpan pathways were representative of positive genetic interactions in S. cerevisiae metabolism based on significant gene correlations across pathway gene sets.
  • E. coli MinSpan pathways share over 6,700 co-regulated gene pairs within metabolism; local regulation is pathway based while global regulation is not captured by any pathway structure. >6,700 co-regulated gene pairs
  • MinSpan outperformed human-defined databases for positive genetic interactions and local/intermediate transcriptional regulation, but for PPIs the advantage was marginal and not statistically significant. PPI: P = 0.0780 vs KEGG, 0.168 vs EcoCyc, 0.901 vs GO
  • MinSpan pathways are significantly dissimilar (depleted) from human-defined pathway databases in both organisms, while KEGG and BioCyc are most similar to each other; MinSpan captures ~88% of E. coli KEGG pathways whereas KEGG captures only ~26% of MinSpan pathways. 88% vs 26%
  • 533 MinSpan pathways are similar to 582 traditional pathways; E. coli has 204 unique MinSpan pathways (54 novel pathways related to ion transport, alternate carbon metabolism, and electron transfer), while S. cerevisiae has none because Gene Ontology captures all model pathways. 533 vs 582; 204 unique; 54 novel
  • Differences with traditional databases stem largely from steady-state flux balance: 26 of 56 pathways missed by MinSpan in E. coli were fatty acid metabolism, and 99 of 204 MinSpan-unique pathways contained the necessary cofactors and precursors, mainly for fatty acid metabolism. 26/56; 99/204
Key statistics
  • count 750 (E. coli) and 332 (S. cerevisiae) MinSpan pathways (Null space dimensions of the two genome-scale metabolic networks)
  • other 80% (Fraction of known yeast two-hybrid metabolic PPIs recovered within S. cerevisiae MinSpan pathways)
  • count over 6,700 co-regulated gene pairs (Shared by E. coli MinSpan pathways within metabolism)
  • pvalue P = 0.0780 vs KEGG; P = 0.168 vs EcoCyc; P = 0.901 vs Gene Ontology (two-tailed t-test) (MinSpan vs other pathway definitions for protein-protein interactions (not significant))
  • pvalue P = 3.16e-3 vs KEGG; P = 0.133 vs YeastCyc; P = 1.47e-3 vs Gene Ontology (MinSpan vs other pathway definitions for positive genetic interactions)
  • pvalue local TRN: P = 2.79e-4 vs KEGG, P = 0.181 vs EcoCyc, P = 2.35e-4 vs GO; intermediate TRN: P = 3.38e-3 vs KEGG, P = 3.97e-6 vs EcoCyc, P = 3.30e-18 vs GO (MinSpan vs other pathway definitions for transcriptional regulation)
  • count Number of pathways — E. coli: KEGG 91, EcoCyc 199, GO 348, MinSpan (filtered) 737; S. cerevisiae: KEGG 74, YeastCyc 121, GO 296, MinSpan (filtered) 298 (Table 1 pathway counts across databases)
  • mean Average pathway length (genes) — E. coli: KEGG 5.4, EcoCyc 6.6, GO 7.3, MinSpan 13.6; S. cerevisiae: KEGG 3.9, YeastCyc 7.1, GO 7.9, MinSpan 13. Average gene usage — E. coli: 1.6, 2.7, 2.4, 8.9; S. cerevisiae: 1.4, 2.2, 6.1, 7.1 (Table 1 average pathway length and gene usage)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper introduces a computational algorithm (MinSpan) for defining metabolic pathways and evaluates it primarily through correlation-based enrichment analyses (Spearman correlations of gene/protein co-occurrence across pathway sets) and ROC/AUC accuracy metrics comparing MinSpan pathways to human-curated pathway databases (KEGG, BioCyc, Gene Ontology) and to alternative computationally generated null-space bases (MaxSpan, RandSpan). Statistical significance between methods was assessed with two-tailed t-tests on accuracy values and, for pathway-similarity classification, a binomial test with Bonferroni correction. Results are reported mainly as exact p-values and descriptive figures (ROC/AUC plots, heatmaps), with dispersion shown for the random-null-space control as mean plus one standard deviation.

Replicationunclear Sample sizeSample sizes are given as network/database sizes (e.g., 750 and 332 MinSpan pathways for E. coli and S. cerevisiae; 100 randomly generated RandSpan null spaces) rather than as a formal power analysis. GroupsMinSpan pathways vs. KEGG, BioCyc/EcoCyc/YeastCyc, Gene Ontology, MaxSpan, and RandSpan pathway definitions, evaluated against PPI, genetic interaction, and transcriptional regulatory network datasets Pairingunclear Randomization/blindingnot stated DispersionSD Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionBonferroni correction
Statistical tests used
Test Applied to n Assumptions
Spearman correlation (co-occurrence/co-absence of genes or proteins across MinSpan pathway sets) Fig 2C-F, comparison of MinSpan pathways to PPI, genetic interaction, and TRN datasets not stated
Two-tailed Student's t-test Fig 2D-F, comparing AUC-based accuracy of MinSpan vs KEGG, EcoCyc/YeastCyc, and Gene Ontology for PPI, genetic interaction, and transcriptional regulation consistency not stated
Binomial distribution test with Bonferroni correction Fig 3C, K-nearest neighbor pathway similarity classification across MinSpan, KEGG, BioCyc, and GO not stated
ROC/AUC analysis (accuracy vs coverage) Fig 2C-F, evaluating how well pathway definitions predict biomolecular interactions na
Approaches that could also have been used
  • MinSpan was compared to each of three pathway databases (KEGG, BioCyc/EcoCyc/YeastCyc, Gene Ontology) separately with two-tailed t-tests, across four interaction-type outcomes, without a stated multiplicity correction for this test family.
    Could also: A one-way ANOVA (or Kruskal-Wallis for non-normal data) across all pathway-definition groups followed by a post-hoc test such as Tukey HSD, or a Benjamini-Hochberg FDR correction applied across the full set of pairwise comparisons — Either approach would control the family-wise or false-discovery error rate across the multiple simultaneous comparisons being made, which can be a useful complement when many pairwise tests are reported together.
  • Pathway biological relevance was quantified using ROC/AUC accuracy against interaction datasets that are typically sparse (few true positive interactions relative to all possible pairs).
    Could also: Precision-recall curves and area under the precision-recall curve — Precision-recall metrics are often considered more informative than ROC/AUC when the positive class (true interactions) is a small fraction of all possible comparisons, and can complement the ROC-based results already shown.
  • The RandSpan control (100 randomly generated null-space bases) was summarized in the figure as mean plus one standard deviation.
    Could also: Reporting an empirical p-value derived directly from the distribution of the 100 random values, or a 95% confidence/percentile interval from that empirical distribution — An empirical percentile-based interval or permutation p-value uses the full simulated distribution directly, which can be a natural complement to a mean ± SD summary, especially if the distribution of random-null-space accuracies is not symmetric.
  • Enrichment of co-occurring genes/proteins across MinSpan pathway sets was assessed via Spearman correlation significance.
    Could also: A hypergeometric or Fisher's exact test for pairwise gene/protein co-membership enrichment, or a permutation-based null model built directly from the network structure — These approaches directly test for over-representation of co-membership rather than relying on correlation significance, and a network-based permutation null can account for structural features of the pathway sets themselves.
  • Significance for pathway-similarity classification (K-nearest neighbor analysis) was assessed with a binomial test plus Bonferroni correction.
    Could also: A Benjamini-Hochberg (FDR) correction in place of Bonferroni — FDR control is generally less conservative than Bonferroni while still accounting for multiple comparisons, which some readers may prefer when many simultaneous tests are involved, though the choice depends on how conservative one wants the correction to be.
  • Differences in AUC-based accuracy between MinSpan and other pathway definitions were tested with t-tests on point-estimate accuracy values.
    Could also: Bootstrap resampling of the underlying interaction data to generate confidence intervals for the AUC differences directly — A bootstrap-based confidence interval on the AUC difference itself would convey both the magnitude and uncertainty of the difference in a single reported quantity, complementing a p-value-only comparison.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

toy_model_linear_algebra_checks
Reported
nnz(S)=309, null shape (84,23) (repo unit-test fixture)
Reproduced
nnz=309, shape (84,23)
m.public.grade.reproduced
ecoli_genomescale_null_space_dimension
Reported
750 raw null-space dim (E. coli iJO1366)
Reproduced
750
m.public.grade.reproduced
ecoli_minspan_full_milp_pathways
Reported
737 filtered MinSpan pathways, avg length 13.6 genes
Reproduced
not completed -- compute-duration bound ~43.7h
m.public.grade.not_run
toy_model_full_minspan_milp
Reported
MILP converges to 0.1% gap for all columns (implicit)
Reproduced
12/23 cols exact at relaxed 5% gap; 0.1% gap not closed on hard cols
partial
yeast_genomescale_null_space_dimension
Reported
332 raw null-space dim (S. cerevisiae iMM904)
Reproduced
431
did not match
yeast_minspan_full_milp_pathways
Reported
298 filtered pathways, avg length 13 genes
Reproduced
not run -- blocked upstream
m.public.grade.not_run
rnaseq_differential_expression_pipeline
Reported
14 samples, GSE48324 Cuffdiff DE analysis
Reproduced
14/14 samples, 24/24 DE files verified; published output used
m.public.grade.not_run

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 30/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

What deviates: only one measured number — the yeast iMM904 null-space dimension came out 431 vs the reported 332 — and that is traceable to a decade of BiGG model drift, since the E. coli branch, fed a bundled period-accurate iJO1366, reproduced 750 exactly. The paper's actual headline outputs (737 E. coli and 298 yeast MinSpan pathways, avg length 13.6 / 13 genes) were never computed: E. coli is compute-duration-bound at 209.8s/col x 750 = ~43.7h against a 12h SLURM ceiling, and yeast is blocked upstream by ATPM's fixed flux violating prepare_model(). Whose side: essentially none of this lands on the authors — code and GSE48324 are public and the dataset profiled clean (14/14 samples, 24/24 DE files); the limits are our compute budget, model-version drift, and a genuine upstream defect (cobra 0.5.4 cglpk.pyx raising GLP_EMIPGAP as an uncaught RuntimeError, which capped the toy MILP at 12/23 columns). Severity: moderate and non-damning — exact agreement wherever the computation finished, no contradicted claim, but the central pathway-enumeration result remains untested, so the core conclusion is only limited rather than fully confirmed. The room's self-imposed-limit checks (catching and correcting timelimit=60 / first_round_timelimit=2 before drawing conclusions) are a point in this reproduction's favour.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.