Improving recombinant protein production by yeast through genome-scale modeling using proteome constraints.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH + 1:1 REPRODUCED. pcSecYeast is a proteome-constrained genome-scale secretory model of S. cerevisiae (Yeast8-based; MATLAB+COBRA+RAVEN+SoPlex). «our HPC» has no MATLAB license, so the full from-scratch simulation could not be re-run; instead the pipeline-derived results were reproduced 1:1 in Python from the shipped model (pcSecYeast.mat, enzymedata*.mat, TableS1/TargetProtein.xlsx) and the Zenodo-deposited SoPlex/FSEOF outputs (a valid reproduction path). Of 11 claims: 4 EXACT (C1 gene split 1639=1156+483; C5 amylase 116 targets, with the FSEOF selection independently recomputed; C6 41 shared targets at the high-confidence threshold; C10 8 proteins), 4 WITHIN-TOL (C3 65% vs ~70% proteome mass; C7 avg 117.9/77% sec vs 117/80%; C8 3.75x vs 3.5x cost; C9 Pearson -0.275..-0.308/P<1e-8 vs <-0.27/P<1e-8), 2 PARTIAL (C2 166 modeled vs ~200 literature secretory proteins; C4 72-template Methods count not a shipped object), 0 MISMATCH, and 1 NOT-ATTEMPTED (C11 basic-GEM 30x over-prediction, needs MATLAB re-run). One minor internal inconsistency flagged: the paper's amylase split 28/88 vs the authors' own deposited table 29/87 (total 116 exact) - not fabrication. NOT attempted: from-scratch MATLAB re-simulation (no license) and the FigS6 basic-GEM control. Both datasets (Zenodo + GitHub) deliver what they promise, grade A.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 84assessed: 2026-06-21 ⛓ 680940cbdc37
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetA proteome-constrained genome-scale model of the yeast protein secretory pathway (pcSecYeast) can simulate and explain phenotypes caused by limited secretory capacity, and can be used to rationally predict engineering targets that improve recombinant protein production.
- ★ pcSecYeast, a proteome-constrained genome-scale model integrating metabolism, translation, and detailed secretory pathway processing (translocation, PTMs, folding, misfolding, degradation), was constructed for S. cerevisiae resource
- ★ pcSecYeast correctly predicts the experimentally observed switch from high-affinity (Hxt7) to low-affinity (Hxt1/Hxt3) glucose transporters at high glucose concentrations based on secretory cost differences finding
- ★ Yeast cells suppress expression of native secretory proteins with high unit secretory cost when the secretory pathway is under pressure from recombinant protein (α-amylase) production finding
- ★ The degree of suppression of costly native secretory proteins scales with the level of recombinant protein production stress finding
- ★ pcSecYeast can predict overexpression targets for recombinant protein production, validated experimentally for α-amylase method
- Misfolded protein accumulation reduces maximum growth rate finding
- Accounting for full secretory-pathway costs (translation, modification, machinery) gives substantially higher cost estimates than simple direct-cost calculations finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| genome-scale metabolic/proteome-constrained flux simulation | S. cerevisiae (in silico, pcSecYeast model) | varying extracellular glucose concentration | glucose uptake rate, ethanol production, specific growth rate, glucose transporter usage (Hxt1/Hxt3/Hxt7) | — |
| unit secretory cost calculation / correlation analysis | S. cerevisiae, 497 native secretory and cell membrane proteins (in silico) | none (computational cost modeling) | unit secretory cost vs direct cost per protein | — |
| mRNA expression correlation analysis | S. cerevisiae strains MH34, B184, AAC (recombinant α-amylase producers) | differing levels of recombinant α-amylase production | correlation of unit secretory cost with mRNA levels of 497 native secretory proteins | — |
| kcat sensitivity analysis | S. cerevisiae (in silico, pcSecYeast model) | varying Hxt1 kcat value | predicted glucose transporter usage at maximum specific growth rate | — |
| gene overexpression / recombinant protein production validation | S. cerevisiae, α-amylase producing strains | overexpression of model-predicted target genes | recombinant α-amylase production level | — |
- – Model correctly predicted the switch from high-affinity transporter Hxt7 to low-affinity transporters Hxt1/Hxt3 at high glucose concentrations, matching experimental gene expression data
- ▲ Unit secretory cost calculated by pcSecYeast is overall higher than simple direct cost calculations 3.5-fold
- ▼ Significant negative correlation between unit secretory cost and mRNA level of native secretory proteins in α-amylase-producing strains Pearson r < -0.27, P < 1e-8
- ▼ Negative correlation is stronger in strains with higher α-amylase production (MH34, B184) than in the lower-producing strain (AAC) P = 0.004
- – pcSecYeast accounts for 1639 protein-coding genes and approximately 70% of total proteome mass 1639 genes; ~70% proteome mass
- correlation Pearson correlation coefficient < -0.27 (unit secretory cost vs mRNA level of native secretory proteins in α-amylase production strains)
- pvalue P < 1e-8 (significance of negative correlation between unit secretory cost and mRNA level)
- pvalue P = 0.004 (comparison of correlation strength between higher-producing (MH34, B184) and lower-producing (AAC) α-amylase strains)
- fold_change 3.5-fold (unit secretory cost overall higher than direct cost across native secretory proteins)
- count 1639 protein-coding genes (genes included in pcSecYeast (1156 metabolic, 483 protein synthesis/secretion-related))
- count 497 native secretory and cell membrane proteins (proteins analyzed for secretory cost calculations)
- other ~70% of total proteome mass (45.7% metabolic, 20.6% ribosome/proteasome/secretory machinery, 4.6% unmodeled secretory) (proteome coverage of pcSecYeast per PaxDb)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is primarily a computational genome-scale modeling paper (pcSecYeast) that simulates protein secretory pathway resource allocation and validates model predictions largely through qualitative agreement with previously published experimental observations (e.g., glucose transporter switching). Quantitative statistical support in the excerpted text is limited to Pearson correlation coefficients (with associated P values) relating model-predicted unit secretory costs to measured mRNA levels of native secretory proteins across three previously characterized α-amylase-producing yeast strains, plus a single-parameter sensitivity analysis. Dedicated inferential statistics for the later experimental validation of predicted overexpression targets (e.g., α-amylase titer measurements) are not detailed in this excerpt.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson correlation | Correlation between model-predicted unit secretory cost and measured mRNA level of native secretory proteins, in α-amylase-producing strain MH34 and across all three strains (Supplementary Fig. 2b, c) | 497 native secretory proteins | not stated |
| Comparison of correlation strength between strains (method not specified) | Comparing the negative correlation strength in higher-producing strains (MH34, B184) versus a lower-producing strain (AAC) | — | not stated |
| One-parameter sensitivity analysis | Varying the kcat value for the Hxt1 transporter in the model to test robustness of the predicted transporter preference | — | na |
-
A Pearson correlation coefficient was used to relate model-predicted unit secretory cost to measured mRNA expression levels of native secretory proteins.↳ Could also: Spearman's rank correlation — Spearman's correlation could also be applied if the cost-expression relationship is monotonic but not strictly linear, or if the distribution is influenced by extreme values — the text itself notes outliers such as Rax2, Tor1, and Tor2 — since rank-based correlation is less sensitive to such outliers and non-linearity.
-
A single P value (P = 0.004) was reported to indicate that the negative correlation was stronger in higher-producing strains than in the lower-producing strain.↳ Could also: A formal statistical comparison of correlation coefficients, such as Fisher's r-to-z transformation with a confidence interval for the difference — This would provide both a quantitative effect size and an interval estimate for the difference in correlation strength between strains, complementing the single reported P value.
-
Correlational and sensitivity analyses were carried out across many proteins and multiple strains without a stated adjustment for multiple comparisons.↳ Could also: A multiple-testing correction such as Benjamini-Hochberg false discovery rate control — This could help manage the overall false-positive rate when many simultaneous or related statistical comparisons (across proteins/strains) are being interpreted together.
-
Model predictions, such as the switch between glucose transporters and secretory cost trends, were validated mainly by qualitative agreement with previously published experimental patterns rather than by a dedicated new statistical test.↳ Could also: Quantitative model-fit metrics, such as R² or RMSE between predicted and observed values, or bootstrap confidence intervals on model outputs — These would provide a numerical measure of how closely the simulated outputs track the experimental data, complementing the qualitative/directional comparison already presented.
-
Sensitivity of the model's transporter-switching prediction was assessed by varying a single kinetic parameter (kcat for Hxt1).↳ Could also: A global or Monte Carlo sensitivity analysis varying multiple kinetic parameters simultaneously — This would characterize how robust the qualitative conclusion is to joint uncertainty across several parameters at once, rather than one parameter in isolation.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35624178
Paper: Li F. et al. (2022) Improving recombinant protein production by yeast through genome-scale modeling using proteome constraints. Nat Commun 13:2969. DOI 10.1038/s41467-022-30689-7 · PMCID PMC9142503.
Code: https://github.com/SysBioChalmers/pcSecYeast (MATLAB, GPL-3.0,
latest commit 9bce335ff6ede75fae67b34e22adc0ee27ad5bcd, pushed 2023-11-03).
Data: Zenodo 10.5281/zenodo.6320643 → Results.zip (395,524,466 bytes,
intermediate simulation results). Model file pcSecYeast.mat (1.6 MB) shipped
inside the repo (ComplementaryScripts/pcSecYeast.mat).
What the pipeline is
pcSecYeast is a proteome-constrained genome-scale metabolic model (pcGEM) of
S. cerevisiae built on Yeast8.3.5, expanded with translation, post-translational
modification and the secretory pathway. The computational pipeline:
buildModel.m— reconstructs the model from collected data (COBRA + RAVEN).ComplementaryScripts/Simulation/*.m— LP/MILP optimisations (SoPlex solver) producing per-protein secretory costs and overexpression-target predictions.- Two Python scripts (
Fig4c_*,FigS7_*) — SHAP feature-importance ML on the simulation outputs.
Required stack: MATLAB (≥7.3) + COBRA toolbox + RAVEN toolbox + SoPlex solver.
In scope (pipeline-derived, attempt to reproduce)
| id | reported result | pipeline | tier |
|---|---|---|---|
| C1 | model genes: 1639 protein-coding (1156 metabolic + 483 synth/secretion) | model structure | T1 (no MATLAB) |
| C2 | ~200 proteins engaged in secretion | model structure | T1 |
| C3 | model covers ~70% of total proteome mass | model + abundance | T1/T2 |
| C4 | 72 template reactions for protein synthesis | model structure | T1 |
| C5 | α-amylase: 116 predicted targets (28 metabolic + 88 secretory) | SimulateTP/Fig5 | T1 (deposited results) / T2 (re-run) |
| C6 | 41 targets shared by all 8 recombinant proteins | Fig5c | T1/T2 |
| C7 | avg 117 targets per protein; ~80% secretory / 20% metabolic | Fig5 | T1/T2 |
| C8 | unit secretory cost 3.5× higher than direct cost | SimulateProteinCost / FigS2a | T2 |
| C9 | unit secretory cost vs mRNA: Pearson < −0.27, P<1e-8 | FigS2 | T2 |
| C10 | 8 recombinant proteins simulated | SimulateTP | T1/T2 |
| C11 | basic GEM over-predicts α-amylase ~30× vs experiment | FigS6 | T2 |
- T1 = reproducible WITHOUT MATLAB: load the shipped
.mat/.xlsx/Zenodo result tables in Python on «infra» and recompute counts/structure. - T2 = needs the full MATLAB+SoPlex re-run (the harder "keep going" target); feasible only if «our HPC» provides a MATLAB module + SoPlex.
Out of scope (not pipeline-derived)
- Wet-lab validation fold-increases (CYS4 2.14×, ERO1/ERV2 etc., 9/14 secretory & 1/4 metabolic targets validated) — experimental, manual strain construction.
- Transcriptome / proteome measurement generation (input data, not pipeline output).
- α-amylase cysteine-enrichment (sequence statistic, trivially checkable but a property of the protein sequence, not a model output) — will spot-check if easy.
Reproduction plan
- T1 first (quick floor): on «infra», load
pcSecYeast.matwith scipy → countrxns/mets/genes; parse the deposited target tables (Results.zip + Supplementary Data) → verify C5/C6/C7 counts. No MATLAB needed. - T2 (keep going): check
module avail matlab+ SoPlex on «our HPC»; if present, build COBRA/RAVEN env, runSimulateProteinCost.m/SimulateTP.mfor ≥1 protein and compare costs/targets 1:1.
Blockers seen so far
- «our HPC» VPN tunnel down at start (ssh timeout to «host») — all «infra» downloads
- compute wait on the central fix. No data may be pulled onto «host».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.