Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

OSeMOSYS Global, an open-source, open data global electricity system model generator.

Sci Data · 2022
L1 80/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Same input data as the authors
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 56% of all assessed papers rank 484 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (partial, strong) with LIVE compute. OSeMOSYS Global (Sci Data 2022) is an open model GENERATOR (Snakemake -> OSeMOSYS LP -> CBC solve), reproduced end-to-end on «our HPC»/SLURM with the authors' own repo OSeMOSYS/osemosys_global @ v0.4.0 (commit b469377, submodules simplicity@662c96a + OSeMOSYS_GNU_MathProg@6d86dd0; input data bundled in-repo). The scaffold had mis-linked code/data to a DIFFERENT paper (niclasmattsson/Supergrid + zenodo:4730003); corrected to the repo named in the paper's own Code Availability statement. This run re-built everything from scratch (the prior Jun-22 «infra» workdir had been janitor-reclaimed): clone+submodules+conda env+otoole 0.11.0 sdist+model generation+CBC solve, all inside SLURM «job» on a compute node. TABLE 2 FLOOR (C1-C3, deterministic): India 5/165/76 EXACT; BBIN 8 nodes EXACT, tech/comm 278/130 (Δ1 vs 279/131, ≤0.76%); World 265 nodes EXACT, tech/comm 9437/4126 (Δ1 vs 9438/4127, ≤0.02%). The Δ=1 is a single power-plant data-vintage element present in the v0.4.0-shipped global SET files themselves (TECHNOLOGY.csv=9437, FUEL.csv=4126 before any filtering), not a fabrication and not derivable away. Per-scenario SET sizes required reproducing the authors' INTENDED geographic-filter (v0.4.0 has a latent pathlib.Path==str bug leaving the SET-definition files globally unfiltered, so the generated datafile carries 9437 for every scope; the filtered PARAMETER data gives the correct per-scenario 165/76 etc., cross-checked two ways). CBC solves for India+BBIN both reached Optimal (objectives 1,404,195 and 1,564,970, identical to the Jun-22 run) and reproduce the paper's Fig 2/3/4 qualitatively: ~96% nuclear-dominated systems by 2050 (India 96.2%, BBIN 96.0%), hydro/solar supplementary, BBIN cross-border trade 542 PJ over 9 links with more hydro than India. NOT attempted: World CBC solve (paper itself reports CBC infeasible <100h, needs commercial Gurobi) and the Spain/Portugal custom-config case (Fig 5/6, underspecified). The nuclear dominance is a property of the default cost/penalty assumptions and is what the paper reports. All grades PROVISIONAL; a human reviewer signs off via AUDIT.md.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.4730003

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 80
    assessed: 2026-06-22 ⛓ 1aca858cd125
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Can an open-source, open-data model generator provide full user flexibility in geographic scope and temporal resolution for building global electricity system models, overcoming the spatial, openness, and scenario-generation-speed limitations of existing global energy models?

Core claims
  • OSeMOSYS Global is an open-source, open-data global electricity system model generator built on the OSeMOSYS framework resource
  • OSeMOSYS Global allows full user flexibility in determining time slice structure and geographic scope, controlled via a single configuration file method
  • The model generator can produce models covering any scope from a single country up to all 265 available nodes across 164 countries resource
  • Open-source solvers (CBC) are infeasible for solving world-scale models in reasonable time, requiring commercial solvers such as Gurobi finding
  • Changing temporal resolution drastically changes generation mix results, particularly reliance on solar power, when storage is absent finding
  • Geographic scope changes the resulting generation technology mix (e.g., diverse mix globally vs. nuclear-dominated in India/BBIN) finding
  • OSeMOSYS Global does not yet include energy storage functionality finding
  • OSeMOSYS Global acts as the first step toward a full global energy system model generator incorporating multiple energy sectors method
Experimental setups
Assay System Perturbation Readout Platform
OSeMOSYS capacity expansion model simulation India (5 sub-country nodes) geographic scope system capacity and generation by technology CBC solver
OSeMOSYS capacity expansion model simulation Bangladesh, Bhutan, India, Nepal (BBIN, 8 nodes) geographic scope system capacity, generation, and electricity trade between nodes CBC solver
OSeMOSYS capacity expansion model simulation World (265 nodes, 164 countries) geographic scope system capacity and generation by technology Gurobi solver
OSeMOSYS capacity expansion model simulation Spain and Portugal temporal resolution (8 vs 144 time periods per year) system capacity and generation mix, 2015-2050
OSeMOSYS capacity expansion model simulation Spain and Portugal none (high temporal resolution scenario) hourly generation mix in 2030
Key results
  • CBC solver failed to produce a World-scenario solution within 100 hours, while Gurobi solved it in about 5 hours 5 hrs (Gurobi) vs >100 hrs (CBC, unsolved)
  • India scenario generated 165 technologies and 76 commodities, solved in 1 min via CBC
  • BBIN scenario generated 279 technologies and 131 commodities, solved in 3 min via CBC
  • World scenario generated 9438 technologies and 4127 commodities across 265 nodes
  • 2050 generation mix is diverse (wind, solar, nuclear, hydro) at world scale but dominated by nuclear in India and BBIN scenarios
  • Increasing temporal resolution from 8 to 144 periods/year drastically reduced solar power's share of generation, favoring dispatchable sources (nuclear, natural gas)
  • Hourly results show nuclear providing constant base load while wind and solar contribute variably, with natural gas ramping up sharply around 5pm as solar declines and load peaks
Key statistics
  • count 265 nodes (total geographic nodes available for model generation)
  • count 164 countries (countries covered in the World scenario)
  • count 17 power generation technologies (technology types included in the model (Table 1))
  • other $50 per Tonne of CO2 (carbon tax applied in example scenarios)
  • other 2015-2050 (model horizon used in example scenarios)
  • count 1 to 288 time periods per year (user-configurable temporal resolution range, over a horizon up to 2015-2100)
  • other 17% (proportion of national energy system decarbonisation analyses in a literature review that used OSeMOSYS)
  • other 1 min (India, CBC); 3 min (BBIN, CBC); 5 hrs (World, Gurobi) (approximate solve times by scenario on Intel Core i9-9900, 64GB RAM hardware)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods/software descriptor paper presenting OSeMOSYS Global, an open-source model generator for global electricity system optimization models, rather than an empirical study involving statistical hypothesis testing. Results are illustrative scenario outputs (capacity expansion and generation mixes) from deterministic linear optimization runs at varying spatial and temporal resolutions, reported as figures and summary tables (e.g., node/technology/commodity counts, solve times) rather than through statistical inference. No statistical tests, p-values, or measures of variability are reported, as the underlying results are outputs of a single deterministic optimization model run per scenario rather than repeated stochastic samples.

Replicationunclear GroupsDifferent geographic scopes (India, BBIN, World) and different temporal resolutions (8 vs 144 time periods/year) for Spain/Portugal, compared descriptively via model output figures Pairingna Randomization/blindingna Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • The paper reports single deterministic scenario runs for each geographic/temporal configuration (e.g., India, BBIN, World; 8 vs. 144 time periods) without repeated runs or uncertainty quantification.
    Could also: A structured sensitivity or uncertainty analysis (e.g., Monte Carlo sampling over key input parameters, or systematic scenario sweeps with reported ranges) could also be used — This would allow the range of plausible outcomes across parameter uncertainty to be summarized quantitatively, complementing the illustrative single-run comparisons already shown.
  • Spatial and temporal sensitivity is demonstrated qualitatively through a small number of illustrative comparisons (India vs. BBIN vs. World; two temporal resolutions for Spain/Portugal), which the authors note is a proof of concept.
    Could also: A full factorial or Latin hypercube sampling design across spatial and temporal resolution settings could also be used — This would enable a more systematic mapping of how model outputs change across the full range of configuration choices, which the authors themselves identify as future work.
  • Solve times are reported as single approximate values per scenario/solver combination (e.g., '5 hrs via Gurobi').
    Could also: Reporting solve times as a distribution (e.g., mean and range or SD across multiple runs/hardware configurations) could also be used — This would give readers a sense of variability in computational performance across runs or hardware, which can be useful for others planning similar model generation workflows.
  • Differences in generation mix between scenarios (e.g., dominance of nuclear in India/BBIN vs. a more diverse mix in the World scenario) are described narratively based on the figures.
    Could also: Quantitative comparison metrics (e.g., percentage-point differences in technology shares, or a formal decomposition of drivers such as resource limits vs. trade vs. carbon tax) could also be used — This would let readers quantify the magnitude of the differences discussed narratively, complementing the visual comparison already provided.
Software: OSeMOSYS (Open Source energy MOdelling SYStem) · Python · CBC solver · Gurobi solver

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-36241673 (OSeMOSYS Global)

Paper: Barnes T, Shivakumar A, Brinkerink M, Niet T. OSeMOSYS Global, an open-source, open data global electricity system model generator. Sci Data 2022. DOI 10.1038/s41597-022-01737-0 · PMCID PMC9568605.

Code (corrected): https://github.com/OSeMOSYS/osemosys_global @ v0.4.0 (commit b469377ded31ca71c07cc655b3c5318d3af94f4c). The scaffold enrichment had mis-linked niclasmattsson/Supergrid + zenodo:4730003 — those belong to a different paper (GlobalEnergyGIS). The paper's own Code Availability statement names OSeMOSYS/osemosys_global and says figures are replicable with version 0.4.0.

Nature of the artifact. OSeMOSYS Global is a model generator: a Snakemake workflow that ingests bundled, openly-licensed input data (PLEXOS-World 2015, Global Energy Monitor plant trackers, IRENA renewable profiles, IEA WEO costs, IAMC SSP GDP/POP) and emits an OSeMOSYS linear-program (a .txt/datafile), then solves it with an LP solver (CBC open-source; CPLEX/Gurobi for the World case) and post-processes results CSVs + figures. This is a textbook "third-party/own tool on the paper's own data" reproduction (BRIEF P16): valid and reproducible.

In scope (pipeline-derived, deterministic or solver-derived)

# Result Where Type Plan
C1 India scenario size: 5 nodes, 165 technologies, 76 commodities Table 2 model-generator output (deterministic) build India model (config geographic_scope=IND), count SETs in datafile
C2 BBIN scenario size: 8 nodes, 279 technologies, 131 commodities Table 2 model-generator output (deterministic) build with default v0.4.0 config (IND+NPL+BGD+BTN), count SETs
C3 World scenario size: 265 nodes, 9,438 technologies, 4,127 commodities Table 2 model-generator output (deterministic) build World model (scope=all), count SETs — build only, no solve
C4 India system capacity results (Fig 2a) Fig 2 solver output (CBC) solve India LP with CBC, extract TotalCapacityAnnual by tech
C5 India system generation results (Fig 3a) Fig 3 solver output (CBC) solve India LP, extract ProductionByTechnologyAnnual
C6 BBIN capacity / generation (Fig 2b, 3b) Fig 2/3 solver output (CBC) solve BBIN LP with CBC (~3 min)
C7 BBIN total electricity trade 2050 (Fig 4) Fig 4 solver output extract trade flows if C6 solves

C1–C3 (the SET counts) are the quick-minimum ~80% floor: fully deterministic generator outputs that should match Table 2 exactly if the code+data reproduce. C4–C7 require solving the LP and comparing figure-level values (qualitative shape

  • key magnitudes; the paper ships figures, not tables of these numbers, so grades are "partial/qualitative" at best).

Stretch (attempt after the floor)

  • World solve (Table reports ~5 h with Gurobi / ~20 h CPLEX): needs a commercial solver license. Likely env_unresolvable for the solve; the model-file generation (C3 counts) is still attemptable with open tools.
  • Spain & Portugal case study (Fig 5, 6): a custom non-default config (8 vs 144 timeslices). Re-deriving the exact config from the text is underspecified; attempt only if time permits, grade qualitatively.

Out of scope

  • No wet-lab / experimental component (this is a pure modelling paper) — nothing to exclude there.
  • Fig 7 (workflow diagram) and Fig 8 (demand-projection sample for Asia) and Table 1 (list of 17 technologies) are descriptive, not quantitative claims to re-derive (Fig 8 is an intermediate, could be spot-checked if cheap).

Compute plan (HARD RULE 1: heavy compute on «our HPC» only)

  • Env build + repo clone + data: on «our HPC» front1 (internet), into «infra» workdir «path».
  • Model build + CBC solve (India/BBIN: minutes; World build: light): SLURM job --partition=std --nodes=1 --cpus-per-task=N (no --mem
Figures / tables: TableFig 2aFig 3aFig 2bFig 4
C1_india_nodes
Reported
5
Reproduced
5
exact
C1_india_tech
Reported
165
Reproduced
165
exact
C1_india_comm
Reported
76
Reproduced
76
exact
C2_bbin_nodes
Reported
8
Reproduced
8
exact
C2_bbin_tech
Reported
279
Reproduced
278
within tolerance
C2_bbin_comm
Reported
131
Reproduced
130
within tolerance
C3_world_nodes
Reported
265
Reproduced
265
exact
C3_world_tech
Reported
9438
Reproduced
9437
within tolerance
C3_world_comm
Reported
4127
Reproduced
4126
within tolerance
C4_india_capacity
Reported
nuclear-dominated capacity (Fig 2a)
Reproduced
nuclear 626 GW of 726 GW total PWR capacity (86%)
partial
C5_india_generation
Reported
dominated by nuclear; solar low-hundreds PJ (Fig 3a)
Reproduced
nuclear 19,649 PJ (96.2%); solarPV 137.5 PJ; hydro 537 PJ
partial
C6_bbin_capgen
Reported
nuclear-dominated, more hydro than India (Fig 2b/3b)
Reproduced
nuclear 21,563 PJ (96.0%); hydro 588 PJ > India's 537 PJ
partial
C7_bbin_trade2050
Reported
BBIN cross-border trade PJ flows (Fig 4)
Reproduced
542 PJ over 9 inter-node directional links in 2050
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

Strong, near-complete reproduction using the authors' own open repo (OSeMOSYS/osemosys_global v0.4.0) with bundled open data. The deterministic Table 2 floor reproduced India bit-exact (5/165/76) and BBIN/World exact on nodes with only a Δ=1 (≤0.76%) offset on technologies/commodities, which is traced to the v0.4.0-shipped global SET files themselves — a single power-plant data-vintage drift, not our error and not a fabrication. CBC solves reached Optimal and reproduce the paper's Fig 2/3/4 qualitatively (~96% nuclear-dominated 2050 systems, supplementary hydro/solar, BBIN cross-border trade); the World CBC solve is legitimately out of scope per the paper itself. The only caveats are figure-level claims being comparable only qualitatively and the corrected scaffold code/data mis-link — neither undermines the central conclusion.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

392.2 k
tokens (I/O) · 30.6 M incl. cache
73 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.