Energy, power, and infrastructure demands from electrifying airport ground support equipment at United States airports.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
salvaged by watchdog from agreement.json (agent omitted ROOM_RESULT.json)
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-07-29
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusHow much new electrical energy, peak power, and charging infrastructure would be required to electrify ground support equipment (GSE) across major U.S. airports, and can behind-the-meter battery storage and solar PV mitigate the resulting grid impacts and costs?
- ★ A bottom-up, agent-based modeling framework can quantify site-specific energy demand, peak power, fleet/charger requirements, and costs for electrifying GSE at 317 U.S. airports method
- ★ Electrifying major GSE at the largest airports can create peak power demand up to ~20 MW and annual electricity consumption approaching ~51,000 MWh finding
- ★ Behind-the-meter battery energy storage systems (BTMS) and solar PV can reduce peak load and lower total system costs by as much as $10 million finding
- ★ Charging strategy and charger power level substantially affect peak power demand and the required number of GSE units and chargers finding
- Per-task energy consumption and battery sizing differ by GSE type and aircraft class (wide-body vs narrow-body), determining duty cycles and charging needs finding
- Off-peak overnight charging yields a distinct load profile with generally higher peak demand but concentrated in low-cost/low-usage hours finding
- The framework provides full-year, airport-level charging load profiles for all 317 airports as a resource resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Agent-based discrete simulation of eGSE service and charging events | 317 major U.S. mainland airports (large/medium/small hub, non-hub) covering 8 GSE types | Electrification (conventional GSE replaced by electric counterparts) under 6 charging scenarios | Charging load profiles, peak power demand, annual energy consumption, fleet size, charger counts, costs | — |
| Flight activity data analysis / scaling | 317 U.S. airports | none | Daily/annual flight arrival counts and arrival-time distributions | BTS Airline On-Time data and BTS T-100 Segment data |
| Per-task energy consumption and battery-size calculation | 8 eGSE types by wide-body vs narrow-body aircraft | none | Per-task energy consumption (kWh) and battery capacity (kWh) | GSE manufacturer specifications; ACRP/TRB report parameters |
| Behind-the-meter storage and on-site generation techno-economic analysis | U.S. airports under charging scenarios | Integration of stationary battery storage (BTMS) and PV systems | Peak demand reduction and charging/total system cost savings | — |
- ▲ Large hub airport peak power demand can reach up to 20 MW up to 20 MW
- ▲ Annual electricity consumption at largest airports approaches 51,000 MWh ~51,000 MWh
- ▼ BTMS and PV integration can lower total system costs up to $10 million
- – Large hub airport load ranges from 1–2 MW up to 10–20 MW across charging scenarios 1–2 MW to 10–20 MW
- – Medium and small hub airports generally require power demand below 5 MW <5 MW
- – Non-hub airports consistently remain under 1 MW peak power demand <1 MW
- ▲ Off-peak charging concentrates demand between 10 p.m. and 8 a.m. with generally higher peak demand
- ▼ Electric baggage tractor operates for under $9/day vs $20+/day for internal combustion equivalent <$9/day vs ≥$20/day
- other up to 20 MW peak power demand (largest airports peak load from eGSE)
- other ~51,000 MWh annual energy consumption (largest airports annual electricity use)
- other up to $10 million cost savings (from BTMS and PV integration)
- count 317 airports modeled (major U.S. mainland airports analyzed)
- count large hubs and medium hubs each 30 airports (9.5%); small hubs 65 (20.5%); non-hubs 60.5% (airport classification distribution)
- other >50% of airport infrastructure emissions from GSE operations (IEA-ETSAP life-cycle assessment)
- other 13% of total NOx emissions at U.S. airports from GSE (GSE air-pollution contribution)
- other GPU per-task energy 150.8 kWh (wide-body) / 50.6 kWh (narrow-body); 310/160 kWh battery (Table 1 per-task energy and battery size)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper presents a bottom-up, agent-based simulation framework to estimate energy demand, peak power, infrastructure requirements, and costs of electrifying ground support equipment (GSE) at 317 U.S. airports. No inferential statistical tests are applied; the study is a deterministic computational scenario analysis comparing six charging configurations (three strategies × two charger power levels). Results are reported descriptively using distributional summaries, load profiles, and scenario-level comparisons, with ranges characterizing variability across airports and conditions.
-
Vehicle specifications (battery capacity, power utilization rates, service durations) were fixed at single representative values drawn from ACRP reports and manufacturer data, applied uniformly to all airports↳ Could also: A probabilistic parameter sweep or Monte Carlo simulation could also be used, sampling input specifications from distributions reflecting observed variability across manufacturers and operators — Propagating input uncertainty would quantify how sensitive modeled peak demand and cost estimates are to the choice of representative specifications, making the range of plausible outcomes explicit
-
A single day (October 20, 2023) was described as randomly selected to illustrate daily load profiles across all airports↳ Could also: Percentile-based representative days (e.g., days at the 10th, 50th, and 90th percentile of annual demand) could also be used to characterize the range of typical, low, and high operating conditions — Percentile selection makes explicit how representative the illustrated day is relative to the full annual distribution, helping readers assess generalizability
-
Airports were grouped using existing FAA categorical classifications (large hub, medium hub, small hub, non-hub)↳ Could also: Data-driven clustering (e.g., k-means on modeled annual energy consumption, peak demand, or fleet size) could also be used to group airports by energy and infrastructure characteristics — Groupings derived from modeled outputs might reveal clusters more directly relevant to infrastructure planning than administrative categories that were not designed with energy demand in mind
-
Intra-group variability in load profiles and power demand is conveyed visually through overlaid line plots and descriptive ranges↳ Could also: Summary statistics with interquartile ranges or credible intervals, or quantile regression across airport characteristics, could also be used to characterize variability more precisely — Structured dispersion statistics would allow quantitative comparison of within-group spread and support downstream planning decisions that depend on worst-case or typical-case estimates
-
The six charging scenarios are compared qualitatively and graphically, without a formal quantitative ranking or sensitivity decomposition↳ Could also: A structured sensitivity or variance decomposition analysis (e.g., Sobol indices or one-at-a-time parameter sweeps) could also be used to attribute output differences to specific scenario dimensions — Decomposing the contribution of charging strategy versus charger power level to output variability would clarify which lever most strongly drives peak demand or cost, supporting targeted planning
-
The study covers 317 airports as a near-census; no extrapolation model is stated for airports outside this set↳ Could also: A regression or scaling model relating airport characteristics (e.g., annual enplanements, fleet size) to modeled energy outcomes could also be used to extend estimates to uncertificated or smaller airports — An explicit scaling model with stated assumptions would allow practitioners to estimate eGSE impacts at airports not directly simulated, broadening the applicability of the framework
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41912553
Paper: He Y, Kelly K, Jeffers M, Vercellino R, Ge Y, Lunacek M. "Energy, power, and infrastructure demands from electrifying airport ground support equipment at United States airports." Nat Commun 2026. DOI 10.1038/s41467-026-71125-4 · PMCID PMC13199391.
Code: https://github.com/NatLabRockies/AthenaGSEPaper
pinned commit 0f81b28905b146c3d05c852948de71088a9a290a (HEAD, "Update markdown with data source link").
Zenodo 10.5281/zenodo.18854761 is only a snapshot of this same repo (4.29 MB zip), not extra data.
Data: NREL/NLR Data Catalog submission 298 (https://data.nlr.gov/submissions/298, ~6.44 GB):
- Inputs:
BTS_on_time_flight_data.csv(313.5 MB),T_T100_all_2023.csv(43.5 MB) - Authors' precomputed OUTPUTS:
all_flight_energy_consumption.zip(1.2 MB),all_flight_GSE_charger_count.zip(15 KB),all_flight_GSE_vehicle_count.zip(33 KB),GSE_load_profiles_S{1..6}_*.zip(~840 MB each, per-airport per-minute charging power). The repo also ships a self-contained Demo for one airport (SLC) with its own input CSVs.
Pipeline (from README + code, all custom Python: pandas/numpy/geopandas)
- Scale up BTS on-time (domestic) flight data to total arrivals using T-100 (incl. international). STOCHASTIC — synthetic flights placed at random hours/days from the empirical arrival-hour distribution; no random seed is set in the code.
- Clean flight data (remove 00:00 outlier spikes).
- Generate GSE service tasks per flight (
get_GSE_tasks.py, deterministic). - Generate GSE service+charge events under 6 scenarios (deterministic given tasks):
- S1/S2 = charge when SOC insufficient (40/20 kW)
- S3/S4 = charge during service gaps (40/20 kW)
- S5/S6 = charge overnight 22:00–08:00 (40/20 kW)
- Postprocess overnight scenarios (S5/S6).
- Build per-minute load profiles (concurrent chargers × charger kW).
- Charger count = max concurrent chargers (per scenario).
- Vehicle count per GSE type (shared per airline; narrow/wide body).
- Plots/tables (
step_9_*.ipynb).
IN SCOPE (pipeline-derived, attempted)
| Claim | Reported (paper) | How reproduced |
|---|---|---|
| Table 3 — eGSE annual energy by airport category (MWh, min/max/mean) | Large 10,819/50,878/26,709; Medium 2941/12,102/5934; Small 577/9,239/1617; Non-hub 8/1204/265 | (A) Recompute from authors' all_flight_energy_consumption + FAA category map; (B) independently regenerate per-airport energy from BTS inputs via steps 1–8 |
| Abstract — peak power up to 20 MW at largest airport | "up to 20 megawatts" | From authors' load-profile zips (max over airports/scenarios) and/or independent run of largest airport |
| Abstract — annual energy approaching 51,000 MWh | "approaching 51,000 MWh" (= Table 3 large-hub max 50,878) | same as Table 3 |
| Per-airport charger counts | implied in Fig 5 | authors' all_flight_GSE_charger_count; independent SLC demo run |
| Per-airport vehicle counts by GSE type | implied in Fig 5 | authors' all_flight_GSE_vehicle_count; independent SLC demo run |
| SLC demo (pipeline validation) | not a named paper value; SLC is a large hub → must fall in large-hub ranges | run shipped Demo end-to-end, 5 stochastic replicates |
OUT OF SCOPE (not attempted / not in this repo)
- Table 4 (behind-the-meter battery + PV optimization: peak-demand & life-cycle-cost reductions, $10M savings): produced by the external EVI-EDGES optimization tool, which is NOT shipped in this repo. Cannot reproduce without that separate tool/config. Marked non-pipeline-for-this-repo.
- Any wet-lab / manual / external-tool values.
Reproducibility caveats
- Step 1 is unseeded-stochastic: exact charger counts / peak power vary per run; annual energy is ~deterministic (depends on total flight count, which is fixed) so it is the cleanest 1:1 target. Charger/peak compared at distribution/tolerance level.
- All heavy compute on «our HPC» («infra» «our HPC»-2), conda env built on compute node.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's headline computational results reproduce essentially exactly from the shipped code+data: abstract ~51,000 MWh = ATL 50,876.8, Table 3 category min/max and Table 4 ATL/LAX peak baselines match to the decimal, and every extreme corresponds to a real airport (no fabrication signal). The only material deviations are per-category means/min (1.7–5.5%) driven by a hub-classification boundary for ANC/BET/JNU — an input/sample-definition artifact on our side, not an energy-computation error. Two caveats keep this at yellow overall rather than green: the Table 4 optimisation layer ($10M savings, 20–50% peak reductions) is out of scope (external EVI-EDGES tool not in repo), and the independent stochastic regen (C17/C18) of charger/vehicle counts was still pending. The central conclusions hold fully.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.