← Back to project catalogue
GP-EE-0W1GUIPElectricalReady

Distribution Load Forecasting Benchmark

A completed electrical engineering benchmark for one-hour and 24-hour aggregate distribution load forecasting using historical demand, weather, calendar features, rolling-origin tests, peak metrics, and prediction intervals.

Distribution Load Forecasting Benchmark project visual
GP-EE-0W1GUIP · Electrical
  • Python
  • NumPy
  • Pandas
  • scikit-learn
  • SciPy
  • Matplotlib
  • Jupyter

Project definition

Problem statement

Distribution demand changes with hour, weekday, season, weather, holidays, and consumer mix. Random train-test splits can leak future patterns and overstate forecasting performance.

The engineering problem is to create a time-correct benchmark that compares simple and learned forecasts, measures peak error, and reports uncertainty without hiding failures in interval coverage.

Project objectives

  • Prepare an attributed hourly demand and weather dataset for four reproducible analytical feeder groups.
  • Create target-aligned lag, rolling, calendar, holiday, temperature, and humidity features.
  • Compare previous-day, previous-week, Ridge, and histogram gradient boosting methods.
  • Use four chronological train, calibration, and test origins at one-hour and 24-hour horizons.
  • Measure point accuracy, scaled error, peak magnitude, peak timing, condition error, interval coverage, and interval width.
  • Retain every prediction and calibration record for independent checking.

Project structure

Project components

01

Data preparation

Downloads the public UCI load archive and NASA POWER weather, creates four deterministic analytical groups, aligns hourly records, and writes validation evidence.

02

Feature pipeline

Builds origin-safe load lags, rolling statistics, cyclic time features, holidays, temperature, and relative humidity.

03

Forecast models

Runs two transparent persistence baselines, regularised linear regression, and histogram gradient boosting.

04

Uncertainty calibration

Uses a recent 60-day calibration window to construct 90 percent split-conformal intervals for each fold.

05

Evaluation pipeline

Retains timestamp-level predictions and calculates fold, aggregate, condition, peak, coverage, and width metrics.

06

Evidence generator

Produces open CSV and JSON results together with twelve PNG and SVG analytical figures.

Methodology

Project workflow

  1. 01
    Prepare the data

    Download, aggregate, align, and validate the public load and weather records.

  2. 02
    Build origin-safe features

    Create only the information that is available at each declared forecast origin.

  3. 03
    Run rolling tests

    Fit four models across four analytical groups, two horizons, and four seasonal origins.

  4. 04
    Calibrate intervals

    Calculate the conformal radius from the held-out calibration period for every model and fold.

  5. 05
    Check the evidence

    Compare point, peak, condition, coverage, and width results and retain all underlying predictions.

Demonstration scenario

Four analytical demand groups are forecast at one-hour and 24-hour horizons through winter, spring, summer, and autumn 2014 origins. Histogram gradient boosting produces the lowest point error at both horizons. The retained interval results also show that nominal 90 percent coverage is not achieved, which makes the uncertainty limitation visible rather than treating the interval as guaranteed.

Engineering

Tools and method

Tools
The project uses Python, NumPy, Pandas, scikit-learn, SciPy, Matplotlib, Jupyter for subject analysis, simulation, and results.
Numerical analysis
Python, NumPy, and Pandas handle source preparation, chronological alignment, features, and retained evidence.
Forecasting
scikit-learn supplies Ridge and histogram gradient boosting under fixed, documented configurations.
Uncertainty
A transparent split-conformal implementation uses absolute calibration residuals and a finite-sample rank.
Visualisation
Matplotlib creates feeder, weather, forecast, accuracy, stability, condition, peak, calibration, and residual figures.
Verification
Automated tests, static analysis, dependency audit, repository checks, Linux container execution, and rendered-document inspection form the release gate.

Testing

Evaluation

Evaluation measures

  • One-hour histogram gradient boosting nMAE of 2.177 percent and MASE of 0.384
  • 24-hour histogram gradient boosting nMAE of 3.446 percent and MASE of 0.621
  • One-hour peak magnitude MAE of 1,211.19 kW and peak timing MAE of 1.161 hours
  • 24-hour peak magnitude MAE of 3,206.78 kW and peak timing MAE of 1.317 hours
  • One-hour interval coverage of 87.770 percent with 8.882 percent normalised width
  • 24-hour interval coverage of 85.370 percent with 13.688 percent normalised width
  • Fold and condition evidence across four seasonal origins and four analytical groups

Project boundaries

  • The four feeder labels are deterministic groups of anonymised UCI client series and do not represent physical feeder circuits.
  • NASA POWER target-time observed weather is an optimistic explanatory input and is not an archived operational weather forecast.
  • The historical period predates recent electric-vehicle, heat-pump, rooftop-solar, storage, and tariff adoption.
  • The project does not model voltage, reactive power, network topology, equipment ratings, protection, dispatch, or customer behaviour.
  • The reported intervals are symmetric marginal intervals and do not guarantee seasonal, peak, feeder-specific, or future coverage.
  • The benchmark is an offline engineering study and is not an operational forecasting service.

Included

  1. 01Prepared public load and weather dataset with provenance and validation records
  2. 02Previous-day, previous-week, Ridge, and histogram gradient boosting forecasts
  3. 03One-hour and 24-hour rolling-origin evaluation across four seasonal origins
  4. 04Retained point forecasts, 90 percent intervals, fold metrics, condition metrics, and peak errors
  5. 05Twelve analytical figures in PNG and editable SVG formats
  6. 06Three attributed literature images and 46 annotated references
  7. 07Eighteen automated tests, static analysis, dependency audit, and Docker verification
  8. 08Complete project files, models, calculations, and analysis material in a private GitHub repository
  9. 0971-page project documentation in PDF and editable Word formats
  10. 1020-page setup and usage guide in PDF and editable Word formats

Project record

No information is collected on this page.

Permanent project ID
GP-EE-0W1GUIP
Catalogued
21 Aug 2026
Completed
28 Aug 2026
Verified
28 Aug 2026
Demonstration
Included in repository

Handover

After purchase

  1. 01
    Payment is confirmed

    The project is marked unavailable and cannot be purchased again.

  2. 02
    Repository access is granted

    The buyer's submitted GitHub account receives access to the private repository.

  3. 03
    The purchase record is delivered

    The certification sheet is prepared from the reviewed buyer details and sent privately by email.