Skip to content

Lab 1 · From notebook to tested package

Foundation Production track

Chapter 1 built a linear program. This lab looks at how that model is engineered: where the code lives, how it is tested, and how CI stops a broken model reaching the book. Every later engineering lab builds on this one.

The maturity ladder

Every model in this book climbs the same ladder. Chapter 1's model is on rung 3.

flowchart LR
    N[Notebook] --> M[Python module] --> T[Unit-tested model] --> A[API]
    A --> C[Container] --> D[CI/CD] --> S[Deployed service]
    S --> Sch[Scheduled pipeline] --> O[Observable service]
    O --> B[Backtested] --> Mon[Monitored] --> R[Re-optimised]
    style T stroke-width:3px

1 · Where code lives

src/energy_or/
├── optimisation/
│   ├── lp.py          # LinearProgram → LPSolution (+ SolveReport)
│   └── geometry.py    # feasible polygon / vertex enumeration for 2-D pictures
├── network/
│   └── shared_connection.py   # Case A: Generator, solve_shared_connection, pro-rata
├── data/
│   └── synthetic.py   # clearly labelled synthetic datasets, seeded
└── viz/
    └── lp_plots.py    # Matplotlib views shared by figures and notebooks

Two rules:

  1. Notebooks call the library; they do not copy it. If a notebook cell grows past a few lines of model logic, that logic belongs in src/energy_or/.
  2. One model, many views. The figures, the Manim animation, the notebook and the tests all import solve_shared_connection. When the model changes, everything that shows it changes too.

2 · Name things, keep units visible

Raw solver output is a vector: x = [20., 80.]. The library returns names and units:

result.dispatch_mw  # {'WF_A': 20.0, 'WF_B': 80.0}       MW
result.value_rate_per_h  # 7400.0                             $/h
result.value_per_interval(5)  # 616.67                             $ per 5-min interval
result.connection_shadow_price  # 50.0                               $/MWh
result.binding_constraints  # ['shared_limit', 'avail_WF_B']

Attribute names carry their units (_mw, _per_h, _per_mwh). In energy modelling a unit slip (MW versus MWh, $/h versus $ per interval) produces numbers that look plausible and are wrong by a factor of 12. Naming is the cheapest defence.

3 · Test the mathematics, not just the code

An optimisation test should assert something you know is true mathematically:

Test What it proves
test_reference_optimum_matches_hand_derivation Solver agrees with the hand-derived corner (20, 80)
test_reference_shadow_prices Duals match the marginal-farm argument (50, 0, 30)
test_shadow_price_predicts_objective_change A dual really is the derivative: +1 MW ⇒ +$50/h
test_network_limit_and_availability_are_respected No solution ever violates a physical limit
test_value_of_network_is_piecewise_linear Three regimes of the value curve
test_participation_coefficients_change_merit_order Merit order is value per MW of limit
test_lp_never_worse_than_pro_rata_on_synthetic_day Optimality holds in every one of 288 intervals
test_infeasible_lp_reports_status_instead_of_crashing Failure is reported, not hidden

A test caught the author

While writing the coefficient test, the expected answer was first derived by reasoning "WF_B is worth more, so it goes first". The solver disagreed, and the solver was right: per MW of limit, WF_A was worth more. Small models with hand-derivable answers are where you learn to trust, and to check, a solver.

Run them:

uv run pytest

4 · Solver status is data (the first piece of OptOps)

Every solve returns a SolveReport:

result.report
# SolveReport(status=<SolveStatus.OPTIMAL: 'optimal'>, solver='HiGHS (scipy.optimize.linprog)',
#             solve_time_s=0.0027, n_variables=2, n_constraints=3, message='Optimization terminated successfully. ...')

A production optimiser can be optimal, infeasible, unbounded, hit an iteration or time limit, or fail outright. Each of those needs a defined response. Later labs record these reports as metrics and alert on them, the same way you would alarm on a plant trip.

5 · Quality gates in CI

Every push and pull request runs (.github/workflows/tests.yml):

flowchart LR
    P[git push] --> L[ruff lint + format] --> Ty[mypy --strict] --> U[pytest] --> NB[execute notebooks]
    P --> D[mkdocs build --strict] --> CF[deploy to Cloudflare Pages]
    P --> AN[render Manim scenes]
  • ruff catches style issues and common bugs.
  • mypy --strict checks the type hints on the library.
  • pytest runs the mathematical tests.
  • Notebook execution runs every notebook top to bottom, so a lab can never silently rot.
  • mkdocs build --strict fails on broken links.
  • The docs workflow publishes the book to Cloudflare Pages, with a preview URL for every pull request.

Exercises

Guided

Change MINUTES_PER_HOUR to 50 in shared_connection.py and run uv run pytest. Which tests fail, and why is that a good thing?

Engineering

Write a test asserting that, for random availabilities and values (use numpy.random.default_rng(seed)), the LP value is always ≥ the pro-rata value. This is a property-based test.

Production challenge

solve_shared_connection raises RuntimeError if the LP is not optimal. In a five-minute control loop, is raising the right behaviour? Sketch a fallback policy and the log line you would want to see at 3 a.m.