Lab 1 · From notebook to tested package¶
Foundation Production track
Chapter 1 built a linear program. This lab looks at how that model is engineered: where the code lives, how it is tested, and how CI stops a broken model reaching the book. Every later engineering lab builds on this one.
The maturity ladder¶
Every model in this book climbs the same ladder. Chapter 1's model is on rung 3.
flowchart LR
N[Notebook] --> M[Python module] --> T[Unit-tested model] --> A[API]
A --> C[Container] --> D[CI/CD] --> S[Deployed service]
S --> Sch[Scheduled pipeline] --> O[Observable service]
O --> B[Backtested] --> Mon[Monitored] --> R[Re-optimised]
style T stroke-width:3px
1 · Where code lives¶
src/energy_or/
├── optimisation/
│ ├── lp.py # LinearProgram → LPSolution (+ SolveReport)
│ └── geometry.py # feasible polygon / vertex enumeration for 2-D pictures
├── network/
│ └── shared_connection.py # Case A: Generator, solve_shared_connection, pro-rata
├── data/
│ └── synthetic.py # clearly labelled synthetic datasets, seeded
└── viz/
└── lp_plots.py # Matplotlib views shared by figures and notebooks
Two rules:
- Notebooks call the library; they do not copy it. If a notebook cell grows past a
few lines of model logic, that logic belongs in
src/energy_or/. - One model, many views. The figures, the Manim animation, the notebook and the
tests all import
solve_shared_connection. When the model changes, everything that shows it changes too.
2 · Name things, keep units visible¶
Raw solver output is a vector: x = [20., 80.]. The library returns names and units:
result.dispatch_mw # {'WF_A': 20.0, 'WF_B': 80.0} MW
result.value_rate_per_h # 7400.0 $/h
result.value_per_interval(5) # 616.67 $ per 5-min interval
result.connection_shadow_price # 50.0 $/MWh
result.binding_constraints # ['shared_limit', 'avail_WF_B']
Attribute names carry their units (_mw, _per_h, _per_mwh). In energy modelling a
unit slip (MW versus MWh, $/h versus $ per interval) produces numbers that look
plausible and are wrong by a factor of 12. Naming is the cheapest defence.
3 · Test the mathematics, not just the code¶
An optimisation test should assert something you know is true mathematically:
| Test | What it proves |
|---|---|
test_reference_optimum_matches_hand_derivation |
Solver agrees with the hand-derived corner (20, 80) |
test_reference_shadow_prices |
Duals match the marginal-farm argument (50, 0, 30) |
test_shadow_price_predicts_objective_change |
A dual really is the derivative: +1 MW ⇒ +$50/h |
test_network_limit_and_availability_are_respected |
No solution ever violates a physical limit |
test_value_of_network_is_piecewise_linear |
Three regimes of the value curve |
test_participation_coefficients_change_merit_order |
Merit order is value per MW of limit |
test_lp_never_worse_than_pro_rata_on_synthetic_day |
Optimality holds in every one of 288 intervals |
test_infeasible_lp_reports_status_instead_of_crashing |
Failure is reported, not hidden |
A test caught the author
While writing the coefficient test, the expected answer was first derived by reasoning "WF_B is worth more, so it goes first". The solver disagreed, and the solver was right: per MW of limit, WF_A was worth more. Small models with hand-derivable answers are where you learn to trust, and to check, a solver.
Run them:
4 · Solver status is data (the first piece of OptOps)¶
Every solve returns a SolveReport:
result.report
# SolveReport(status=<SolveStatus.OPTIMAL: 'optimal'>, solver='HiGHS (scipy.optimize.linprog)',
# solve_time_s=0.0027, n_variables=2, n_constraints=3, message='Optimization terminated successfully. ...')
A production optimiser can be optimal, infeasible, unbounded, hit an iteration or
time limit, or fail outright. Each of those needs a defined response. Later labs record
these reports as metrics and alert on them, the same way you would alarm on a plant
trip.
5 · Quality gates in CI¶
Every push and pull request runs (.github/workflows/tests.yml):
flowchart LR
P[git push] --> L[ruff lint + format] --> Ty[mypy --strict] --> U[pytest] --> NB[execute notebooks]
P --> D[mkdocs build --strict] --> CF[deploy to Cloudflare Pages]
P --> AN[render Manim scenes]
- ruff catches style issues and common bugs.
- mypy --strict checks the type hints on the library.
- pytest runs the mathematical tests.
- Notebook execution runs every notebook top to bottom, so a lab can never silently rot.
- mkdocs build --strict fails on broken links.
- The docs workflow publishes the book to Cloudflare Pages, with a preview URL for every pull request.
Exercises¶
Guided
Change MINUTES_PER_HOUR to 50 in shared_connection.py and run uv run pytest.
Which tests fail, and why is that a good thing?
Engineering
Write a test asserting that, for random availabilities and values (use
numpy.random.default_rng(seed)), the LP value is always ≥ the pro-rata value.
This is a property-based test.
Production challenge
solve_shared_connection raises RuntimeError if the LP is not optimal. In a
five-minute control loop, is raising the right behaviour? Sketch a fallback policy
and the log line you would want to see at 3 a.m.