Skip to content
akshay
All work

04 Customer platform Energy utility 2019 — 2022

The day the platform was built for

My first three years were the backend of a US energy utility's public customer platform: C#/.NET services on Azure App Service behind API Management, serving the website and the contact centre's tools. Then a storm arrived, outage traffic ran to several times a normal day, and the platform met the day it existed for. I was one of twenty-plus people on the account. This is the part of the response that was mine, and the part that was not.

services and jobs shipped
20+
API calls a day, normal load
100K+
environments, released per service
3+

01

Context

The shape was conventional and sound. A CMS frontend and the contact centre's IVR tools called .NET APIs through API Management on subscription keys; the services read and wrote SQL Server on Azure, the system of record for usage, outage and service-request data; and account verification and meter data came from client-owned SOAP services fronting a legacy mainframe none of us touched directly.

During my time there the single backend was split into ten to twenty services, one per downstream dependency, so that a change to one could ship without redeploying the rest and a hot path could not take unrelated features down with it. The boundaries were the architect's; I built several of the services and, later, the trigger-driven ones on Azure Functions. Application Insights was wired throughout — telemetry existed, alerting did not.

Normal load was hundreds of thousands of API calls a day. A storm is the reason a utility's outage pages exist, and it is also the day every customer opens them at once.

02

Constraints

The surge is the point
A utility customer platform is judged on its worst day, not its average one. Capacity for the mean is capacity for the wrong day.
Fixed downstream
The SOAP services and the mainframe behind them were the client's, could not be scaled or changed by us, and set a hard ceiling on anything that called them.
A junior seat
I built to a design I did not own. The service boundaries, the queueing and the circuit breaking were others' calls; mine had to fit inside them.
Telemetry without alerting
Application Insights recorded everything and told nobody. The surge was found by watching, not by being paged.

03

Architecture

Scroll the diagram sideways

Traffic on a storm day rises to several times normal. Callers reach API Management, then a tier of ten to twenty services on App Service, then SQL Server; a client-owned SOAP layer in front of a mainframe sits to the side with fixed capacity, and Application Insights records every hop without alerting anyone. Nodes are numbered 1 to 6 to match the legend below. Storm day normal outage traffic, several times a normal day Callers web · IVR API Management keys · cache · throttle App Service 10–20 services · autoscale SQL Server record · purge jobs SOAP services client-owned · mainframe fixed capacity Application Insights telemetry · no alerts seen, not paged
  1. 1 Callers: the CMS frontend and the contact centre's IVR tools. On a storm day, several times the usual number of them at once.
  2. 2 API Management on subscription keys — where the throttling tripped, and where reference reads are now answered from cache without going further.
  3. 3 The service tier on App Service: ten to twenty services, one per downstream. Instances maxed out here first; CPU-based autoscale rules now add them before that.
  4. 4 SQL Server on Azure, the system of record — kept to the size of the work by scheduled purge jobs.
  5. 5 Client-owned SOAP services fronting a legacy mainframe: fixed capacity, and the timeouts and circuit breaking in front of it were a colleague's.
  6. 6 Application Insights: every hop reported; nothing alerted. The surge was seen, not announced.
The request path under surge. The numbers mark where it gave way and where the response landed; the two fixes that were mine sit at 2 and 3.

04

Decisions

What was considered, what was chosen, and what that choice cost. The last line is the one that matters.

Decision 01

Scale out on CPU, and accept that it lags

Considered
  • a larger fixed App Service plan
  • schedule-based scaling around forecast weather
  • metric-based autoscale on CPU percentage
Chose
Autoscale rules on CPU percentage, adding instances as load climbs and releasing them as it falls.
Because
A bigger fixed plan pays for the storm every day of the year and still has a ceiling. Scheduling on forecasts guesses. CPU is a lagging signal — by the time it climbs the queue is already long — but it was the signal the platform reported reliably, and instances maxing out was the first thing that failed, so instance count was the first lever.
Gave up
Cost that spikes with demand rather than being planned, and a few minutes at the start of every surge during which the platform is under-provisioned by design. A leading signal such as HTTP queue length is the better rule, and the one I would write today.

Decision 02

Cache the reads a storm does not change

Considered
  • a higher SQL tier for the duration
  • a read replica
  • response caching at the gateway for reference and lookup data
Chose
API Management response caching on the reference and lookup operations — the data every page requests and no storm alters.
Because
Under surge, much of what reached the services was the same handful of reference reads behind every page, and each one travelled through the gateway to a service to a database to return the answer it returned an hour earlier. Answering at the gateway took them off the service tier and the database entirely, which is cheaper than making either faster — and it was a policy on operations that already existed, not a build.
Gave up
A staleness window on data that does not, in practice, go stale — but the window has to be stated and defended, and the temptation to extend the same cache to outage status, which changes by the minute, had to be refused.

Decision 03

Purge on a schedule rather than buy a bigger database

Considered
  • a larger SQL tier as the tables grew
  • archiving to cold storage
  • a retention-window purge job
Chose
WebJobs that purged outage history and request and audit records past a retention window, on a schedule.
Because
The hot tables were slowing the APIs as they grew, and the growth was records nobody would read again. A larger tier treats the symptom and bills for it monthly; archiving preserves data with no reader. A purge job keeps the working set the size of the work.
Gave up
History past the window is gone, so the window is a decision with an owner rather than a default — and a job that deletes has to be the most carefully tested code on the platform.

Decision 04

One pipeline per service, with a test gate in each

Considered
  • one release pipeline for every service
  • a pipeline per service, each with automated test stages
Chose
Per-service Azure DevOps pipelines, each gated on its own automated tests, so a service could be released the moment it was ready.
Because
The point of splitting the backend into services was independent release, and a single shared pipeline gave that back. Per-service pipelines with their own test stages made on-demand release across three-plus environments the normal way to ship.
Gave up
As many pipelines as services to keep in step. A change to how the team builds anything is now a change in a dozen places.

05

Trade-offs accepted

  • Caching reference data at the gateway means accepting a staleness window during the event when everyone most wants fresh data — defensible only because the cached data is the part that does not change.
  • Scaling on CPU accepts a few slow minutes at the start of every surge. The right fix is a leading signal; this was the one available.
  • The queue-based decoupling of outage reports and the timeouts and circuit breaking on the SOAP layer were colleagues' work, and they mattered at least as much as mine. A case study that claimed them would be a better story and a worse record.

06

Stack

  • App Service
  • WebJobs
  • Azure Functions
  • API Management
  • Azure SQL
  • Application Insights
  • Azure DevOps
  • C# / .NET

All work