Skip to content
akshay
All work

03 Estate turnaround Automotive 2022 — 2024

An estate nobody had documented

In August 2022 I joined the team that had just taken over the Azure integration estate of a global 4x4 vehicle manufacturer from its previous vendor. Nobody could say with confidence what interfaces existed, deployments were done by hand in the portal, and a standing population of failures recurred often enough to be normal and were never explained. The turnaround that followed was the lead's and the architect's to order and the team's to do. I was a mid-level developer on it for the whole of the engagement, built most of what it called for, and made five decisions of my own inside it. This is the estate, the order the team took it in, and the five.

fewer incidents, estate-wide
~30%
lower Azure spend, estate-wide
~15%
interfaces built or changed
30+

01

Context

The platform, as it stood when I left, was the middleware behind the manufacturer's public customer website — the configure, reserve and order journeys, on a frontend another vendor built. The website called APIs published on API Management; behind them, .NET Functions read and updated customer, order and vehicle data in SAP S/4HANA over OData, straight to the SAP gateway. Website events — leads, reservations, cases — landed on Service Bus and a Function delivered them to Salesforce over REST, a path that had been synchronous when the team arrived. Logic Apps — Consumption workflows when the team arrived, ten to twenty of them on one Standard plan by the end — carried the orchestrated and scheduled flows, and Storage queues and blobs the file exchange. Under ten thousand calls and messages a day. This was never a scale problem; it was a correctness one, because each of those calls was somebody configuring, reserving or ordering a vehicle.

The estate had come from a previous vendor in the state the team found it in. Interfaces existed that no list recorded, releases were done by hand in the portal, and a standing population of failures was absorbed by a support team that knew which interfaces to re-run in the morning. There was no outage to point at, which is what makes this kind of situation hard to argue for fixing. The order of the turnaround was the lead's and the solution architect's, and it was the right one — instrument every failure path before redesigning anything, put every release through a pipeline before touching the design, and take cost out as the inventory revealed what was unused rather than as a programme. I built most of what those calls asked for: the Azure Monitor alerts, scheduled KQL over Log Analytics into action groups, that reach a named owner on failures and dead-letter depth; most of the YAML and Bicep pipelines — one artefact promoted through Dev, Test and Prod, zip deploy for the Functions — and the app registrations and service connections behind them. I documented my own interfaces to the team's template. Those were not my decisions, and the decisions below are not those.

My seat was a mid-level one, under a named lead, in the integration workstream of a team of six to fifteen on the account. A Cognizant solution architect owned the design; I built and refined against it and contributed to the internal design reviews. Two incidents shaped the decisions that were mine. An expired secret took an interface down, which is where the identity work began. And SAP or Salesforce had slow days that the website felt directly — degraded pages, timeouts — which is where the cache, the asynchronous Salesforce path and the fail-fast timeouts came from. Later I added Event Grid, firing Functions off blob events for the file-exchange flows. Delivery of new interfaces never stopped for any of it, and when the lead was stretched, delivery was still mine to answer for.

02

Constraints

No stopping for repairs
New interface delivery continued throughout; none of this was a funded remediation programme.
A mid-level seat
I built to an order I did not set. Instrument first, automate first, cost as a by-product were the lead's and the architect's calls — and when the lead was stretched, delivery was still mine to answer for.
Downstream fixed
SAP and Salesforce were the manufacturer's systems, not ours to speed up or change. The only question open to us was what the website saw meanwhile.
Failures without a path
Retries unbounded or missing, exceptions swallowed, no dead-letter queue. A failing message retried for ever or vanished, and nobody was told either way.
Low volume, high stakes
Under 10,000 calls and messages a day, and every one of them a customer configuring, reserving or ordering a vehicle.

03

Architecture

Scroll the diagram sideways

Two paths, under ten thousand calls and messages a day. The request path: a public website calls API Management, which calls .NET Functions, which read and write SAP S/4HANA over OData, with a Redis cache above the Functions for reference data. The event path: the Functions put website events on Service Bus, and a Function delivers them to Salesforce; failures on either path reach an alert route to a named owner. Logic Apps Standard, one plan for ten to twenty workflows, sits to the side. Nodes are numbered 1 to 6 to match the legend below. Request path under 10K calls and messages a day Website configure · reserve · order API Management OAuth 2.0 · validate-jwt .NET Functions .NET 8 isolated · managed identity SAP S/4HANA OData · timeouts · bounded retry Redis SAP reference data Service Bus retry → dead-letter → owner Salesforce REST · async via Function Alert → owner KQL · failures · DLQ depth Logic Apps Standard one plan · 10–20 workflows
  1. 1 API Management, the gateway the website vendor's frontend calls. It had no OAuth before; now validate-jwt against Entra ID tokens, a client-credentials app for the website vendor, and APIM's own managed identity to the Functions in place of function keys.
  2. 2 The .NET Functions: moved from .NET 6 in-process to the isolated worker on .NET 8 across most of the estate, one app at a time. They reach Key Vault, Service Bus and Storage with managed identity in place of connection strings.
  3. 3 Azure Cache for Redis on the SAP reference data, added after a slow day at SAP became a slow day on the website.
  4. 4 SAP S/4HANA over OData: the manufacturer's system, fixed. The timeouts, bounded retries and idempotent writes on this edge are decision 01.
  5. 5 Service Bus to Salesforce: website events accepted at once and delivered by a Function, with a dead-letter route per interface that has a named owner and a runbook.
  6. 6 Logic Apps: ten to twenty Consumption workflows consolidated onto one Standard plan, for cost and cold start.
The request and event paths as they stood when I left. Numbered nodes are keyed to the legend below and carry the five decisions that were mine; the alert route to a named owner, unnumbered, was the lead's decision and my build.

04

Decisions

What was considered, what was chosen, and what that choice cost. The last line is the one that matters.

Decision 01

Bound every retry, and give every dead letter an owner

Considered
  • longer retry windows and higher retry counts
  • catch, log and move on
  • bounded retry with backoff and a timeout, a dead-letter route with a named owner, one error contract end to end
Chose
Bounded retry with exponential backoff and a timeout on every SAP and Salesforce call; a Service Bus dead-letter route per interface with a named owner and a runbook, wired into the alerting; one error contract from API Management and the Functions back to the website; idempotent writes to SAP and Salesforce.
Because
Unbounded retry hides a failure rather than handling it — a message that will never succeed becomes indefinite load and an alert nobody receives — and a missing retry does the same with less noise. Once retry is bounded, the only real question is how fast a human finds out and whether they can act, which makes the dead-letter queue, the owner and the runbook the same decision as the retry, not three separate ones. The error contract is that decision seen from the website: standard error responses, a correlation ID on every hop and structured logs, so a failure can be followed from the customer's page to the dead letter it became. And retrying a write promises the downstream a duplicate unless the write is idempotent; a duplicate order or lead is worse than a failed one.
Gave up
An operational obligation that had not existed. Every interface now has a person and a runbook, and a dead-letter queue without either is just a second silent failure, so the owner and the runbook were part of the change, not an afterthought. And a contract only holds if the caller honours it: part of this decision was another vendor's team building the website against it, on their release calendar rather than ours.

Decision 02

Take SAP and Salesforce off the website's critical path

Considered
  • faster SAP and Salesforce — not ours to buy
  • retry harder from the website
  • cache the reads that rarely change and queue the writes that can wait
Chose
Azure Cache for Redis in front of the SAP reference-data reads behind API Management and the Functions; Salesforce writes taken off the request — the event lands on Service Bus and a Function delivers it to Salesforce REST, dead-lettered if it cannot.
Because
When a downstream had a slow day the website had one too, and the downstreams were the manufacturer's to run, not ours. Most of the reference reads never needed to travel to SAP at all: the same lists came back unchanged hour after hour, and every trip was another chance to time out. Nothing about a lead or a reservation needed Salesforce to answer before the customer could be told "accepted"; the website needed the event kept, not delivered. Cache the first and queue the second, and a slow SAP becomes a cache miss while a slow Salesforce becomes a longer queue — neither of which the customer waits on.
Gave up
A staleness window on the reference data, with an expiry that has to be stated and defended rather than left as a default. An "accepted, not yet in the CRM" state that the business had to understand: a later Salesforce rejection is now a dead letter and a human, not an error the customer sees. And the asynchronous path is exactly where the first decision had to hold, because a queued write that fails with no owner is the failure this estate started with.

Decision 03

Identity in place of secrets, at the gateway and behind it

Considered
  • rotate on a schedule, and remember to
  • Key Vault references, with the keys and connection strings staying
  • OAuth 2.0 on API Management and managed identity both ways, so most secrets stop existing
Chose
validate-jwt on API Management against Entra ID tokens, with a client-credentials app for the website vendor; API Management calling the Functions with its own managed identity in place of function keys in policy; the Functions reaching Key Vault, Service Bus and Storage with theirs in place of connection strings.
Because
An expired secret took an interface down, and the failure was a date, not a bug: nothing had changed except the calendar. API Management had no OAuth in front of it, and the estate was held together by keys in policies and strings in app settings, each one a date waiting to arrive. Key Vault references move a secret; managed identity removes it, and a credential that does not exist cannot expire, leak or be pasted into the wrong environment. A client-credentials app makes the boundary between two companies a token with a lifetime and an audience, rather than a shared key both sides have to keep.
Gave up
The one credential that could not go: SAP's technical user, which stayed in Key Vault because the gateway wanted basic auth. And a change at a two-vendor boundary needs the other vendor to change too — the website had to start sending tokens on a release date that was not ours to set, so the boundary was only as done as their next release.

Decision 04

Ten to twenty Consumption workflows onto one Standard plan

Considered
  • leave them on Consumption and pay per run
  • a Standard plan per workflow
  • consolidate onto one Standard plan
Chose
One Logic Apps Standard plan carrying the ten to twenty small Consumption workflows — and the same arithmetic on tiers: API Management, Service Bus, Redis and the hosting plans stepped down from premium SKUs the traffic never justified.
Because
At under ten thousand calls and messages a day, spread across many small workflows, Consumption paid twice: an unpredictable per-action bill, and a cold start on every workflow that woke rarely enough for it to matter. A Standard plan has a fixed cost and warm workers, and at this volume the fixed cost was the lower one. The premium tiers elsewhere were the same mistake in the other direction — capacity bought for a load that never came, billed every month it did not.
Gave up
One plan is one blast radius. A runaway workflow now competes with its neighbours for the same workers, and a plan-level change takes all of them down together where Consumption would have taken one. And every workflow that moved had to be re-homed and retested, on an estate that was shipping features in the same sprints.

Decision 05

In-process .NET 6 to the isolated worker on .NET 8, across most of the estate

Considered
  • stay on .NET 6 until forced
  • upgrade the framework and keep the in-process model
  • move to the isolated worker on .NET 8
Chose
The isolated worker on .NET 8 for most of the estate's function apps, one app at a time through the pipelines rather than as one release.
Because
End of support is a failure with a date on it — the same kind as the expired secret, and just as avoidable. Moving app by app kept every interface's failure path testable as it moved: each app went through the same pipeline and the same alerts as any other release, so a fault showed up in one interface rather than across the estate. The isolated worker decouples the app's .NET version from the host's, which makes the next upgrade ours to schedule rather than the platform's to force.
Gave up
Months of work that changed nothing a customer could see, competing for the same sprints as features that did. And a stretch with two hosting models running side by side, which is two sets of behaviour to hold in mind for every incident until the last app moved.

05

Trade-offs accepted

  • Instrumenting first meant the reported failure count rose before it fell. That was the lead's decision and the architect's order, not mine; I built the alerts, and I would make the same call.
  • Automating releases before improving designs delayed every user-visible improvement. Also not my call — and it is the reason decisions 03 to 05 could ship one app at a time through a pipeline instead of by hand.
  • The inventory, the documentation template and the removal of unused resources were the team's, and they carry a good share of the ~15%. My share of it is the Logic Apps consolidation and the tier downgrades.
  • Caching and queueing bought independence from SAP and Salesforce at the price of a staleness window and an "accepted, not yet in the CRM" state, both of which someone has to be able to explain to the business.
  • The ~30% is the estate's number — failure and alert counts from Log Analytics and Application Insights across the team's interfaces, roughly the first year against the second. It is not a figure I can carve my share out of, and I have not tried.

06

Stack

  • API Management
  • Azure Functions
  • Logic Apps
  • Service Bus
  • Azure Cache for Redis
  • Event Grid
  • Key Vault
  • Managed Identity
  • Microsoft Entra ID
  • Azure DevOps
  • Log Analytics
  • C# / .NET 8

All work