Skip to content
akshay
All work

02 Integration platform Construction · ERP 2024 — 2025

Thirty interfaces, one go-live date

An Australian construction, property and finance group was putting in Dynamics 365 Finance and Operations, and the applications that ran its business — jobs, estimating, procurement, forecasting — had to feed it. I was the technical lead on the integration layer: I designed the chain every interface is built on, reviewed by the engagement's senior architects, and led the two or three developers who built it with me. Thirty-odd interfaces went live on one date.

Service Bus messages, peak day
500K+
interfaces live at go-live
30+
new-interface delivery
3 wks → 1 wk

01

Context

The applications were bespoke, built in-house over years on .NET and SQL Server, and they were where the business actually happened. The ERP was going in beside them, built by another team on the same programme, and the integration layer was the thing that decided whether the two halves were one system or two. A job created in the field had to become a project in the ERP. A purchase order raised against it had to arrive with its lines intact. A cost forecast revised on Friday had to be the forecast the ERP reported on Monday.

The chain I designed is the same for every interface, which is most of why it worked. The on-premises application calls an API Management endpoint over a private path. API Management validates the request against its contract and puts it on Service Bus, and the caller gets an acknowledgement and a correlation identifier rather than a wait. Behind the queue, either a Logic Apps Standard workflow or a Function picks the work up, maps it — in C# where the logic has conditions in it, in Liquid where it is a shape change — and delivers it to the ERP: a single record over OData, or a batch assembled into a data management package and imported. Business events come back the other way, onto Service Bus and out to the applications through a Function. A supplier order exchange runs beside it over AS2 with EDIFACT documents, through an Integration Account.

The programme's constraints were the ordinary bad ones. The date was fixed by the ERP go-live and everything else moved around it. Requirements changed continuously, because the ERP configuration was itself being decided while we integrated to it. The on-premises applications had no current documentation, so the first task on most interfaces was working out what the source actually did. And QA was a small team covering a programme much larger than the integration layer. None of that is unusual. What it does is remove every mechanism except discipline about contracts.

02

Constraints

One go-live date
Set by the ERP programme, not by us. Thirty-odd interfaces had to be live on the same day, so there was no interface that could be finished later.
Moving requirements
The ERP configuration was being decided while we built against it, so a field, an entity or a rule could change after an interface was written and tested.
Undocumented sources
The on-premises applications were years of in-house development with no current documentation. Every interface started with reverse-engineering the source.
A small QA team
Covering the whole programme, not just us. Anything we did not catch before handover was likely to reach user acceptance testing rather than be caught in it.
A pod of two or three
Two or three developers at a time, four over the engagement as people rotated. The design had to be teachable and reviewable, not merely correct.

03

Architecture

Scroll the diagram sideways

The inbound path: bespoke on-premises applications call API Management over a private connection, which validates the request against its contract and places it on Service Bus. A Logic Apps Standard workflow or a batching Function takes the work off the queue, maps it in C# or Liquid, and delivers it to Dynamics 365 Finance and Operations — as a data management package for bulk work, or over OData for a single record. Business events return from the ERP through Service Bus and a Function to the applications. A supplier order exchange runs beside it over AS2 with EDIFACT documents through an Integration Account. Nodes are numbered 1 to 6 to match the legend below. Inbound path 500K+ Service Bus messages on a peak day On-prem apps jobs · POs · forecast private path API Management contract first · 202 + ID Service Bus duplicate detection · DLQ Workflow or Function maps in C# and Liquid Batch → one package bulk · status and errors read back package · OData Dynamics 365 F&O upsert on a business key Return Function events out to the apps business events, back to the apps Integration Account AS2 · EDIFACT orders Suppliers two to four partners
  1. 1 API Management, the boundary every on-premises application calls over a private path. The contract is agreed and signed before the interface is built, and the caller is acknowledged with a correlation ID rather than made to wait.
  2. 2 Service Bus, the default transport. Duplicate detection catches the accidental redelivery; a dead-letter route per interface alerts a named owner when a message cannot be delivered at all.
  3. 3 A Logic Apps Standard workflow or a Function takes the work off the queue and maps it — in C# where the logic has conditions in it, in Liquid where it is a shape change.
  4. 4 The batching Function: it drains a batch of messages and assembles one data management package rather than importing per message, then reads the execution status and error file back so a failed row is visible as a row.
  5. 5 Dynamics 365 Finance and Operations, the target. Bulk work arrives as a package, single records over OData, and every write is an upsert against a business key so the same message twice is the same record once.
  6. 6 The supplier order exchange, beside the main path: AS2 with EDIFACT documents through an Integration Account, for two to four trading partners.
The chain every interface was built on. Numbered nodes are keyed to the legend below; the dashed path is the business events returning from the ERP to the applications.

04

Decisions

What was considered, what was chosen, and what that choice cost. The last line is the one that matters.

Decision 01

The contract is written, versioned and signed before the implementation

Considered
  • build first and document after
  • a shared canonical data model across all the applications
  • a signed interface design document per interface, with API Management as the boundary
Chose
An interface design document per interface — source, target, entity, field-level mapping, error behaviour — agreed and signed before build, re-baselined and re-signed whenever a requirement changed, with the API Management contract as the enforced boundary in front of it.
Because
Requirements that move are not a documentation problem, they are an attribution problem: without a signed baseline, every change is a disagreement about what was asked for, and on a fixed date those disagreements are what actually consume the schedule. Versioning the document turned each change into a visible, costed, agreed event instead. A canonical model was the tempting alternative and the one I rejected hardest — it makes every application wait on a modelling decision none of them own, and we had neither the time nor the authority to own it.
Gave up
Real time spent agreeing and re-agreeing documents rather than building, and a process that only works while both sides keep honouring it. It also makes the design documents the system of record, which means they go stale the moment anyone fixes something directly in code.

Decision 02

Asynchronous by default, synchronous by exception

Considered
  • synchronous calls from the applications straight through to the ERP
  • messaging for everything, including where the caller needs an answer
  • a queue behind the gateway, with a synchronous path only where the caller cannot proceed
Chose
Service Bus behind API Management as the default: the application is acknowledged with a correlation identifier the moment its message is accepted, and the delivery to the ERP happens behind the queue. A synchronous path only where the caller genuinely cannot continue without the answer.
Because
The ERP was the slowest and least predictable participant, and it was being configured while we integrated to it — availability we did not control and could not plan around. A synchronous chain is only as available as its least available link, so every synchronous interface would have made an on-premises application's usability depend on an ERP environment that was routinely down for a deployment. Accepting the message and owning it from there moves that dependency off the user.
Gave up
"Did it work?" stops being a response code. It becomes a question the platform has to have been instrumented to answer, and a support conversation whenever it has not been. The correlation identifier exists to make that conversation possible, and it only works because it is on every hop.

Decision 03

One package per batch, not one call per message

Considered
  • one data management import per message, as each arrives
  • a scheduled extract on the source side that bypasses the queue entirely
  • a Function that drains a batch off the queue and assembles one package from it
Chose
A Function that takes a batch of messages off Service Bus, maps them together and assembles a single data management package for the import, rather than one import per message. Single-record work — a customer created, a job updated — keeps going over OData as it arrives.
Because
The ERP's import is a job, not a call: it stages a package, runs it and reports on it, and the fixed cost of that cycle dwarfs the marginal cost of the rows inside it. One import per purchase order means the import queue becomes the bottleneck long before the data does, and on a peak day the message volume is well beyond what per-message imports could absorb. Batching moves the same rows through a fraction of the runs. The split matters as much as the batching: writes a user is waiting on stay single-record over OData, where the feedback is immediate.
Gave up
Latency, deliberately — a message now waits for its window to close. Failure granularity as well: a package that fails fails for rows it contains, so the import's execution status and error file have to be read back and the failing rows surfaced individually, or a batch becomes an all-or-nothing loss. That feedback path was part of the decision, not a follow-up to it.

Decision 04

Idempotency keyed on the business identifier, not the delivery

Considered
  • a central deduplication table the whole estate writes to
  • broker-level exactly-once semantics
  • handlers keyed on an identifier the business already recognises
Chose
Duplicate detection on Service Bus for the accidental redelivery, and every write to the ERP expressed as an upsert against a business key — the order number, the job number — so that the same message twice is the same record once.
Because
At-least-once is the only honest assumption once there is a queue, and exactly-once at the broker is a promise that dissolves the moment a handler has a side effect. Duplicate detection catches the same message arriving twice; it does not catch the same purchase order being sent again on Tuesday because somebody re-ran a job on Monday. Only the business key catches that, because only the handler knows what "the same thing twice" means for its domain. A duplicate purchase order in an ERP is a real financial event, not a tidiness problem.
Gave up
Every handler carries the burden and there is no central place to enforce it, so it lives on code review. It also constrains the mapping: the business key has to be present and stable in the source, and where it was not, that had to be fixed in the source application rather than worked around in the interface.

Decision 05

Templates, not a shared framework

Considered
  • an internal integration framework the team builds and owns
  • generating interface code from the design documents
  • reusable workflow, policy and pipeline templates a developer copies
Chose
Standardised templates — a Logic Apps Standard workflow skeleton, an API Management policy, a Bicep module set and a pipeline — that a developer copies and adapts, rather than a framework they take a dependency on.
Because
A framework becomes a product: it needs an owner, a backlog and a deprecation policy, and on a fixed-date delivery with two or three developers there is nobody to be that owner without stopping delivery. Templates gave most of the same benefit immediately — after the first few interfaces a new one was a copy, a mapping and a test rather than a design, which is what took a new interface from about three weeks to about one across the sprints. They are also teachable in an afternoon, which mattered more than elegance with people rotating through the pod.
Gave up
No single place to fix a systemic bug: a flaw in a template has to be chased across every interface that copied it, and the copies have already diverged. That debt is real and it is inherited by whoever maintains the estate.

05

Trade-offs accepted

  • Zero defects in user acceptance testing is the claim this engagement can make, and it is the one I make. My pod handed the estate to a separate support team at go-live after knowledge transfer, so I have no visibility of how it behaved in production afterwards and nothing here should be read as a production record.
  • The landing zone, the subscription layout and the platform standards were the engagement architect's, not mine. I designed the integration layer inside them, and the parts of this design that look like platform decisions were mostly constraints I was given.
  • Batching bought throughput with latency and with failure granularity, and both had to be paid back in engineering: a window that delays every message in it, and an import feedback loop that reads the execution status and pulls the error file so a failed row is visible as a row rather than as a failed package.
  • Contract-first absorbed the requirement churn, but it put the weight on a document. When the pressure came on at the end, the temptation to fix something in code and update the document later is exactly the failure mode this discipline exists to prevent, and it is only a discipline while someone is checking.

06

Stack

  • API Management
  • Service Bus
  • Logic Apps Standard
  • Azure Functions
  • Integration Account
  • Liquid
  • Key Vault
  • Bicep
  • Azure DevOps
  • Application Insights
  • C# / .NET

All work