Agents that dispatch help
A multi-agent roadside-emergency platform on Azure AI Foundry — and why the agent boundaries sit where they do
Akshay Kakoriya
Three or four agents on GPT-4o, built in about four weeks by a team of three for a hackathon. I architected it and wrote most of it. Every external service it calls is simulated, which I will say again at the bottom, because a claimed integration that was not real is a much worse thing to carry into an interview than a mocked one honestly described.
The scenario is a vehicle incident — a collision, a breakdown, a medical emergency in a car — and a caller who is frightened, possibly injured, and not going to answer questions in a helpful order. The system’s job is to work out what happened, how urgent it is, and who to send.
Why multi-agent rather than one prompt
A single model with a long instruction and every tool attached is the fastest thing to build and the wrong shape for this problem. Three reasons:
The tasks have different failure costs. Getting intake slightly wrong means asking a follow-up question. Getting triage wrong means an ambulance doesn’t get sent. Those shouldn’t share a prompt, a temperature, or an evaluation standard.
They have different tool access. Intake needs nothing but the conversation. Dispatch reaches the external services — geolocation, telemetry, facility lookup. Collapsing them would give the conversational surface, the part exposed to unpredictable input, direct access to the systems that summon emergency vehicles. That the services are simulated here does not change the boundary: the topology is the thing being designed, and in production those calls are real.
They need to be evaluated separately. “Did triage assign the right severity?” is a question you can build a test set for. “Did the system behave well?” isn’t.
So the work splits along those seams:
Intake
Talks to the caller and produces structured incident data from unstructured, distressed speech. People in emergencies don’t report location and injury status in a clean order; they describe what they can see. The agent’s job is to extract what’s needed and ask for what’s missing, without interrogating someone who is in trouble.
Its output is a structured record, not a decision. That boundary matters — intake never decides severity.
Intake is text chat, not voice. That was a scope decision rather than a design preference — speech-to-text in front of a distressed caller is its own problem, and it would have consumed the build without teaching us anything about the agent boundaries.
The record it produces carries six fields: location, incident type, vehicle, injuries, hazards and caller contact. The agent tracks which of those are still empty and asks only for those, which is the whole trick to not interrogating someone in trouble: a caller who leads with “there’s smoke and I can’t open the door” has just filled in hazards and told you something about injuries, and being asked next for their registration number is the behaviour that makes these systems hated. Order is the caller’s; completeness is the agent’s problem.
Where a field cannot be filled — no answer, no signal, a caller who cannot say where they are — it stays empty and moves on rather than blocking. An incomplete record is a valid input to triage, because the alternative is a system that stalls precisely when the emergency is worst. Geolocation and telemetry exist partly to fill those gaps without the caller.
Triage
Takes the structured record and assigns response priority. This is the agent that most needs constrained, auditable behaviour: a decision that determines whether someone gets an ambulance should be explainable after the fact.
Three severity levels, assigned against a rubric carried in the prompt rather than left to the model’s judgement. The rubric is the artefact: it states what puts an incident in each level, so a decision can be argued with afterwards by pointing at a clause rather than at a model.
At a boundary, the higher level wins. An incident that could be read either way is read as the more urgent of the two. That rule is not a tie-breaker for convenience, it is the design: the two errors are not symmetrical. Over-triage sends more help than was needed and costs money. Under-triage does not send an ambulance to someone who needed one. Any system where those are weighted equally has been built by someone who has not thought about which mistake they would rather explain.
Dispatch and routing
Selects and routes to the right responder — ambulance, tow, roadside mechanic — based on incident type and location. This is where the tool calls live:
- Geolocation to establish where the vehicle actually is, which the caller frequently cannot tell you.
- Vehicle telemetry for what the car itself reports — impact data, whether it’s drivable, fault codes.
- Nearest-facility lookup to find the closest appropriate responder, where “appropriate” depends on what triage decided.
Telemetry is the interesting one, because it’s the input that doesn’t depend on the caller being able to speak. A system that only knows what it’s told fails exactly when the emergency is worst.
All three are simulated. They are mock services returning realistic shapes, not live integrations, and calling them real would be the kind of claim that survives exactly one follow-up question. What the hackathon build demonstrates is the agent topology and the tool contracts — which agent may call what, with what arguments, and what it does when the answer is missing. Swapping a mock for a real endpoint behind the same contract is the easy half.
What I’d do differently at production scale
A hackathon build establishes that the agent topology works. Getting it into production would need:
- A triage evaluation set with labelled incidents and a measured confusion matrix, weighted so under-triage costs far more than over-triage.
- Deterministic escalation — any confidence below threshold routes to a human dispatcher rather than the model committing.
- Full decision audit, since every dispatch decision is potentially reviewable.
- Graceful degradation when telemetry or geolocation is unavailable, which in a crash is precisely when it might be.
The general point
Agent boundaries should follow failure cost and tool access, not conversational convenience. The temptation is to split agents by topic because that’s how you’d describe the system to someone. The useful split is by what happens when each part gets it wrong, and by what each part is allowed to reach.
Azure AI Foundry and its agent framework. Three or four agents on GPT-4o, about four weeks, a team of three; I architected it and built most of it. Written for Capgemini Hackathon 2026, which was a co-branded event with Microsoft. It was not carried forward afterwards, and every external service it calls is simulated.