Skip to content
akshay
All work

Independent project Retrieval-augmented generation 2026

Retrieval that cannot leak, and an upgrade refused

This one is mine. No client, no employer, no team: I built it alone on my own Azure subscription between May and August 2026, over a corpus of public documents I assembled, and the only people who have used it are me and a few testers. It is here because two of its decisions are ones I would defend on any engagement — enforcing a retrieval boundary structurally rather than conversationally, and refusing a model upgrade because the measurement said to.

chunks indexed
1.5M
groundedness, 87 cases
4.32/5
the upgrade refused
19.7% vs 16%

01

Context

The system answers questions about product claims and the research behind them over roughly 1.5 million indexed chunks, drawn from several document collections I assembled plus one approved external web source. Every answer carries citations and an explicit sufficiency verdict, because in this kind of material "the sources do not say" is a correct answer and a confident paraphrase of something adjacent is a defect.

Scale was never the interesting part. The interesting constraint was that evidence from the open web must never appear in an answer about the internal collections — and "must never" has to mean something stronger than a sentence in a system prompt. The second interesting part arrived later, when a newer model appeared in the catalogue and the obvious move was to upgrade.

Because it is my own subscription, the engineering standards are the ones I would apply on a client engagement rather than the ones a side project usually gets: Bicep modules for every component, pipelines across three environments, and managed identity rather than connection strings for the database. That was partly discipline and partly self-interest — I pay this bill.

02

Constraints

One person
No team to divide the work across and nobody to catch a bad assumption, so anything that could be measured had to be, because judgement alone had no second opinion.
A public corpus
Documents I gathered myself. That removes any confidentiality question and also removes the convenience of a corpus someone else had curated and deduplicated.
Citations or nothing
An answer without support is worse than no answer in this material, so the sufficiency verdict and the citation requirement are part of the contract, not presentation.
My own bill
Every design choice has a monthly cost I personally pay, which is a surprisingly good discipline for noticing which components are actually earning their place.

03

Architecture

Scroll the diagram sideways

Documents are ingested per source — extraction that preserves tables and layout, vectorisation, and a vision model that describes embedded images so a number inside a chart is retrievable. Each space has its own knowledge base carrying only its own sources; the approved web source is attached to one shared web-only knowledge base and nowhere else, and a space that may use web evidence reaches it through a separate call routed by an explicit registry. Agentic retrieval plans the query, runs hybrid vector and keyword search with semantic reranking, and a synthesis model composes an answer with mandatory citations and a sufficiency verdict. An 87-case evaluation harness measures groundedness against each answer's own citations. Nodes are numbered 1 to 6 to match the legend below. Ingestion, per source 1.5M chunks of public documents Documents decks · reports · web Extract · vectorise tables and layout preserved Verbalise images a chart becomes searchable text Space knowledge bases Internal base its own sources only Another space a different attachment set Web-only, attached nowhere else Approved web source reached by a separate call, routed by registry Agentic retrieval → synthesis hybrid search · reranking · citations · verdict 87-case evaluation harness groundedness against each answer's own citations
  1. 1 Ingestion runs per source: extraction that preserves tables and layout rather than flattening a document to plain text, then vectorisation with a large embedding model.
  2. 2 A vision model describes each embedded image and the description is embedded beside the surrounding text. This is what makes a number that exists only inside a chart retrievable at all.
  3. 3 A knowledge base per space, each attached only to the sources that space may see. The guarantee is structural: a source not attached to a base cannot be retrieved from it.
  4. 4 The approved web source is attached to one shared web-only base and nowhere else. A space permitted web evidence reaches it through a separate call, routed by an explicit registry.
  5. 5 Agentic retrieval plans the query, runs hybrid vector and keyword search with semantic reranking, and the synthesis model answers under instructions that require citations and a sufficiency verdict.
  6. 6 The evaluation harness: 87 cases, each answer judged for whether its claims are supported by the sources it chose to cite, plus out-of-scope, over-refusal and prompt-injection checks.
Ingestion along the top, the isolation boundary in the middle, retrieval and the evaluation harness below. The frames are the point: a source outside a knowledge base cannot be retrieved from it.

04

Decisions

What was considered, what was chosen, and what that choice cost. The last line is the one that matters.

Decision 01

Isolation by attachment, not by instruction

Considered
  • one index for everything, with a prompt rule about which sources to use
  • metadata filters applied at query time to exclude web evidence
  • separate knowledge bases, each carrying only the sources it is allowed to see
Chose
Separate knowledge bases per space, each attached only to its own sources, with the approved web source attached to one shared web-only knowledge base and nowhere else. A space that may use web evidence reaches it through a separate retrieve call, routed by an explicit registry.
Because
Prompts are advisory and filters are code that can be wrong. A source that is not attached to the knowledge base being queried cannot be retrieved from it — not "should not" — and that property holds regardless of how the question is phrased, how many hops it takes, whether the user is adversarial, or whether the model is replaced entirely. It also fails loudly rather than silently: a leak through a prompt rule is invisible until somebody reads a citation closely.
Gave up
More knowledge bases to provision and keep in step, and a routing layer the API resolves on every request. Sources are shared by reference rather than re-ingested, so storage barely moves, but the operational surface is genuinely larger. The topology is authored as one declarative file and compiled into both the knowledge bases and the per-environment registry, which is what keeps the two from drifting.

Decision 02

Verbalise the images, or lose what is inside them

Considered
  • text extraction only, accepting that charts are lost
  • OCR over the images and index the extracted characters
  • a vision model describing each image, embedded alongside the surrounding text
Chose
A vision model describes embedded images during ingestion and the description is embedded next to the text around it, alongside extraction that preserves tables and layout rather than flattening documents to plain text.
Because
Much of this corpus is decks and reports where the number that answers the question exists only inside a chart. Without verbalisation that content is invisible to search while appearing perfectly well indexed, which is the worst combination: the system looks complete and quietly cannot answer a category of question. OCR recovers characters and not meaning, and a chart's meaning is in its shape.
Gave up
Ingestion cost and ingestion time, and a dependency on the vision model's description being faithful — a wrong description is now retrievable evidence. Per-document failures are recorded rather than halting the run, because one content-filtered image should not stop an ingestion of hundreds of thousands of documents.

Decision 03

Build the instrument before touching the model

Considered
  • assess answers by reading them, which is what everyone actually does
  • gold-answer matching against a reference set
  • a fixed case set judged for groundedness against each answer's own citations
Chose
An 87-case evaluation set — 71 fact-specific cases and 16 behavioural ones, including 8 out-of-scope questions it must decline, 2 boundary questions it must answer and 3 prompt-injection attempts it must refuse — with each answer judged for whether its claims are supported by the sources it chose to cite.
Because
"Does this look right?" scales to about ten questions and then stops working. Judging against gold answers misses the failure mode that actually matters, which is a fluent answer citing the wrong document. And the over-refusal guard is the case people forget to write: a system that declines everything scores beautifully on safety and is useless, so the set has to be able to fail in that direction too.
Gave up
The harness took real time that could have gone into features, and it is itself a thing to maintain — every change to the corpus or the prompt can invalidate a case. It also established that single runs are too noisy for small deltas: one retrieval tweak looked like an improvement and turned out to sit inside run-to-run variance.

Decision 04

The newer model was measured, and not adopted

Considered
  • upgrade on release, since the newer model is presumed better
  • upgrade behind a flag and watch what happens
  • run the case set against both and let the result decide
Chose
Staying on the deployed model. It scored 4.32 out of 5 mean groundedness with 97% of answers grounded and a 16% hallucination rate; the candidate scored 4.17, 88% and 19.7%, and it over-declined questions the baseline answered correctly.
Because
Worse on every axis that mattered, and over-refusal is not a safe failure in a research tool — it teaches people the system is unreliable and they stop asking. Reading the failure cases individually mattered more than the scores: most were retrieval misses, where the right document was never fetched and the model filled the gap. No amount of model capability fixes a system that hands it the wrong evidence, so the lever was chunking, reranking and source routing rather than model selection.
Gave up
A negative result reads as nothing happening, and it is the hardest kind of work to show for a month. I also ran a frontier model from outside the catalogue against the twelve baseline failures as a ceiling — it fixed all twelve, and it sat outside the region I had pinned the data to, so it was never a deployable option. Useful as a measurement, not as a choice.

05

Trade-offs accepted

  • This is a personal build and should be read as one. There is no production incident record, no operational load and no user population beyond me and a handful of testers, so it evidences design judgement and engineering standards rather than anything about how the system behaves under real use.
  • The figures come from my own evaluation runs against my own case set. That makes them reproducible and it also makes them mine to mark — an 87-case set written by the person whose system it judges has a blind spot exactly where his assumptions are.
  • Structural isolation buys a guarantee and costs flexibility: adding a source to a space is a topology change that goes through CI rather than a configuration toggle. That is the right ratio for a boundary that matters and the wrong one for a corpus that changes daily.
  • Holding a personal project to client engineering standards — nine Bicep modules, fifteen pipelines, three environments, managed-identity-only database access — is defensible for the practice it gives and indefensible on cost per user, of which there is one.

06

Stack

  • Azure AI Foundry
  • Azure AI Search
  • Azure Functions
  • API Management
  • Front Door + WAF
  • Azure SQL
  • Managed Identity
  • Bicep
  • Azure DevOps

All work