Skip to content

A Local Hybrid Retrieval Pattern for Health Records

Vector search is excellent at finding related language. It is less reliable when a question hinges on an exact identifier, an uncommon drug name, or a clinical code. Full-text search has the opposite personality: it rewards literal overlap but can miss paraphrases. Health-record retrieval needs both – and it needs a disciplined way to combine them without pretending their native scores mean the same thing.


Part 5 builds the retrieval half of an on-premises RAG system: conversation-aware rewrite and expansion, local BGE-M3 embeddings, concurrent vector and lexical search in DocumentDB, reciprocal rank fusion in Rust, row-level deduplication, adaptive cross-encoder reranking, a calibrated operational gate, and citations passed into local generation.


Retrieval is also where a cloud dependency is easiest to leave in by accident. Embedding and reranking are usually reached through an API, and a hosted reranker sees the retrieved passages – the most sensitive text in the pipeline – even when the chat model is local. Some countries require processing connected to primary or secondary health care to run on a server and data centre inside their borders, as Kenya, where the research behind this series was carried out, does; and where no residency rule applies, health data still leaves the country only under an adequacy finding or comparable safeguards. Either way, shipping passages to an offsite reranker is a transfer, however local the generation step is. BGE-M3 and the cross-encoder therefore run in process, on the same machine as generation. Part 1 sets out the boundary this follows from.


The prototype uses synthetic records. It is not clinical guidance, a compliance certification, or evidence that a score threshold transfers to another corpus.


Rust is used for the series implementation, but the pattern can be built with the Foundry Local SDK Reference in C#, JavaScript, Python, or Rust; Part 2 links to each language.


Problem statement


Neither vector search nor full-text search is sufficient across health-record questions. Dense retrieval can miss an exact code, identifier, or uncommon medication name, while lexical retrieval can miss a paraphrase that expresses the same event. Choosing only one creates predictable blind spots.


Naively combining both creates new errors. Cosine and text scores are not directly comparable, repeated query variants can over-reward the same record, and overlapping chunks from one row can crowd out independent evidence. Passing every weak candidate to generation also converts poor retrieval into confident-looking output.


The retrieval pattern must preserve exact and semantic recall without inventing a common score scale. It must fuse rankings, deduplicate at the source-row level, rerank a bounded candidate set locally, and apply an operational relevance gate before evidence reaches Foundry Local. It must do all of that without a network call leaving the facility, because the passages handed to a reranker are record content and the residency obligation follows them.


Success criteria



  • Exact terms and semantic paraphrases both receive a retrieval path.

  • Vector and lexical results are fused by rank rather than raw-score arithmetic.

  • Duplicate variants and overlapping chunks cannot inflate one row’s evidence.

  • A local cross-encoder reranks a bounded, diverse candidate set.

  • Weak evidence triggers an explicit no-relevant-records outcome instead of generation.

  • Accepted passages retain stable citations to their source records.


Start with the correct data boundary


PostgreSQL, MySQL, and SQL Server are grouped as External Clinical Databases. They remain operational systems of record. Ingestion copies selected, allowed fields into a separate Internal DocumentDB Hybrid Store, which contains documents, vectors, metadata, and application state. Retrieval operates on active copied generations in that internal store; it does not perform vector search against the source databases.


 



 


Foundry Local appears only at generative stages such as rewrite and answer generation. fastembed owns embedding and reranking. That distinction matters operationally: changing the chat model does not change the vector index contract.


Stage 1: make the question standalone


A follow-up such as “What happened after that?” is a poor retrieval query without conversation context. The server loads bounded working memory – a rolling summary plus a chronological verbatim tail – and asks a local model to produce a standalone rewrite. When expansion is useful, the same local call returns a small set of variants; the default total is three queries.


Expansion is not always beneficial. Short queries, quoted phrases, and identifier-like questions skip it because paraphrasing can erase exactly what matters. If model preparation fails or produces malformed output, retrieval falls back to the original question rather than failing the request.


Conceptually:


let (standalone, variants) = prepare_queries(
foundry,
&memory.rewrite_turns(),
question,
).await;

let queries = if should_expand(&standalone) {
unique_bounded(variants, 3)
} else {
vec![standalone.clone()]
};

Rewrite is a recall aid, not evidence. Conversation memory can clarify pronouns, but only retrieved records become citable context later.


Stage 2: batch and cache BGE-M3 query embeddings


Every query variant is embedded locally with fastembed BGE-M3. The current contract is:



  • 1,024 dimensions;

  • the same prefix-free encoder for records and queries;

  • normalized query text as the cache key; and

  • a process-local LRU cache capped at 1,024 entries.


Cache misses for all variants are submitted in one embedding batch and one blocking-task boundary. fastembed sessions are synchronous, so they run outside Tokio’s async reactor and remain guarded by a process-level mutex. This protects correctness, although concurrent requests can queue behind the shared model.


Batching matters because expansion should not multiply model setup overhead. Caching matters because follow-up turns and repeated test questions often reuse normalized text. Neither cache stores final answers or prompts.


Stage 3: run dense and lexical searches concurrently


For each query variant, hybrid mode creates two independent searches:



  1. Vector: DocumentDB cosmosSearch over contentVector, using an IVF cosine index.

  2. Lexical: DocumentDB $text over the projected text field.


Both sides filter for active: true, so incomplete ingest generations stay invisible. Every variant-by-side operation runs concurrently and returns a bounded ranking.


for (query, vector) in queries.iter().zip(query_vectors) {
searches.push(vector_search(db.clone(), vector, per_side));

if mode == RetrievalMode::Hybrid {
let lexical = enhance_exact_terms(query);
searches.push(text_search(db.clone(), lexical, per_side));
}
}

let ranked_lists = futures::future::join_all(searches).await;

The lexical enhancement adds quoted copies of ICD-like tokens and all-caps drug-like terms. It does not replace the user’s query; it gives exact tokens another chance to rank.


Concurrency reduces one request’s wall-clock latency, but it multiplies DocumentDB work. A retrieval admission semaphore is therefore part of the pattern. Parallel fan-out without a capacity boundary is an outage amplifier.


Stage 4: fuse ranks, never raw scores


A cosine score and a full-text relevance score are not probabilities and do not share a scale. A weighted average such as 0.7 * cosine + 0.3 * textScore looks scientific but is meaningless without careful normalization and corpus calibration.


Reciprocal rank fusion (RRF) ignores raw scores and combines positions:


 



 


Here,ri(d) is the one-based rank of document d in list i. The default k=60 dampens one-list spikes and rewards evidence that appears across variants or search modes.


 


The Rust implementation is deliberately small:


pub fn fuse(rankings: &[Vec<String>], k: f64) -> Vec<(String, f64)> {
let mut scores = HashMap::<String, f64>::new();
for ranking in rankings {
for (index, id) in ranking.iter().enumerate() {
*scores.entry(id.clone()).or_default() +=
1.0 / (k + index as f64 + 1.0);
}
}
let mut fused: Vec<_> = scores.into_iter().collect();
fused.sort_by(|a, b| b.1.total_cmp(&a.1).then_with(|| a.0.cmp(&b.0)));
fused
}

A stable ID tie-break makes repeated results deterministic when fused scores match. RRF runs in application code because the tested DocumentDB path supports cosmosSearch and $text, but not a native rank-fusion stage suitable for this design.


Stage 5: deduplicate by source row before truncation


Ingestion chunks long rows with overlap. Without row-level deduplication, several adjacent chunks from one source record can occupy most candidate slots. The system therefore keeps the strongest chunk per (source_id, table, row_pk) before truncating to the rerank candidate set.


That ordering is important. Deduplicating after top_k can return fewer than top_k unique rows even when more diverse candidates exist just below the cutoff. A second defensive dedupe occurs after reranking.


This design chooses diversity over parent reconstruction. The selected chunk carries cloned row fields for citation metadata, but the current implementation does not fetch sibling chunks or reconstruct the complete row. That missing parent-expansion stage is an explicit limitation, not an implied capability.


Stage 6: rerank only an adaptive prefix


RRF produces a useful candidate order, but it does not jointly reason over the question and passage. A local bge-reranker-v2-m3 cross-encoder scores the primary standalone question against a bounded candidate prefix.


The prefix is adaptive. It starts with at least max(top_k, 8) candidates and can stop at an elbow in fused scores, never exceeding the configured rerank maximum. This reduces expensive pairwise scoring when the fused list has a clear tail.


The reranker’s raw logits are converted with sigmoid:


 



 


This yields a stable operational range (0,1) , but it does not make the score a calibrated probability of clinical relevance. That distinction prevents a common documentation error.


 


After reranking, the weak tail can be trimmed below half the configured answer gate while retaining at least one passage. The default final context cap is six passages, also constrained by per-row and total approximate-word budgets.


Stage 7: gate before generation


A local model should not be asked to produce a grounded answer when retrieval found no adequate evidence. For pointed questions, the current default gate refuses when the top sigmoid-transformed reranker score is below 0.30. The threshold is configurable and can be disabled.


Broad questions – “list recent records” or “give an overview” – are handled differently. They may intentionally skip reranking, leaving RRF scores at the top. Comparing those values to a reranker threshold would mix scales again. Broad requests therefore refuse only when retrieval is empty.


fn should_refuse(question: &str, top_score: Option<f64>, gate: Option<f64>) -> bool {
match top_score {
None => true,
Some(_) if is_broad_question(question) => false,
Some(score) => gate.is_some_and(|floor| score < floor),
}
}

The gate is a risk-control mechanism, not proof of accuracy. It must be calibrated with representative answerable and no-answer questions from the intended corpus. The Microsoft architecture guidance similarly emphasizes evaluating retrieval and generation stages independently, then judging end-to-end behavior.


Stage 8: ground citations, then generate locally


Accepted passages become numbered context. The system prompt permits citations only to those records. Conversation memory provides continuity but is explicitly non-citable. Citations are emitted before answer tokens, and the successful assistant message persists the exact evidence used.


Foundry Local generates the answer locally through its embedded Rust SDK. There is no cloud-model fallback. A missing local model returns an unavailable response rather than broadening the trust boundary.


An optional verifier runs after streaming. It first checks citation indexes deterministically, then can ask a local verifier model to label claims supported, partial, or unsupported against the supplied passages. Verification is disabled by default, adds latency after the visible answer, and fails open as “skipped.” It annotates the answer; it does not retroactively make generation safe.


Missing pieces that matter


Three gaps should remain visible in architecture reviews:


General metadata filters are absent. There is no shared allow-listed source/table/patient/date filter contract applied to both vector and lexical search. The RetrievalFilter above is the planned answer; until it is built and verified, do not claim patient-safe prefiltering.


Parent expansion is absent. One strongest chunk per row survives, but sibling chunks are not reconstructed after ranking.


Chunking is approximate. The 384/64 windows and prompt budgets use whitespace counts rather than the embedding model’s tokenizer.


The router can classify hybrid cohort-plus-narrative questions, but it does not yet execute a structured cohort and inject its row keys into both retrieval sides. Such requests currently fall back to ordinary semantic retrieval.


Tradeoffs and anti-patterns


Hybrid search increases recall but doubles search operations per variant. Expansion increases that multiplier again. Bound both query count and per-side depth, and enforce admission.


Reranking improves precision but uses a serialized local model. An adaptive prefix controls cost, while broad-query handling avoids applying a pointed-question scorer where it does not fit.


Avoid these anti-patterns:



  • comparing vector and text scores directly;

  • deduplicating only after top_k truncation;

  • treating sigmoid output as a calibrated probability;

  • letting conversation memory become citable evidence;

  • generating despite an empty or weak pointed retrieval;

  • claiming metadata filtering because metadata happens to be stored;

  • assuming “hybrid” means one atomic database operation.


Prototype status


Implemented today: rewrite and three-query expansion; BGE-M3 query batching and caching; concurrent cosmosSearch and $text; Rust RRF with


k=60; pre-truncation row dedupe; adaptive bge-reranker-v2-m3; sigmoid operational scores; a configurable default 0.30 pointed-question gate; bounded context; citations; local grounded generation; and opt-in post-stream verification.


 


Still incomplete: corpus-specific gate calibration, general metadata filters, parent/sibling expansion, tokenizer-accurate chunking, and a full structured-cohort-to-hybrid-search executor. The scoped-retrieval and cohort designs above are specified in plans/new but not implemented. Validation uses synthetic health data and does not establish clinical safety or compliance.


Key takeaways



  1. Use dense and lexical retrieval for complementary failure modes.

  2. Fuse rank positions with RRF instead of mixing incomparable scores.

  3. Deduplicate source rows before candidate truncation.

  4. Spend cross-encoder work on an adaptive bounded prefix.

  5. Apply a score gate on the correct score scale and calibrate it empirically.

  6. Make only retrieved passages citable; memory is continuity, not evidence.


Continue the series



Learn more on Microsoft Learn




The code behind this series is open source: github.com/kevin-gatimu/onprem-health-rag. It is a synthetic-data research prototype, not a clinical or compliance-ready system — issues and corrections are welcome.


Microsoft Tech Community originally posted this article on 5 October 2026 at 8:00 AM.

Leave a Reply