A RAG store is often introduced as a vector database. That description is too narrow for a real application. Retrieval needs text and vectors, but ingestion also needs checkpoints and publication state. Structured planning needs schema metadata. Chat needs conversations and evidence. Operations need users, settings, and audit events.
This prototype uses a local DocumentDB engine as one internal hybrid store for those responsibilities. The application connects through the MongoDB wire protocol with the Rust mongodb driver. Inside the engine, PostgreSQL and pgvector provide implementation foundations, but application code does not connect to that internal PostgreSQL layer or issue SQL against it.
The distinction is important: this DocumentDB is neither Azure Cosmos DB nor MongoDB Atlas. A compatible wire protocol does not imply identical search operators, index syntax, or product behavior. The implementation and its dated verification target the local DocumentDB engine used by this repository.
Although this series implements the system in Rust, the pattern works with the Foundry Local SDK Reference for C#, JavaScript, Python, or Rust; see Part 2 for language links.
Where the store runs matters as much as how it is shaped. Chunk text, embeddings, schema samples, and conversation history are all derived from patient records, so a data-residency obligation reaches them exactly as it reaches the source rows. Where a country requires processing for primary or secondary health care to run on a server and data centre inside its borders – as Kenya, the setting of the research behind this series, does – a serving copy of the index has to sit there too. A managed cloud vector database is therefore not a drop-in substitute for this design, whatever its encryption posture: moving the index offsite is a cross-border transfer of health data in its own right. Running DocumentDB inside the facility keeps the derived copy in the same jurisdiction as the record it came from. Part 1 sets out the boundary in full.
Prototype status
Collection ownership, active-generation ingestion, cosmosSearch IVF cosine retrieval, $text retrieval , Rust reciprocal rank fusion, local reranking, conversations, schema catalog state, and readiness checks are implemented. The exact search syntax was live-verified against documentdb-local:latest on August 23, 2026. Corpus-scale index tuning, broad metadata filtering, complete backup/restore evidence, failure injection, and production security hardening remain open. The project uses synthetic data and makes no clinical-advice or compliance claim.
Problem statement
A hybrid RAG store is easy to mislabel as either the clinical database or merely a vector index. Both descriptions hide essential responsibilities. Operational PostgreSQL, MySQL, and SQL Server data has a different authority and freshness model from copied chunks, while retrieval depends on text, vectors, metadata, and application state moving together.
If that boundary is unclear, stale derived records can be presented as live truth, rebuilds can publish mixed generations, and citations can lose their source-row provenance. Separate ad hoc stores can also leave schema cards, conversations, ingest checkpoints, and search indexes without a coherent ownership or backup contract.
The design must make DocumentDB an explicitly internal, application-owned hybrid store. It must preserve source and row provenance, publish complete generations atomically, support both vector and lexical retrieval, and define contracts for metadata and application state without confusing MongoDB-wire compatibility with product equivalence. It must do so in a store that runs inside the facility, because a derived copy of a health record carries the same residency obligation as the original.
Success criteria
- External Clinical Databases and the Internal DocumentDB Hybrid Store have distinct ownership.
- Every searchable chunk identifies its source row, chunk, and ingest generation.
- Vector and full-text indexes operate only over successfully published data.
- Active-generation pointers prevent incomplete refreshes from becoming searchable.
- Metadata, conversations, jobs, settings, and audit state have explicit collection owners.
First separate source truth from retrieval state
PostgreSQL, MySQL, and SQL Server are external clinical systems of record. The internal store contains selected, transformed copies designed for application workflows.
This separation creates two legitimate freshness models:
- guarded live SQL sees the external source at query time;
- semantic retrieval sees the last successfully published ingest generation.
Neither should masquerade as the other. The internal copy is not the canonical patient record, and top-k retrieval is not an exact counting engine.
Why one hybrid store is useful
The internal store combines four related data planes:
- Retrieval data: chunk text, BGE-M3 vectors, structured row fields, source identity, and generation state.
- Metadata: source registrations, indexed-table summaries, versioned schema cards, relationships, and administrator overrides.
- Application state: users, settings, ingestion jobs, conversations, messages, citations, and verification results.
- Operational evidence: best-effort audit events and durable ingestion snapshots.
Keeping these together simplifies local deployment and allows application-level workflows to refer to consistent identifiers. It does not remove the need for explicit ownership or consistency rules. DocumentDB remains schemaless at the storage layer; Rust structs, validation, indexes, and write protocols form the actual contract.
The RAG solution design and evaluation guide is useful here because it treats retrieval as more than a vector query. Data preparation, search, ranking, grounding, and evaluation must be designed together.
The record is a chunk with row provenance
One document in records represents one embedded chunk from one source row in one ingest generation. Its identifier is:
{source_id}:{table}:{row_pk}:{chunk_index}:{ingest_generation}
The generation suffix allows the current and staged forms of the same row chunk to coexist safely. A conceptual record looks like this:
{
_id: “<source>:encounters:<row>:0:<generation>”,
source_id: “<source UUIDv7>”,
table: “encounters”,
row_pk: “<stable source-row identity>”,
chunk_index: 0,
fields: {
patient_id: “SYN-2024-0001”,
encounter_type: “outpatient”,
// selected source fields
},
text: “patient_id: SYN-2024-0001 …”,
contentVector: [/* 1024 BGE-M3 values */],
ingest_generation: “<generation UUIDv7>”,
active: false,
extracted: {/* optional local annotations */},
ingested_at: ISODate(“…”)
}
fields preserves selected structured values for display and constrained aggregation. text is a retrieval-oriented projection that omits configured noise such as opaque IDs or audit fields where appropriate. contentVector is produced locally by fastembed BGE-M3. Optional extraction annotates the record; it is not redaction.
Chunking currently uses approximately 384 whitespace-delimited words with 64-word overlap. That is deliberately described as an approximation, not tokenizer-accurate segmentation. Microsoft’s RAG enrichment guidance provides a broader framework for deciding how chunking and enrichment should support downstream retrieval.
Storage is chunk-grained, but provenance remains row-grained. Retrieval keeps the strongest chunk per (source_id, table, row_pk) before final candidate truncation, preventing overlapping chunks from one row from crowding out all other evidence.
Collections have explicit owners
The current store is more than records:
|
Collection |
Responsibility |
|
users |
Password hashes, roles, token-version revocation state, and profile data |
|
sources |
External database definitions and encrypted credentials |
|
records |
Active and staged chunks, vectors, text, structured fields, and provenance |
|
indexed_tables |
Per-table active-generation pointer, counts, and refresh state |
|
jobs |
Durable ingestion request, checkpoint, progress, bounded log, and outcome |
|
schema_catalog |
Versioned source-table cards, relationships, samples/profiles, text, and vectors |
|
schema_catalog_state |
Last-known-good active schema version and refresh health |
|
schema_catalog_history |
Bounded schema refresh outcomes |
|
schema_metadata_overrides |
Administrator aliases and undeclared relationships |
|
chat_conversations |
User-owned conversation metadata and rolling memory state |
|
chat_messages |
Transcript, citations, structured/SQL evidence, and optional verification |
|
settings |
Persisted model-role overrides and server settings |
|
audit_log |
Best-effort security and administrative events |
|
schema_bindings (planned) |
One SchemaBinding per source: each table’s hospital concept, column roles, service lines, and shortest join path to the patient table, versioned against the active catalog |
|
schema_binding_history (planned) |
Bounded prior binding versions for the administrator diff view |
The last two rows are design direction, not shipped code. They come from the hospital-agnostic service-line design in plan 01 – service-line ontology and schema binding: the agent roster is a fixed ontology, and the binding records how one deployment’s physical schema maps onto it. Storing the binding beside the schema catalog keeps the two on the same refresh timeline and lets every consumer – router, deterministic SQL, linker, retrieval, UI – read one document instead of hardcoded table names.
There is no separate raw-row collection, dedicated vector database, analytics warehouse, or telemetry database in the current implementation. That simplicity is useful, but it makes backup scope broad: vectors, source-derived fields, schema samples, conversations, and model outputs are all sensitive derived health data.
Publish ingestion by generation
Refreshing a searchable corpus in place creates a dangerous interval: old chunks disappear before new embedding and storage succeeds. The project instead uses last-known-good publication.
For each selected table, ingestion:
- validates the requested table against live schema;
- reads a bounded, ordered page from the external source;
- removes explicitly excluded fields;
- projects and optionally annotates each row locally;
- chunks and embeds text in local batches;
- writes all chunks with a new ingest_generation and active: false;
- advances a durable checkpoint only after the page is written; and
- publishes the complete generation in a transaction.
The cutover transaction deactivates the previous current records, activates the staged generation, and updates indexed_tables.active_generation with counts and timestamps. Older inactive data is deleted afterward on a best-effort basis.
This protocol provides a clear reader rule – query only active: true – and avoids exposing a half-refreshed table. Resume can replay a checkpoint page safely by deleting matching inactive row keys in the same generation before reinserting them.
The same principle appears independently in the schema catalog. A complete new catalog version is written before schema_catalog_state.active_version moves. Record generations and schema versions are separate timelines and must not be conflated.
Build two search sides, not one blended score
The records collection has two search indexes:
// Vector index command shape used by this engine.
{
createIndexes: “records”,
indexes: [{
name: “records_contentVector_cosmos”,
key: { contentVector: “cosmosSearch” },
cosmosSearchOptions: {
kind: “vector-ivf”,
numLists: 100,
similarity: “COS”,
dimensions: 1024
}
}]
}
// Lexical text index.
{
createIndexes: “records”,
indexes: [{
name: “records_text”,
key: { text: “text” }
}]
}
The vector side uses cosmosSearch as the first aggregation stage, over-fetches, filters for active records, limits the results, and projects searchScore. The lexical side uses $text, restricts results to active: true, and sorts by textScore.
// Simplified vector search pipeline.
[
{
$search: {
cosmosSearch: {
vector: queryVector,
path: “contentVector”,
k: requestedCount * 2
}
}
},
{ $match: { active: true } },
{ $limit: requestedCount },
{
$project: {
text: 1,
fields: 1,
source_id: 1,
table: 1,
row_pk: 1,
chunk_index: 1,
score: { $meta: “searchScore” }
}
}
]
The stage ordering and index options are engine-specific. They were verified against the dated local image, not inferred from MongoDB Atlas or Azure Cosmos DB documentation. Atlas-style $vectorSearch and a native $rankFusion stage were not available on the tested path.
Scoping search (planned)
The pipeline above filters only on active: true. The service-line design in plan 04 – structured execution and fallbacks adds a typed RetrievalFilter:
pub struct RetrievalFilter {
pub tables: Vec<String>, // tables bound to the active service line
pub source_ids: Vec<String>,
pub row_pks: Option<Vec<String>>, // structured cohort keys, capped
pub patient_key: Option<String>, // one patient’s row plus one-hop child rows
pub explicit: bool, // user named the scope, or it was inferred
}
Because cosmosSearch must be the first stage, the filter cannot simply be a $match placed in front of it. The intended shape passes the constraint through the cosmosSearch filter option, for example filter: { table: { $in: […] } }. Whether the local engine honours that option for the IVF index has not yet been verified live; if it is rejected, the fallback is to over-fetch roughly four times the requested k (capped) and $match afterwards, accepting weaker recall inside a tight scope. The $text side needs no such workaround because its filter fields sit in an ordinary $match.
Either variant needs compound indexes that do not exist today: { table: 1, active: 1 }, { source_id: 1, table: 1, active: 1 }, and a per-deployment index on the patient business-identifier field named by the binding. Treat this subsection as the design target; the dated verification note above still describes what is proven.
Fuse ranks in Rust
Cosine search scores and text-search scores do not share a meaningful scale. Adding or averaging them would manufacture precision. The server instead fuses positions with reciprocal rank fusion:
For each rewritten or expanded query, vector and lexical searches run concurrently. Their ordered document IDs become ranked lists. A small Rust implementation captures the essential operation:
use std::collections::HashMap;
fn fuse(rankings: &[Vec<String>], k: f64) -> Vec<(String, f64)> {
let mut scores = HashMap::<String, f64>::new();
for ranking in rankings {
for (zero_based_rank, id) in ranking.iter().enumerate() {
let rank = zero_based_rank as f64 + 1.0;
*scores.entry(id.clone()).or_default() += 1.0 / (k + rank);
}
}
let mut fused: Vec<_> = scores.into_iter().collect();
fused.sort_by(|a, b| {
b.1.partial_cmp(&a.1)
.unwrap_or(std::cmp::Ordering::Equal)
.then_with(|| a.0.cmp(&b.0))
});
fused
}
RRF rewards evidence that ranks well across query variants or search modes without pretending their raw scores are comparable. The deterministic ID tie-break makes repeated results easier to test.
Hybrid retrieval in this design is therefore two DocumentDB queries plus application-side fusion, not one atomic database operation.
Rerank, deduplicate, and gate before generation
RRF is the recall-oriented merge, not the end of retrieval. The server then:
- keeps the strongest chunk for each source/table/row identity;
- selects an adaptive, bounded candidate prefix;
- runs local bge-reranker-v2-m3 against the primary standalone query;
- applies a defensive second row deduplication;
- trims very weak tail passages; and
- applies a relevance gate before generation.
For pointed questions, the project compares sigmoid-transformed reranker scores with an operational floor, currently 0.30 by default. Broad overview questions do not compare RRF values with that reranker threshold because the scales differ; they refuse only when retrieval is empty.
This mechanism is implemented, but the threshold is not claimed to be clinically calibrated. It requires representative answerable and unanswerable evaluation for each corpus and model combination.
Foundry Local receives only the selected numbered passages and bounded conversation memory. It generates the answer locally. DocumentDB does not run the LLM or reranker, and Foundry Local does not own the vector index.
Use structured fields without letting chunks inflate counts
Copying row fields onto every chunk is convenient for evidence display and constrained aggregation, but it creates a trap: counting documents counts chunks, not source rows.
The aggregation subsystem avoids that error by matching the selected physical table and active: true, then grouping by row_pk before counts or metrics. Logical table and field names are validated against an in-memory catalog. Rust constructs the pipeline; the model cannot submit arbitrary BSON stages. Limits and maxTimeMS bound execution.
This path answers over the latest active ingested snapshot. For questions requiring current operational values, the separate guarded live-SQL path remains the correct choice.
Store provenance with the answer
A hybrid store is valuable when it preserves how an answer was produced:
- semantic answers store ordered passage citations with source/table/row/chunk identity and ranking metadata;
- live SQL answers store source ID, executed SQL, columns, and exact rows;
- constrained internal aggregations store their validated specification, rows, and pipeline;
- conversation-only answers carry no record citations.
The transcript is server-owned and user-scoped. A user turn is persisted before generation; the assistant turn is persisted only after successful completion. This means a stream failure can leave an unmatched user message, which is more honest than recording a partial answer as successful.
Conversation memory uses a rolling local summary plus a bounded verbatim tail. Compaction applies compare-and-swap to its summary cursor so concurrent work does not interleave summaries.
The planned conversation model adds a second, typed memory beside that summary. Plan 06 – conversation memory and suggestions stores a focus document on the conversation record – current patient, place, time range, service line, last query specification, and a PHI-minimised digest of the last result – updated by code after each turn with no model call. Assistant messages would also gain provenance (the execution path actually taken), spec (the structured query IR), suggestions, clarify, focus_used, and mode. Each is a persisted field first and a streamed event second, so reloads show the same evidence the live stream did. None of these fields exist in the current schema.
Make readiness check the data contract
A live process is not necessarily ready for retrieval. The server’s liveness endpoint can remain healthy while reporting DocumentDB down. Readiness is stricter: it requires a successful database ping, both required records search indexes, available Foundry Local, completed or disabled warmup, and free admission capacity.
Search-index creation occurs as part of ingestion preparation rather than unconditionally at boot. A fresh store can therefore be live but not retrieval-ready until those indexes exist. This distinction is useful operational truth, not an inconvenience to hide.
Security and lifecycle caveats
The internal store contains sensitive text, structured fields, vectors, conversations, schema samples, generated output, and encrypted source credentials. Vector data is still sensitive derived data. Protect storage files, backups, administrative access, and logs accordingly.
The development connection currently allows invalid TLS certificates. Production must validate the DocumentDB certificate authority and use non-default credentials. Source passwords are encrypted with AES-256-GCM before storage, but key rotation needs an explicit migration workflow.
Cross-collection operations are not universally atomic. Record publication is transactional, while source deletion cascades, conversation deletion, cleanup, and audit writes are application-managed sequences. Backup, restore, interrupted destructive operations, and retention therefore need deployment-specific testing.
Current semantic retrieval also lacks a general allow-listed metadata-filter contract applied consistently to vector and lexical searches. The RetrievalFilter and structured-cohort-then-semantic executor described above are planned in plans/new; the current hybrid cohort route falls back to ordinary semantic retrieval. These limits should remain visible in product claims.
Key takeaways
- Treat the internal RAG store as documents, vectors, metadata, and state – not merely a vector index.
- Keep External Clinical Databases authoritative and show the Internal DocumentDB Hybrid Store separately.
- Do not confuse local DocumentDB with Azure Cosmos DB or MongoDB Atlas because the wire protocol is compatible.
- Store row and generation provenance on every chunk.
- Publish refreshes through inactive staging and transactional active-generation cutover.
- Run vector and $text retrieval separately, then fuse ranks in Rust.
- Deduplicate by source row before truncation so chunk overlap does not monopolize evidence.
- Keep reranking and generation local but outside DocumentDB.
- Preserve execution provenance with each answer and report unvalidated thresholds or scale limits honestly.
For deeper implementation detail, see DocumentDB Architecture, Data Architecture, RAG Patterns, Retrieval, chat memory, and concurrency, and Verified Platform Notes.
Learn more on Microsoft Learn
- Design and evaluate a RAG solution
- Plan chunking and other enrichment work
- Understand Foundry Local’s place in an on-device RAG application
- Foundry Local SDK Reference – Rust on Windows
https://techcommunity.microsoft.com/blog/educatordeveloperblog/foundry-local-with-rust-from-catalog-discovery-to-in-process-streaming/4558179?previewMessage=true&WT.mc_id=MVP_406617
Continue the series
Previous: Part 2 – Foundry Local with Rust: From Catalog Discovery to In-Process Streaming.
Next: Part 4 – Resumable Health-Record Ingestion Without Sacrificing the Last-Known-Good Index.
The code behind this series is open source: github.com/kevin-gatimu/onprem-health-rag. It is a synthetic-data research prototype, not a clinical or compliance-ready system — issues and corrections are welcome.


