RAG with Microsoft Foundry: Azure AI Search and Foundry IQ explained

Microsoft Foundry RAG connects generative models with enterprise knowledge. The benefit, however, does not arise from simply connecting a search index. Data quality, permissions, content division, search methods, and evaluation determine whether answers are relevant and up-to-date.

The guide places classical retrieval with Azure AI Search and agent-based approaches around Foundry IQ in context. It shows which architectural decisions should be made before the first index and how a team can permanently control answer quality and access limits.

Microsoft Foundry RAG begins with a domain-specific question

A RAG system needs a clear information task. Should it guide service technicians through manuals, find contract clauses, or explain internal policies? From this follow permissible sources, required metadata, update goals, and the type of source citation. "Search all company data" is neither a testable nor a secure task.

Cases are also defined in which no answer should be provided. If a source is not released, documents contradict each other, or the permission context is unclear, the system should report uncertainty. This rule prevents linguistic plausibility from being confused with domain certainty.

What retrieval augmented generation does

During retrieval, the application searches for relevant information sections and passes them along with the question to the model. The model then formulates an answer. The approach can update knowledge and make sources visible without retraining the base model for every document change.

RAG does not ensure a correct answer. The search can deliver irrelevant sections, miss relevant content, or favor outdated versions. The model can misrepresent context. Therefore, retrieval and generation should be measured separately. Only then can it be determined whether an error originates from the index, search query, or answer formation.

Distinguish classical and agent-based retrieval

Classical retrieval follows a largely predefined process: prepare the question, perform the search, select hits, and generate the answer. The process is well controllable and sufficient for many knowledge applications. Filters, hit count, and combinations of full-text and vector search can be tested specifically.

Agent-based retrieval breaks down more complex questions, plans multiple search steps, or combines sources. This can yield better results for multi-part tasks, but increases runtime, consumption, and variance. Foundry IQ is positioned in this context as a knowledge and retrieval layer. Its suitability depends on the specific service status, region, and operational requirements at the time of introduction.

Azure AI Search as a controllable search layer

Azure AI Search provides an index, full-text search, vector search, semantic functions, and filters. For enterprise RAG, it is particularly important that metadata and permission characteristics can be modeled in the index. The search service remains a separate architectural component with capacity, network, scalability, and costs.

The search configuration should not be hidden in the prompt. Index schema, search profile, filters, top-k, and weighting are versioned and tested against a reference dataset. The Foundry architecture shows how project, models, data access, and operational services interact.

Plan data intake with ownership and currency

Before technical intake, an owner is named for each source. They are responsible for content, release, storage, and currency. Without this role, the system indexes files but cannot professionally clarify contradictory or outdated statements.

The pipeline needs rules for new, changed, and deleted content. Deletions are particularly important: a removed document must not influence answers as an orphaned index entry. For each source, expected update time, error handling, and a comparison between origin and index are defined.

Split documents meaningfully

Chunking splits documents into searchable sections. Too small chunks lose context, too large dilute the relevant signal and consume context. Headings, tables, chapter boundaries, and document types should influence the split. A strict character length is only a starting value.

Each chunk needs stable metadata: document ID, version, section, source, validity, language, and permission characteristics. Overlap can preserve context but creates duplicates. The appropriate strategy is evaluated with real questions, not just technical metrics.

Consciously design index schema and metadata

An index contains more than text and vector. Filterable fields allow restrictions by tenant, area, document status, or validity date. Searchable titles, synonyms, and domain-specific characteristics increase accuracy. Uncontrolled growing metadata makes intake and query unnecessarily complex.

Schema changes require a migration path. Some changes require a new index and a controlled switch. An alias or a configuration outside the application code simplifies the transition. Beforehand, it is checked whether new and old versions meet the same permission and quality requirements.

Enforce permissions up to the hit

A RAG system must not retrieve content that the querying user is not allowed to see. The permission model must therefore be translated into a reliable filter before the search. A later instruction to the model to not mention confidential hits is not access protection.

Depending on the source, user, group, or attribute-based characteristics are taken into the index and applied at query time. Changes in Entra or the source system must take effect promptly. Test accounts with different roles check positive and negative access. Particularly critical collections can be in separate indexes or services.

Place hybrid search, vectors, and re-ranking in context

Vector search finds semantically similar content, while full-text search reliably finds exact terms, article numbers, or legal references. Hybrid search combines both signals. A re-ranker can then improve the order. More stages do not automatically mean better answers.

Tests should include different question types: natural paraphrases, exact product codes, ambiguous terms, and questions with missing answers. Evaluation checks whether relevant sections are in the top hits. Only then is it examined whether the model correctly formulates it.

Treat source attribution as a product requirement

A source citation must lead to the actually used document and preferably to the relevant section. Model-generated literature references are unsuitable. The application passes stable source information from retrieval and renders it separately from the generated answer.

Users should be able to recognize how current a source is and whether multiple documents are used. In the case of conflicting sources, the system must not smooth out the conflict. It names the discrepancy and refers to the responsible owners or the manual clarification process.

Assess Foundry IQ based on its use

Foundry IQ aims at a managed knowledge layer for agents and applications. For teams, this can simplify the reusable deployment of knowledge sources and agent-based retrieval. The decision should still depend on required data sources, permission model, regional availability, interfaces, and lifecycle.

Those already operating a controlled Azure AI Search setup should measure the migration benefit specifically. Those starting from scratch compare development effort, transparency, operational model, and product maturity. Preview functions should not be used untested in critical processes; their status is checked again before architecture approval and live deployment.

Build RAG evaluation in two levels

The retrieval level measures whether expected sources are found and ranked appropriately. Metrics like coverage or rank position require a domain-curated reference set. The generation level evaluates source fidelity, completeness, clarity, and behavior in the absence of evidence.

Automated evaluators speed up comparisons but do not replace a domain sample. For particularly risky answers, strict rules apply, such as no statement without a source or no mixing of tenants. The building of a Foundry agent should already consider these quality boundaries in tool design.

Connect operations, costs, and error diagnosis

Observations include intake errors, index age, search latency, empty hits, used sources, token consumption, and answer interruptions. A trace connects the question, search parameters, hit IDs, and model call. Sensitive content is minimized or masked so that diagnosis does not itself become a data risk.

Costs arise from data intake, embeddings, search capacity, models, and monitoring. Caches can reduce consumption, but must consider permissions and currency. An answer from an old cache must not bypass a revoked release.

Introduce in controlled stages

The first stage includes a few high-quality sources and a representative question set. After passing permission and quality tests, a limited user group follows. New sources are individually taken in and tested against the same checks. This ensures that any change causing a quality loss is clearly identifiable.

RAG supports the scaling of a AI transformation, when knowledge products are built with ownership, quality goals, and operations. The solid core is not a as large an index as possible, but a verifiable chain from released source to cited answer.

Build a gold set according to question types

The reference set contains exact terms, semantic paraphrases, multi-part questions, outdated documents, conflicting sources, and cases without a permissible answer. For each case, expected documents and unwanted hits are stored. This allows the team to objectively compare retrieval changes.

Questions come from real work tasks and are checked by source owners. New content or search profiles run against the same core. A separate security part checks user roles and manipulated documents, so quality gain does not compromise access limits.

Organize index switching without knowledge loss

Larger schema, chunking, or embedding changes are built in a new index. Data completeness, permission filters, and gold set results are checked before switching the application. The previous version remains available for a defined fallback period.

An alias or central configuration prevents hard-coded index names in multiple applications. After the switch, the team observes empty hits, latency, and source distribution. Only after a stable phase is the old index controlledly decommissioned.

Treat permission changes as a separate service goal

Quality and currency do not only refer to document text. If a user is removed from a group or an external release is ended, the search layer must consider this change within a defined time. The team therefore measures the delay from the source system through data intake, indexing, and to the filtered hit.

Negative test cases run regularly with different roles. A revoked document must not appear as an answer source or as an indirect hint. If the pipeline exceeds the service goal, affected sources are restricted or the agent switches to a secure mode. This control is particularly important when multiple data sources have different synchronization paths and permission models.

Are you planning a RAG solution with Microsoft Foundry?
We structure data sources, retrieval, permissions, evaluation, and the path to operations. Discuss RAG architecture

All articles