What is Microsoft Foundry? architecture, projects, and limits

What is Microsoft Foundry in a business context?

What is Microsoft Foundry? The platform bundles development, deployment, and governance of AI applications and agents in Azure. It connects model access, agent service, evaluation, tracing, and integrated services under a common resource and project model. This makes it especially suitable for teams that need more control than a ready-made productivity copilot or a pure low-code configuration provides.

Microsoft Foundry is the current name for the evolved platform that was previously known as Azure AI Studio or Azure AI Foundry. Older searches and documentation still use these names. For new architectures, it should be checked whether they already describe the current Foundry resource model or still older hub-based structures.

Problems the platform solves

A direct model call can quickly enable a prototype. However, for production use, identities, network boundaries, model deployments, tools, data access, testing, logging, and cost control come into play. If these tasks are solved anew in each project, inconsistent security and operations models emerge.

Foundry creates a common platform layer. Teams can compare models, develop agents, run evaluations, and connect telemetry with Azure services. This does not replace any domain-specific architectural decision, but reduces the number of self-built platform components. The article on scaling AI transformations describes why exactly these operational questions become relevant after a successful pilot.

Separate foundry resources from projects

The Foundry resource forms the overarching management boundary. There, common model access, policies, network, and access decisions are anchored. Projects create separate workspaces for applications, teams, or lifecycle stages. This separation is crucial when multiple business departments use the same platform.

A project is not a complete security strategy. Responsible parties must define which resources can be shared, how development and production should be separated, and who can deploy models or publish agents. Naming conventions, tags, and cost centers should already be established for the first production project.

Models are interchangeable architecture building blocks

Foundry provides a model catalog with models from different providers and deployment types. The selection is not based solely on general model quality. Context windows, language, latency, tool usage, data residency, cost, and availability in the target region influence the use case.

A model is treated as a versioned technical building block. Changes or version updates require regression testing because identical instructions can produce different results. For critical processes, the application should log model name, deployment, and relevant parameters in a traceable manner.

Agent Service connects model, instructions, and tools

An agent complements the model call with instructions, tools, conversation state, and optionally knowledge. Foundry Agent Service provides managed runtime functions. Prompt Agents can be configured declaratively; Hosted Agents allow custom code and frameworks when orchestration or runtime needs to be more customized.

The platform takes on hosting tasks, but not the domain-specific responsibility. Each tool requires limited permissions and a clearly defined interface. An agent that can modify data must handle errors, retries, confirmation, and cancellation just like any other production application.

Tools and integrated Azure services

Tools connect the agent with search, functions, APIs, or business systems. Common building blocks are Azure AI Search, Azure Functions, Logic Apps, OpenAPI interfaces, or MCP servers. Depending on the use case, storage, Key Vault, and Application Insights are added.

These components should not be connected as a random collection of functions. Every connection increases the scope of permissions, the area of potential errors, and costs. An interface is only accepted if its purpose, authentication, owner, and behavior in case of failure are documented.

Identity and RBAC as the core of the architecture

Developers, operators, applications, and agents need different rights. Microsoft Entra ID and Azure RBAC control who can manage resources, deploy models, or call endpoints. Production applications should, if possible, use managed identities instead of distributing keys in configurations or source code.

Roles are assigned at the smallest meaningful level. A developer does not automatically need owner rights at the subscription level, and an agent does not need access to all data sources of a project. Regular reviews remove outdated group memberships and check privileged changes.

Plan networks, data flows, and secrets

For sensitive workloads, a mere login is often not enough. Private endpoints, virtual networks, controlled outgoing traffic, and regional availability determine how data flows between application, model, and integrated services. The chosen architecture must also consider development tools, CI/CD, and support access.

Keys and connection data belong in an appropriate secret store. Logs should not include complete prompts, documents, or personal content if they are not required for diagnostics and evaluation. Data minimization applies to telemetry and test datasets as well.

Evaluation and observability belong to the product

With generative AI, a successful HTTP status is not sufficient as evidence of quality. Evaluations measure relevance, groundedness, task completion, tool calls, and security aspects. Traces show which model and tool steps led to a response.

Application Insights and OpenTelemetry can make operational data, latencies, and error paths visible. Access to conversation content is limited and logged. Thresholds and alerts must be linked to a support process; a dashboard without a clear response process does not improve operations.

When Microsoft Foundry is useful

Foundry is suitable when a team wants to develop its own AI applications or agents with controlled runtime, model selection, APIs, and Azure governance. Typical triggers are custom user interfaces, complex retrieval, custom orchestration, private networks, or integration into existing software products.

For a simple personal knowledge agent in Microsoft 365, the platform can be disproportionately complex. There, Agent Builder or Copilot Studio are often closer to the needed channel and reduce development effort. The decision is similar to the general trade-off between low-code and pro-code on Microsoft 365: control is valuable, but also incurs operational responsibility.

Operational costs arise in multiple services

The Foundry resource alone does not describe the total cost. Model tokens, agent runtime, search, storage, network, monitoring, and possibly additional AI services are billed according to their respective models. A cost model must therefore map the complete data and call path.

Budgets, tags, and warnings are prepared per project or cost center. Load tests consider not only the average but also long queries, repeated tool calls, and extensive retrieval contexts. Unused deployments and test resources are systematically removed.

Use foundry as a platform rather than a playground

Microsoft Foundry creates a common framework for models, agents, evaluation, and Azure operations. The benefit arises only when resource and project boundaries, identities, data paths, and responsibilities are consciously designed. A playground result is the beginning of development, not its production release.

The platform is especially robust when multiple teams need repeatable standards. Common guidelines for projects, models, telemetry, secrets, and releases prevent each agent from creating its separate infrastructure. The technical freedom of pro-code can be combined with a controlled operations model.

Align project boundaries with responsibility and risk

A project should bundle a traceable group of applications and responsible parties. Separate data classifications, operational models, or cost centers speak for separate projects. Placing every small experimental idea into its own project, however, creates unnecessary role, network, and monitoring work.

The decision is documented as a standard: allowed regions, naming schema, tags, owners, connection types, and transition to production. This allows teams to start independently without having to renegotiate basic security decisions for every initiative.

Check the network path from user to data source

Architecture includes not only Foundry endpoints. The complete path includes the calling application, identity service, model, search or storage service, tools, and telemetry. For each connection, it is clarified whether it is public, via service endpoints, or private connection, and how name resolution and firewall rules function.

Private networking increases control but brings dependencies in DNS, routing, and development access. An architecture test should therefore not only check a model call from the studio. It must use the same path that the production application will later use.

Separate the platform team from the product team

The platform team is responsible for guiding principles, identity patterns, network, model releases, observability, and cost standards. The product team is responsible for domain-specific utility, prompts, tools, evaluation data, and support for the specific use case. Both sides share responsibility for production release.

This separation prevents two extremes: a central platform that must guess domain-specific quality, and product teams that must build security and operational foundations anew. A short service catalog describes what the platform provides and what the product team must deliver as proof.

Use a reference architecture as a verifiable contract

The reference architecture should be more than a diagram. It names allowed resource types, identity patterns, network zones, logging standard, model approval, data classification, and production criteria. For each requirement, it is clear whether it is technically enforced, checked in a pipeline, or approved organizationally. Teams thus recognize early which parts the platform provides and which they must implement themselves.

Departures are handled through a time-limited decision process. The request describes business reason, risk, alternative controls, owner, and return to standard. Insights from projects regularly flow back into the reference architecture. This way, it remains connected to new Foundry features and real operational problems, without every product team inventing its own security and operational architecture.

Confirm region and service availability early

Models, agent functions, and dependent Azure services are not identically available in every region. Before the target architecture, data residency, latency, capacity, network options, and current product status are jointly checked. A pilot in a convenient region must not become the production standard unnoticed if domain-specific or regulatory requirements demand another region. The decision also includes handling regional bottlenecks and new model versions. This ensures that the chosen dependency is traceable and that any change that requires a new architecture approval is clearly documented.

Assess Microsoft Foundry for a specific use case
If model selection, agent runtime, data access, and Azure operations must be planned together, the appropriate Foundry architecture can be defined before the pilot. Discuss your Foundry architecture

All articles