Creating a Microsoft Foundry agent does not start with choosing a language model. First, define its task, allowed actions, data access, and responsibilities. Then decide whether a declaratively configured Prompt Agent is enough or whether a Hosted Agent with custom runtime code makes sense.
This guide covers the Foundry resource and project, model, tools, and identity, then deployment, endpoints, and operations. It is for technical leads who want to move beyond demonstrating an agent and integrate it securely into an enterprise architecture.
Create a Microsoft Foundry agent: define its purpose and limits
An agent needs a precisely formulated task. 'Supporting procurement' is too broad for that. If the agent checks supplier documents, answers questions about contracts, or creates purchase requisitions, for each use case, expected inputs, allowed data sources, possible actions, and a measurable outcome should be documented.
Equally important are negative rules. These include data the agent is not allowed to see, decisions that must be reserved for people, and actions that require approval. These boundaries later determine system instructions, tool scope, identity model, and test cases. The architecture of Microsoft Foundry provides the technical framework for this.
Define resource and project boundaries
The Foundry resource bundles central management and security functions. Projects create separate workspaces for applications, deployments, and involved teams. A meaningful project boundary is based on responsibility, data risk, and lifecycle, not on every individual idea. Too many projects increase operational effort; one project for everything makes it harder to allocate costs and permissions.
For development, testing, and production, separate environments or at least clearly separated deployments with their own configurations should exist. Names, regions, tags, budget assignment, and responsible persons should be included in a standard before the first deployment. This ensures that it is clear which endpoint serves which process.
Choose a model based on task, not popularity
The model selection begins with quality requirements: Does the agent need to process long documents, generate structured outputs, reliably select tools, or handle multiple languages? These are joined by latency, available region, throughput, content filtering, context window, and cost. A large model is not automatically the most cost-effective choice.
For a reliable comparison, a representative test set is needed. Multiple candidates answer the same cases under identical instructions and tools. Evaluation focuses on domain-specific accuracy, source reference, format stability, tool selection, runtime, and consumption. Only the measurement values justify a model decision.
When a Prompt Agent is appropriate
A Prompt Agent is suitable when behavior can largely be described by instructions, knowledge, and standardized tools. Typical cases include research in approved sources, summarization, classification, or the execution of fewer clearly defined actions. The platform takes on a large part of the orchestration and accelerates changes.
The smaller codebase makes a pilot easier, but does not eliminate the need for architecture work. Instruction versioning, access to tools, error handling, and evaluation remain necessary. As soon as complex state, custom libraries, or very specific process control are required, the declarative model reaches its practical limits.
When a Hosted Agent becomes useful
Hosted Agents allow custom agent logic and frameworks in a managed runtime. This is interesting when the workflow coordinates multiple specialized components, contains extensive business logic, or requires dependencies that a Prompt Agent cannot represent. Deterministic preprocessing and postprocessing can also be more tightly controlled this way.
The additional degree of freedom creates responsibility for code quality, dependencies, secrets, scaling, and error patterns. A Hosted Agent should therefore not be chosen simply because Pro-Code generally appears more powerful. The key question is whether the use case actually needs the additional control and whether the team can maintain it permanently.
Comparing Prompt Agent and Hosted Agent
Six questions can help with the decision: Can the workflow be described declaratively? Are there custom libraries? Is state managed outside of a session? How complex is tool control? What level of runtime control is required? Who takes on maintenance and on-call support? A Prompt Agent wins in simplicity and flexibility, while a Hosted Agent excels in specialized logic and runtime control.
A step-by-step approach reduces risk. The team starts by implementing the smallest fully functional workflow and measures where real limits occur. Only proven gaps justify the switch. The comparison Copilot Studio versus Microsoft Foundry is also helpful when considering Low-Code as a platform option.
Designing tools as controlled capabilities
Tools connect the agent to search systems, APIs, or business processes. Every capability needs a clear description, a concise input schema, and validatable outputs. Overlapping tool descriptions lead to the model making unreliable choices. Business terms and examples are often more effective than technical abbreviations.
Writing tools requires additional guidelines: Inputs are validated server-side, repeated calls must be handled, and critical changes require confirmation or approval. Errors are returned to the agent in a structured format. A friendly text message without an error code complicates diagnosis and can trigger false success messages.
Identity follows the business task
An agent should not work with broad application rights by default. First, decide whether an action is performed in the name of the logged-in user or as a separate workload. User delegation preserves the user’s permission context; an agent identity is suitable for clearly defined, user-independent tasks.
When application rights are involved, the required permissions are justified individually and administratively approved. Tips on control are provided by the Review of Entra App Permissions and Admin Consent. Secrets should be avoided or managed in a secure service; where possible, managed identities and certificates are preferable.
Versioning and securing instructions
The system instruction describes role, allowable scope, source rules, output format, and behavior in case of uncertainty. It should not contain any secret values and be treated as a versioned artifact. Changes should go through review and testing because a small formulation can change tool selection or response boundaries.
Prompt injection cannot be resolved by a single sentence like "Ignore external instructions." Protection is achieved through minimal permissions, separation of data and control instructions, validation of tool parameters, content controls, and testing with intentionally manipulated documents. The agent must clearly reject unauthorized requests.
Plan deployment and endpoint
A deployment fixes the model, capacity, and configuration for a usable version. The calling service should not depend on random studio configurations, but instead use a clearly named endpoint with controlled authentication. Configuration values for development, testing, and production are managed separately.
Before moving to production, test for load, time limits, quotas, and repeated behavior. The application needs a clear response when the model or tool is unreachable. Queues, limited retries, and a manual fallback process prevent an external bottleneck from becoming an uncontrolled business risk.
Assess before release
The test set should represent normal, difficult, and unauthorized cases. In addition to answer quality, source attribution, tool selection, parameters, permission limits, and termination behavior are evaluated. For important fields, expected values or categories must be defined. A single average value can hide critical individual violations.
A release requires thresholds. For example, no other customer’s account number may be disclosed, while a small linguistic deviation is tolerable. The business team, IT, and information security must know which criteria they accept. Test results are stored together with prompt, model, and tool versions.
Observability for model and tools
In operations, request, runtime, model consumption, tool calls, error class, and result status should be correlated. Personal or confidential content must not uncontrollably appear in logs. A common trace identifier connects the application, agent, and downstream API without duplicating the entire conversation content.
Dashboards show error rates, latency, consumption, and business success indicators. Warnings need a clear response: Who checks an increase in faulty tool calls, who locks a deployment, and who informs affected process owners? Monitoring without responsibility only generates data.
Operations handover and fallback
Before handover, owners for product, model configuration, tool APIs, permissions, and support are named. Known limits, quotas, data sources, dependencies, and emergency contacts are documented. The monitoring of app secrets and certificates also belongs in the lifecycle.
A fallback could be a previous agent version, read-only mode, or forwarding to a human. It must be technically prepared and tested. If shutdown is only considered during an incident, dependencies and impacts are often unclear.
From a reliable pilot to a production service
A good pilot handles a real but limited process with representative users. It provides measurements on quality, time saved, errors, and operational costs. The results decide whether scope, model, or platform needs adjustment. A successful demo is not yet proof of production readiness.
A production service takes shape when architecture and organization align: a bounded task, minimal rights, traceable versions, verified tools, defined quality thresholds, and an operations team that can be reached. Then the agent experiment becomes a controllable application.
Production check before endpoint approval
Before approval, model and tool versions, minimal rights, network paths, quotas, cost warnings, traces, evaluation results, and fallback paths are confirmed in a checklist. The calling application uses the same identity and endpoint as in later operations. A test alone from the studio is not sufficient.
Additionally, a controlled failure is tested: A tool does not respond, the model quota is exhausted, or an identity loses its permissions. The agent must stop, provide a traceable status, and must not report any incomplete action as a success. Only this test shows whether architecture and runbook align.
Stabilize API contract for calling applications
The agent endpoint should have a clearly versioned input, output, error structure, and correlation ID. Calling applications must not depend on internal intermediate steps or freely changing text formats. Structured result fields are validated server-side; human-readable text remains a representation, not a technical status.
If a tool or agent framework changes, the outer contract can remain stable. Incompatible changes receive a new version and a transition period. Load and security tests occur through exactly this interface. This decouples the user interface and agent implementation, simplifies rollback, and prevents an internal experiment from silently breaking multiple production clients.
Would you like to design a reliable Microsoft Foundry agent?
We support you with architecture, identity, tool design, evaluation, and production transition. Plan your Foundry agent with us