Evaluating a Microsoft Foundry Agent involves assessing response quality, tool behavior, security, and operations together. A good answer in a few demonstration questions is not enough. Production readiness requires representative datasets, measurable release thresholds, and signals that make errors during operations visible.
The guide combines offline evaluation, tracing with OpenTelemetry, Application Insights, content safety, and red teaming into an operating model. It also shows when a deployment must be stopped or rolled back to a known version.
Evaluating a Microsoft Foundry agent: establish a quality agreement
Before metrics, there is a quality agreement. It describes the tasks the agent performs, which errors are tolerable, and which events are considered critical. The business team, product ownership, IT, and information security must understand the same boundaries. Otherwise, each team will optimize for different goals.
The agreement distinguishes response quality, source reference, tool execution, access protection, security, and user experience. For each dimension, evaluation method, minimum value, and escalation are defined. Hard rules like "no action without required approval" must not disappear into an average value.
Derive evaluation data from real tasks
A useful dataset includes typical, rare, difficult, and inappropriate queries. It covers various formulations, languages, document types, and user roles. Production data may only be used under clarified data protection and retention rules; often, cleaned or synthetically augmented cases are useful.
Each case includes input, relevant context, expected behavior, and evaluation criteria. For knowledge questions, reference sources can be included, and for actions, expected tool parameters and status changes are specified. The version and origin of the dataset remain traceable.
Use rubrics instead of vague letter grades
A rubric describes what a rating specifically means. Domain accuracy, completeness, source fidelity, and clarity are evaluated separately. For tool calls, the selection, parameters, order, permissions, and result handling are considered. This turns "it looks good" into a repeatable evaluation.
Automated evaluators can compare large variations but also introduce their own uncertainty. A portion curated by subject-matter experts is regularly evaluated by humans and serves as a control point. Differences between automated and manual evaluations are investigated, rather than selecting the more convenient value.
Rate retrieval and response separately
In RAG applications, a wrong answer can arise from missing hits or from poor processing of good hits. Retrieval metrics check whether expected sources are found and ranked high enough. Answer metrics check whether the output is covered by the provided context.
This separation shortens the diagnosis. A prompt change cannot fix an incomplete index, and a larger model does not replace permission filters. The contribution to Microsoft Foundry RAG describes data ingestion, search layer, and source reference in detail.
Test tool calls as transactions
An agent may appear linguistically correct but still use the wrong tool or incorrect parameters. Tests check successful flows, validation errors, timeouts, duplicate calls, and unavailable dependencies. Writing actions require additional approval and retry tests.
The expected business state matters more than the formulated answer. If the agent reports success, but the API has failed, the test fails. Idempotency keys, status checks, and structured errors prevent duplicate or hidden changes.
Test permissions and tenant boundaries
Test identities with different roles ensure that retrieval and tools only provide allowed data. Negative tests deliberately attempt to access other users’ datasets, hidden fields, or administrative functions. Indirect hints in error messages and source references are also considered.
Tests run again after changes to identity, tool, index, or group model. A test that passed once does not remain valid indefinitely. The creation of a Foundry Agent should therefore document permission model and tool boundaries as versioned architectural decisions.
Test prompt injection and manipulated sources
Red-Teaming datasets include instructions in documents, disguised requests, role switches, data exfiltration attempts, and overly long inputs. The goal is not only to produce a secure text response. It is critical that the agent does not call extended tools or transfer confidential content to other channels.
Protective measures are on multiple levels: minimal rights, trusted tool schemas, separation of instruction and data, input and output controls, and human approvals. A prompt alone is not a security boundary. Found attacks are added as permanent regression tests.
Configure content safety according to risk
Content filters and security services can detect harmful inputs or outputs. Thresholds must align with the use case. An internal workplace protection agent processes different terms than a public support agent. Too strict filters block legitimate tasks, while too loose filters allow risks to pass through.
Blocks are measured by category and version, without unnecessarily storing sensitive content. Business-allowed exceptions require a documented process. Changes to filters are tested using the same reference dataset as model or prompt changes.
Structure tracing with OpenTelemetry
A trace connects incoming requests, agent steps, model calls, retrieval, and tools. OpenTelemetry provides a cross-vendor model for traces, metrics, and logs. A correlation ID enables the reconstruction of an error across multiple services.
Spans should include duration, status, model and tool version, and consumption. Full prompts, documents, or tool responses do not automatically belong in telemetry. The scope of data collection, masking, and access are determined based on diagnostic needs, data privacy, and retention goals.
Use Application Insights as an operational view
Application Insights can capture telemetry, represent dependencies, and enable queries and alerts. Relevant views show end-to-end latency, error rates, token consumption, tool failures, and affected versions. A dashboard per team without a shared definition complicates comparison.
Technical signals are connected to business outcomes. These include successful case closures, handoffs to humans, or corrected responses. If latency increases while quality remains the same, the response is different than when increasing permission errors.
Online monitoring detects drift and outliers
Offline tests check known cases; in production, new formulations, data changes, and load patterns occur. Samples and quality indicators that minimize data collection show whether responses are getting worse. Frequent follow-ups, empty retrieval results, or unexpected tool sequences are early warning signals.
Production cases are collected according to a fixed procedure into the evaluation dataset. Personal data is removed and business references are added. This cycle prevents tests from permanently only representing the original pilot world.
Make thresholds and alerts actionable
An alert requires a signal, threshold, time period, recipient, and action. Examples include a sharp increase in failed tool calls, unusual token consumption, or a critical security violation. Warning and shutdown thresholds are defined separately.
Each alert is assigned a runbook. It outlines initial checks, responsible parties, communication channels, and decision authority. Alert fatigue is mitigated through tiered severity levels, deduplication, and regular reviews. Unaddressed warnings are considered an operational defect.
Put changes through an evaluation gate
Prompt, model, tool, index, filter, and runtime can change behavior. Every relevant change creates a new, uniquely named version and goes through automated and business tests. A comparison against active production shows improvement and regressions.
Releases are based on defined thresholds. A cheaper model may not go live if critical tool errors increase. Conversely, a small stylistic improvement does not justify significantly higher latency. The decision is documented with measurement values and approval by the responsible owner.
Prepare rollback technically and organizationally
A rollback requires that previous model, prompt, tool, and configuration versions are known and deployable. Data or schema changes need compatibility rules. When using writing tools, additional business-specific correction of already executed actions may be necessary.
The rollback path is tested: switching to a known version, performing read-only operations, deactivating individual tools, or handing over to humans. The exercise demonstrates whether monitoring, permissions, and communication actually function. A theoretical switch without testing is not a reliable contingency plan.
Bring lessons from incidents back into product development
After an incident, the cause, scope, affected versions, and actions taken are documented. The focus is on system improvement: new regression tests, tighter permissions, more robust tool validation, or better alerts. Pure prompt corrections often fall short when dealing with architectural errors.
A production-ready agent platform connects evaluation with the scaling of the AI transformation. Common categories, telemetry standards, and release processes reduce effort for further use cases. At the same time, each agent retains its business-specific risk boundaries.
The controlled path to operations
Before go-live, data set, minimum values, security checks, dashboards, alerts, runbooks, owners, and rollback path are complete. A limited user group and tight monitoring reduce the impact of unknown errors. Expansion occurs only when technical and business signals are stable.
Evaluation is thus not a one-time test step. It forms the feedback loop for every change and every incident. A Foundry agent remains manageable when quality is measurable, behavior is traceable, and a secure response is always possible.
Define service goals for agent workflows
A service goal can combine successful operations, maximum end-to-end latency, allowable tool errors, and the proportion of unsupported answers. For critical security incidents, a zero-tolerance policy applies instead of an average. Goals are differentiated by question type or process risk.
The measurement considers multi-step workflows. A fast model call does not help if the downstream tool regularly fails. Business and technical goals appear in the same operational view and lead to clearly named actions.
Establish regular operational reviews
Weekly, notable errors, costs, and alerts are reviewed; monthly, quality samples, model changes, and open risks are followed. At longer intervals, data sets, categories, storage, and ownership are checked. The frequency depends on criticality and change rate.
The review ends with decisions: add test case, adjust threshold, improve tool, reduce permissions, or reset version. Documented decisions create traceability and prevent known quality issues from being permanently considered unavoidable.
Rules for data privacy and access to evaluation data
Evaluation datasets, traces, and transcripts may contain confidential inputs, model responses, and tool results. For each artifact, purpose, allowed fields, masking, access, storage location, and retention period are defined. Developers do not automatically have insight into full production conversations. For many diagnoses, metadata, IDs, and selectively released samples suffice.
Access is logged and regularly reviewed. Cases for regression testing are cleaned or synthetically recreated before permanent inclusion. In deletion and disclosure processes, it must be known which systems retain copies. These rules protect individuals and simultaneously improve data quality: A curated evaluation set is more reproducible than an uncontrolled archive of all conversations.
Would you like to make a Foundry agent production-ready?
We develop evaluation frameworks, telemetry, security audits, and a reliable operations process. Prepare agent operations