An Azure service that provides artificial intelligence algorithms that detect, recognize, and analyze human faces in images.
Yes, production agents need a slightly different testing model from traditional deterministic applications. The pattern I've found most useful is to separate it into pre-production evaluation, runtime tracing, and continuous production monitoring.
Before deployment, maintain a representative evaluation dataset containing normal requests, edge cases, invalid inputs, tool-selection scenarios, and safety cases. Microsoft Foundry supports agent-specific evaluation, so you can measure task completion and tool-call accuracy alongside quality and safety metrics. You can also use those evaluations as quality gates before promoting a new version.
For tool-enabled agents, don't test only the final answer. Test the execution path:
User request → agent decision → tool selected → tool arguments → workflow/API result → final response
That's where tracing becomes important. Foundry tracing integrates with Azure Monitor Application Insights and can capture the sequence of agent actions, tool calls, latency, exceptions, and other execution details. That makes it much easier to diagnose cases where the final answer looks wrong because the agent selected the wrong tool or supplied incorrect arguments.
In production, monitor two categories of signals: traditional operational metrics such as latency, errors, and token consumption, and AI-specific quality signals through continuous evaluation of sampled production interactions. Foundry's Agent Monitoring Dashboard integrates these with Application Insights, and continuous evaluation can surface quality/safety degradation from real traffic.
For changes to prompts, models, tools, or agent instructions, I would treat the agent as a versioned application artifact:
Version change → regression evaluation → compare against baseline → deploy → monitor → evaluate sampled production traffic → rollback if necessary.
I recommend saving meaningful agent changes as versions, debugging them with tracing, running repeatable regression evaluations, and then monitoring the published version in production.
One important production consideration is privacy. Traces can contain prompts, outputs, tool arguments, and results, so redact or exclude credentials and sensitive information before telemetry reaches Application Insights, and apply appropriate access and retention controls to the trace data.
So don't try to make an AI agent completely deterministic. Instead, define measurable acceptable behavior, continuously evaluate it, and make the agent's decision/tool path observable enough to investigate failures.
Useful Microsoft references:
Microsoft Foundry - Observability in Generative AI
Help make this community better for everyone: If this answer helped or resolved your issue, please accept it or upvote it. If not, share more details in a comment so we can continue the discussion and find the right solution. Thank you.