A unified Azure platform for creating and managing AI models, agents, and applications with built‑in enterprise security, monitoring, and governance
Yes, there is now a cleaner option for the authentication side of this architecture: OAuth identity passthrough for the MCP connection.
For your scenario, I would separate two concerns that are related but not identical:
- Who is invoking the MCP tool?
- How do I attribute and enforce the total cost generated by that user, including nested calls?
For #1, Microsoft Foundry Agent Service supports OAuth identity passthrough for MCP servers. Unlike a key-based connection, agent identity, or project managed identity, OAuth identity passthrough preserves the individual user's context. This avoids creating one agent or static API key per user.
With a custom MCP server, you can configure custom OAuth using your own Microsoft Entra app registration. The user authenticates and consents, and the MCP server receives authentication associated with that user's session. For an Azure Functions-hosted MCP server, Microsoft also documents OAuth identity passthrough as the production-oriented option when each user must authenticate individually and user context must persist.
For #2, I would not treat identity passthrough by itself as a complete cost-accounting solution.
If your flow is roughly:
User → Foundry Agent → MCP/Azure Function → LiteLLM → image/model endpoint
then propagate a stable user/correlation identity into your own metering layer after authentication. At the LiteLLM/gateway layer, record usage against that authenticated principal or an internal opaque user ID. This gives you a place to enforce per-user quotas/budgets for the downstream calls.
I would also keep authorization and accounting separate. The identity/token proves who the caller is and what they are allowed to access. Your gateway/metering system should decide how much of a resource or budget that identity may consume. Avoid using a user-supplied ID/header as the authoritative accounting identity unless it is validated against the authenticated principal.
There is one remaining boundary to be careful about: the Foundry agent's own base-model consumption and downstream MCP/LiteLLM consumption are generated at different layers. OAuth passthrough gives you user context at the MCP boundary, but I would not assume that this automatically provides a single native per-user budget covering both layers. If you require one hard budget across the entire chain, correlate telemetry from both layers using a request/user correlation ID and enforce the combined policy in your own control/metering plane.
So architecturally I would prefer:
User
→ Foundry Agent
→ OAuth identity passthrough
→ authenticated MCP/Azure Function
→ validated user/correlation identity
→ LiteLLM/API gateway
→ per-user quota + usage ledger
→ downstream model/tool
That scales much better than duplicating an agent per user, while keeping authentication, authorization, observability, and cost governance as explicit concerns.
Microsoft's documentation on MCP authentication is here:
https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/mcp-authentication
For a deeper treatment of the identity/RBAC side, Microsoft Learn also has an intermediate learning path covering authentication, authorization, managed identity, and RBAC for Azure AI workloads:
One additional security note: Microsoft documents tenant and trust restrictions around OAuth identity passthrough, including same-tenant requirements for the Foundry project and restrictions on sending Microsoft-audience tokens to custom/third-party MCP endpoints. For a custom MCP server, use an audience/app registration that you control rather than designing around forwarding a Microsoft service token downstream.