AI Foundry TPM Limits Prevent Stable, Mission-Critical Use

Schönwald, Alexander 120 Reputation points
2026-01-27T14:39:38.6866667+00:00

Hello everyone,

I have already raised concerns several times in this forum about how Microsoft handles the rollout and provisioning of models and the associated token limits in AI Foundry. The number of TPM provided is so limited that even a single developer can quickly hit the ceiling. Requests for additional TPM often take an unreasonably long time — if you receive a response at all. And by the time increased limits are finally granted, a newer version of the model is already available in Foundry, forcing the entire process to start over again.

Today’s outage of the Sweden region in Foundry highlights this issue once more. Due to the restricted TPM limits, we are forced to rely on a single region and we couldn't reroute the requests to other regions, forcing us to shut down our services until Sweden is fixed. It is simply not feasible to wait for approval requests every time we want to use another region. As a result, it becomes obvious that building stable and reliable services — especially for mission-critical production workloads — is practically impossible under these conditions.

Within our company, we rely heavily on AI Foundry as a core part of our AI strategy. However, the current approach to token allocation clearly needs to be reworked for professional and enterprise use cases. I fully understand limiting access for private individuals, but why are the same restrictions applied to paying companies and long-standing customers?

Foundry Tools
Foundry Tools

Formerly known as Azure AI Services or Azure Cognitive Services is a unified collection of prebuilt AI capabilities within the Microsoft Foundry platform


Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.