Foundry/OpenAI TPM slider jumps from 1K to 2.01M, cannot set capacity within 200K quota which it said is available.

Ashika Hussain 0 Reputation points
2026-09-18T10:25:22.9+00:00

Hi,

When creating or editing a Standard deployment of gpt-4.1-mini in Azure OpenAI. The Tokens per Minute slider does not allow mid-range values.

What I see

  • Available quota for the model/region is about 200K TPM (usage 0).
  • Slider at minimum: 1K TPM.
  • A slight move jumps to about 2.01M TPM (previously I also saw a jump toward hundreds of thousands / millions).
  • I cannot select something like 100K or 200K in the UI.

Is this expected behaviour? Is there any other work around than the slider?

Azure OpenAI in Foundry Models
0 comments No comments

1 answer

Sort by: Oldest
  1. Andriy Bilous 12,276 Reputation points MVP
    2026-09-21T04:28:52.7333333+00:00

    Hello Ashika Hussain

    I see that slider jumping from 1K directly to approximately 2.01M is not documented by Microsoft as expected behavior. I would treat it as a Foundry portal/UI issue unless the backend APIs show otherwise.

    For a Standard Azure OpenAI deployment, Microsoft documents that:

    • sku.capacity = 1 means 1,000 TPM
    • sku.capacity = 100 means 100K TPM
    • sku.capacity = 200 means 200K TPM

    So you can bypass the slider with Azure CLI:

    az cognitiveservices account deployment create -g <RG> -n <resource> --deployment-name <deployment> --model-name gpt-4.1-mini --model-version "2025-04-14" --model-format OpenAI --sku-name Standard --sku-capacity 100

    Microsoft documentation: Automate Azure OpenAI deployments with quota

    Having 200K quota available does not necessarily mean 200K deployment capacity is currently available in that region. Microsoft now exposes these separately:

    • Usages API → assigned quota and current consumption.
    • Model Capacities API → how much capacity can actually be deployed for that model, version, region and SKU.

    You can check capacity with:

    GET https://management.azure.com/subscriptions/<subscription-id>/providers/Microsoft.CognitiveServices/modelCapacities?api-version=2024-10-01&modelFormat=OpenAI&modelName=gpt-4.1-mini&modelVersion=2025-04-14

    Official documentation: Manage Azure OpenAI quota and check capacity

    Microsoft introduced quota tiers and subscription-level quota changes in 2026, so generic quota values in documentation might differ from the allocation shown for a specific subscription.

    So using CLI/REST with sku-capacity 100 or 200 is the supported workaround. If the APIs show at least that much quota and deployment capacity but creation still fails or the portal continues jumping from 1K to 2.01M capture the API response/error and open an Azure Support request, as that would indicate a portal or backend quota/capacity inconsistency rather than an expected slider restriction.

    Azure OpenAI quotas and limits

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.