A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference
Hey,
The first thing I would check is the Fireworks Global Standard quota on the original subscription. Microsoft lists both FW-GLM-5.3-Flash and FW-DeepSeek-V4.1-Flash as Global Standard models.
Fireworks has separate Global Standard and Data Zone Standard quota pools, each shared across Fireworks models. Deploying GLM-5.3 successfully does not mean that sufficient Global Standard quota remains for Flash. An increase applied to the Data Zone pool would not resolve a Global Standard shortage.
I would suggest:
- Compare the successful and failed deployments using the same model version, region, deployment type, and requested TPM.
- Verify that the approved increase was applied specifically to Fireworks Global Standard on the original subscription in East US 2, and check the remaining quota after existing allocations.
- Retry with a smaller TPM allocation that fits the available quota.
If the correct pool has sufficient quota and the deployment still fails, follow up on the existing case or open an Azure support request, including the details and ask support to verify the effective quota and deployment eligibility on the original subscription.
Also, if the quota increase approval was recent, it may take up to 24h to reflect changes in your system.
References: