A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference
Hi Alexis Peralta,
Thanks for the detailed feedback. Some of the concerns you raise are reasonable, but a few assumptions need qualification based on the currently published Microsoft documentation.
- Open-source model availability
Microsoft does not publish a guaranteed onboarding timeline for every third-party or open-source model release in Microsoft Foundry. The documented model catalog and availability pages are the authoritative source for what is currently supported. Neither the Foundry documentation nor the FAQ provides a commitment that newly released models such as future DeepSeek, Kimi, or GLM versions will appear within a specific timeframe.
It is therefore fair to say that customers may perceive a gap between provider releases and availability in Foundry, but there is no public documentation that defines an expected rollout SLA for third-party model additions.
If a specific model is not currently available, the best path is typically to use the model catalog feedback or request mechanisms provided through Foundry rather than assuming a release date.
- Cache behavior
It is important not to assume that Azure OpenAI prompt-caching behavior applies identically to every provider hosted in Foundry.
Microsoft's prompt-caching documentation explicitly describes Azure OpenAI caching requirements, including minimum prompt size, prefix matching behavior, cache statistics, and GPT-5.6 cache enhancements such as prompt_cache_key and cache breakpoints. [Prompt cac...soft Learn | Learn.Microsoft.com], [GPT-5.6 no...ion agents | Azure.Microsoft.com]
The documentation does not state that the same requirements, matching logic, or cache-hit characteristics apply to DeepSeek or other partner/community models. Therefore, comparing a 90% cache-hit rate on GPT with a 50-60% cache-hit rate on another provider does not by itself prove a platform defect. Cache efficiency should be validated using the usage metrics returned by the model responses and the provider-specific documentation where available.
- Pricing and billing
The GPT-5.6 announcement published pricing for Sol, Terra, and Luna models and states that the models are generally available in Microsoft Foundry. [GPT-5.6 no...ion agents | Azure.Microsoft.com]
However, differences between an announcement, the public pricing page, and an individual customer's invoice cannot be diagnosed from a forum post alone. Microsoft documentation states that billing details are available through Cost Management + Billing and are not shown directly in the Foundry portal.
If usage that occurred after a pricing change appears to have been charged at a previous rate, the correct action is to open a billing support case and provide:
- Subscription ID
- Billing account or enrollment information
- Deployment type
- Azure region
- Usage dates
- Meter details
- Expected versus observed charges
This enables the billing team to investigate the specific metering records.
- Data privacy and partner/community models
The concern about governance and compliance differences between first-party and partner/community offerings is understandable.
Microsoft documentation states that partner/community models are Non-Microsoft Products under the Product Terms.
At the same time, Microsoft also states that for partner/community models deployed through Foundry Models, Microsoft hosts the infrastructure, acts as the data processor, does not use prompts or outputs to train models, and does not share prompts or outputs with the model provider.
Customers should nevertheless review the specific terms, compliance posture, deployment type, and regional processing model associated with each provider offering to determine whether it meets their organization's requirements.
- Is this a "global feeling"?
There are clearly customers discussing concerns around model availability, pricing propagation, and rollout speed. However, neither Microsoft documentation nor public support channels provide evidence that these concerns represent the view of the broader customer base. Without published telemetry or survey data, it is not possible to conclude that this is a universal or global sentiment. This is ultimately anecdotal rather than documented.
Let me know if you have any additional questions. If the answer is helpful, please click "Accept Answer" and kindly upvote it. If you have extra questions about this answer, please click "Comment".