Several Microsoft Foundry Issues - Cache, open-source, pricing

Alexis Peralta 25 Reputation points
2026-08-14T08:28:47.8766667+00:00

Hi,

I am writing after working for a long time now with many Azure Foundry models and have come to the conclusion they are not reliable and difficult to work with.

Cache issues and lack of transparency

The only model whose cache seems to be working fine are OpenAI models, and maybe Claude. For the rest, you cannot rely on a proper cache. While I was getting 90% cache hits on GPT I was getting ~50-60% on DeepSeek. This only happening last month as DeepSeek had no cache at all while the main provider had it implemented for a long time now.

Pricing not updated

For every price change the main providers do, we have to wait for a long time to see that change reflected on our bills. While GPT 5.6 Luna price reductions happened the 30th of July we are still being billed the old pricing. Some examples of people complaining.

The main Microsoft blog mentions the prices are updated https://azure.microsoft.com/en-us/blog/gpt-5-6-now-available-in-microsoft-foundry/ but they are neither reflected on the Azure Foundry pricing page https://azure.microsoft.com/en-us/pricing/details/azure-openai/ nor our bills.

Why is this taking so long ? All other AI providers updated this in no more than 1 day...

Open-source models we don't even get to try

Many people use Foundry models to keep their data somehow private within their own Microsoft Tenant. That means they use Foundry as their main source of AI in their organisations. However, Foundry seems to have delegated providing the latest open-source models purely to Fireworks. If I use Fireworks my data leaves Microsoft and other data policies apply. So in terms of privacy it makes no sense to use those models compared to using other providers such as Openrouter or Opencode.

  • The latest DeepSeek V4 Flash update is still not present while, again, other providers are giving access at very little pricing, nearly instantly and with no cache issues.
  • How much time will we have to wait to get accesss to the new DeepSeek V4 Pro if we do not even have the new Flash ?
  • We have been waiting for Kimi K3 for a long time and is still not available, only through Fireworks...
  • Today GLM 5.3 released an I am not expecting to even see that model in Foundry.

Amount of time to update things

I am very worried Azure Foundry falls apart due to how long everything takes to be part of the speed the market is in. Only models that are deployed quickly are OpenAI's and Anthropic's but in terms of new models. Not for pricing as we could see.

Question

So my questions here are:

  • Am I right with my assumptions ?
  • Why is all of this happening ?
  • Is this a global feeling ?
  • Will this continue to be like this for a long time ?
Foundry Models
Foundry Models

A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference

0 comments No comments

Answer accepted by question author
Thanmayi Godithi 11,825 Reputation points Microsoft External Staff Moderator
2026-08-14T10:26:48.1733333+00:00

Hi Alexis Peralta,

Thanks for the detailed feedback. Some of the concerns you raise are reasonable, but a few assumptions need qualification based on the currently published Microsoft documentation.

  1. Open-source model availability

Microsoft does not publish a guaranteed onboarding timeline for every third-party or open-source model release in Microsoft Foundry. The documented model catalog and availability pages are the authoritative source for what is currently supported. Neither the Foundry documentation nor the FAQ provides a commitment that newly released models such as future DeepSeek, Kimi, or GLM versions will appear within a specific timeframe.

It is therefore fair to say that customers may perceive a gap between provider releases and availability in Foundry, but there is no public documentation that defines an expected rollout SLA for third-party model additions.

If a specific model is not currently available, the best path is typically to use the model catalog feedback or request mechanisms provided through Foundry rather than assuming a release date.

  1. Cache behavior

It is important not to assume that Azure OpenAI prompt-caching behavior applies identically to every provider hosted in Foundry.

Microsoft's prompt-caching documentation explicitly describes Azure OpenAI caching requirements, including minimum prompt size, prefix matching behavior, cache statistics, and GPT-5.6 cache enhancements such as prompt_cache_key and cache breakpoints. [Prompt cac...soft Learn | Learn.Microsoft.com], [GPT-5.6 no...ion agents | Azure.Microsoft.com]

The documentation does not state that the same requirements, matching logic, or cache-hit characteristics apply to DeepSeek or other partner/community models. Therefore, comparing a 90% cache-hit rate on GPT with a 50-60% cache-hit rate on another provider does not by itself prove a platform defect. Cache efficiency should be validated using the usage metrics returned by the model responses and the provider-specific documentation where available.

  1. Pricing and billing

The GPT-5.6 announcement published pricing for Sol, Terra, and Luna models and states that the models are generally available in Microsoft Foundry. [GPT-5.6 no...ion agents | Azure.Microsoft.com]

However, differences between an announcement, the public pricing page, and an individual customer's invoice cannot be diagnosed from a forum post alone. Microsoft documentation states that billing details are available through Cost Management + Billing and are not shown directly in the Foundry portal.

If usage that occurred after a pricing change appears to have been charged at a previous rate, the correct action is to open a billing support case and provide:

  • Subscription ID
  • Billing account or enrollment information
  • Deployment type
  • Azure region
  • Usage dates
  • Meter details
  • Expected versus observed charges

This enables the billing team to investigate the specific metering records.

  1. Data privacy and partner/community models

The concern about governance and compliance differences between first-party and partner/community offerings is understandable.

Microsoft documentation states that partner/community models are Non-Microsoft Products under the Product Terms.

At the same time, Microsoft also states that for partner/community models deployed through Foundry Models, Microsoft hosts the infrastructure, acts as the data processor, does not use prompts or outputs to train models, and does not share prompts or outputs with the model provider.

Customers should nevertheless review the specific terms, compliance posture, deployment type, and regional processing model associated with each provider offering to determine whether it meets their organization's requirements.

  1. Is this a "global feeling"?

There are clearly customers discussing concerns around model availability, pricing propagation, and rollout speed. However, neither Microsoft documentation nor public support channels provide evidence that these concerns represent the view of the broader customer base. Without published telemetry or survey data, it is not possible to conclude that this is a universal or global sentiment. This is ultimately anecdotal rather than documented.

Let me know if you have any additional questions. If the answer is helpful, please click "Accept Answer" and kindly upvote it. If you have extra questions about this answer, please click "Comment".

Was this answer helpful?

1 person found this answer helpful.
0 comments No comments

0 additional answers

Sort by: Newest

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.