A unified Azure platform for creating and managing AI models, agents, and applications with built‑in enterprise security, monitoring, and governance
The key point is that "cached_tokens" in the model response and the Azure billing meter are related signals, but they are not sufficient by themselves to prove what rate Azure actually applied to those tokens.
I would troubleshoot this by first identifying exactly which DeepSeek V4 Pro offer/deployment you are using.
Microsoft Foundry currently exposes DeepSeek V4 Pro through different provider/deployment paths, including Fireworks-hosted models. The caching behavior and billing contract should therefore be checked against the specific offer rather than assuming that every deployment named "DeepSeek V4 Pro" has identical caching semantics.
If this is the Fireworks DeepSeek V4 Pro offer
Microsoft's current Foundry documentation explicitly documents prompt-cache optimization for Fireworks models.
It recommends maintaining affinity for related requests using one of:
x-session-affinity
or the request parameter:
user
or:
prompt_cache_key
with "prompt_cache_key" taking priority over "user".
So if your deployment is "FW-DeepSeek-V4-Pro", I would first make sure your repeated requests are using stable cache-affinity information and an identical reusable prompt prefix.
However, your observation is actually more interesting than a simple cache miss, because you are already receiving something like:
"input_tokens_details": {
"cached_tokens": ...
}
with 50-80% of the input reportedly cached.
That suggests the provider/model layer is reporting cached-token activity.
The next question is therefore not:
«"Is caching happening?"»
but:
«"Which Azure billing meter is receiving those cached tokens?"»
Check the Azure meter, not only the API usage object
I would compare a controlled test against Azure Cost Management rather than calculate cost only from total API input tokens.
Run two controlled workloads:
Test A:
unique/non-repeated prompt prefixes
Test B:
identical long prompt prefix
stable session/cache affinity
same model/deployment
similar output length
Record for every request:
timestamp
deployment/model ID
total input tokens
cached_tokens
output tokens
request/correlation ID if available
Then inspect the Azure usage/cost data for the same interval and determine whether usage is being split into separate normal-input and cached-input meters.
Microsoft's Foundry pricing pages distinguish Input, Cached Input, and Output for supported DeepSeek/Fireworks offerings.
Therefore, if the API consistently reports substantial cached-token usage but Cost Management records those same tokens exclusively against the normal input meter, that is no longer simply a prompt-cache hit-rate problem. It becomes a billing/metering discrepancy worth escalating.
Do not calculate the expected Azure bill from DeepSeek's direct API pricing
Another important distinction is that the Foundry deployment is billed according to the Microsoft Foundry offer you deployed.
The upstream provider's direct API price is not necessarily the Azure price.
For this issue, compare:
Azure API usage telemetry
↓
Azure meter quantities
↓
Microsoft's price for that exact Foundry offer
rather than:
Azure API usage
↓
DeepSeek direct API price
Those are different commercial services.
I would collect this evidence before opening a billing case
For a small reproducible workload, capture:
- Exact model/deployment ID
- Provider/offer, especially whether it is FW-DeepSeek-V4-Pro
- Azure region/data-zone deployment type
- Request timestamps
- input_tokens
- cached_tokens
- output_tokens
- Azure Cost Management meter name
- Meter quantity
- Effective rate
Then calculate:
non_cached_input = input_tokens - cached_tokens
and compare the expected split against the actual Azure meter quantities.
If, for example, the API reports:
input_tokens: 100,000
cached_tokens: 70,000
then the important billing question is whether Azure records approximately:
30,000 normal input tokens
70,000 cached input tokens
or:
100,000 normal input tokens
0 cached input tokens
The second result, if reproduced consistently on an offer for which Microsoft publishes a Cached Input meter, would be strong evidence to provide to Azure billing/Foundry support.
One caution about Azure OpenAI caching documentation
I would not use the Azure OpenAI GPT prompt-caching rules as the authoritative behavior for DeepSeek.
Azure OpenAI and partner models can expose similar OpenAI-compatible response structures while still having different provider-specific caching implementations and billing behavior.
For DeepSeek through Fireworks, use the Fireworks-on-Foundry documentation and the pricing contract for that exact Foundry offer.
Microsoft's Fireworks model documentation, including the cache-affinity guidance, is here:
https://learn.microsoft.com/en-us/azure/foundry/how-to/fireworks/enable-fireworks-models
The current Foundry DeepSeek pricing page is here:
https://azure.microsoft.com/en-us/pricing/details/ai-foundry-models/deepseek/
For deeper Microsoft Learn material on deploying and working with models in Microsoft Foundry, the Foundry learning material is available here:
So I would not yet conclude that Azure is ignoring the cache simply because the calculated invoice does not match the API's "cached_tokens" field.
First establish the exact Foundry offer and inspect the actual Azure meter split.
But if this is a Fireworks DeepSeek V4 Pro deployment, the API repeatedly reports substantial cached usage, Microsoft's pricing for that offer provides a Cached Input meter, and Azure Cost Management still puts all of those tokens onto the normal input meter, you would have a well-isolated billing/metering discrepancy rather than merely a prompt-caching configuration issue.