Why Azure foundry seems not 'cache' token for deepseek v4 pro?

jinxjer lee 40 Reputation points
2026-09-18T05:35:31.0466667+00:00

I'm using DeepSeek Harness to configure the Azure AI Foundry DeepSeek V4 Pro model through the standard OpenAI Chat Completions API. In the API response, I can see that about 50–80% of the tokens are reported as cached. However, Azure bills all consumed tokens at the non-cached rate. Is Azure failing to apply cached-token billing for DeepSeek, or is there a bug in its cost calculation?

Microsoft Foundry
Microsoft Foundry

A unified Azure platform for creating and managing AI models, agents, and applications with built‑in enterprise security, monitoring, and governance

0 comments No comments

1 answer

Sort by: Most helpful
  1. Walker Pollitt 0 Reputation points
    2026-09-19T15:32:40.63+00:00

    The key point is that "cached_tokens" in the model response and the Azure billing meter are related signals, but they are not sufficient by themselves to prove what rate Azure actually applied to those tokens.

    I would troubleshoot this by first identifying exactly which DeepSeek V4 Pro offer/deployment you are using.

    Microsoft Foundry currently exposes DeepSeek V4 Pro through different provider/deployment paths, including Fireworks-hosted models. The caching behavior and billing contract should therefore be checked against the specific offer rather than assuming that every deployment named "DeepSeek V4 Pro" has identical caching semantics.

    If this is the Fireworks DeepSeek V4 Pro offer

    Microsoft's current Foundry documentation explicitly documents prompt-cache optimization for Fireworks models.

    It recommends maintaining affinity for related requests using one of:

    x-session-affinity

    or the request parameter:

    user

    or:

    prompt_cache_key

    with "prompt_cache_key" taking priority over "user".

    So if your deployment is "FW-DeepSeek-V4-Pro", I would first make sure your repeated requests are using stable cache-affinity information and an identical reusable prompt prefix.

    However, your observation is actually more interesting than a simple cache miss, because you are already receiving something like:

    "input_tokens_details": {

    "cached_tokens": ...
    

    }

    with 50-80% of the input reportedly cached.

    That suggests the provider/model layer is reporting cached-token activity.

    The next question is therefore not:

    «"Is caching happening?"»

    but:

    «"Which Azure billing meter is receiving those cached tokens?"»

    Check the Azure meter, not only the API usage object

    I would compare a controlled test against Azure Cost Management rather than calculate cost only from total API input tokens.

    Run two controlled workloads:

    Test A:

    unique/non-repeated prompt prefixes

    Test B:

    identical long prompt prefix

    stable session/cache affinity

    same model/deployment

    similar output length

    Record for every request:

    timestamp

    deployment/model ID

    total input tokens

    cached_tokens

    output tokens

    request/correlation ID if available

    Then inspect the Azure usage/cost data for the same interval and determine whether usage is being split into separate normal-input and cached-input meters.

    Microsoft's Foundry pricing pages distinguish Input, Cached Input, and Output for supported DeepSeek/Fireworks offerings.

    Therefore, if the API consistently reports substantial cached-token usage but Cost Management records those same tokens exclusively against the normal input meter, that is no longer simply a prompt-cache hit-rate problem. It becomes a billing/metering discrepancy worth escalating.

    Do not calculate the expected Azure bill from DeepSeek's direct API pricing

    Another important distinction is that the Foundry deployment is billed according to the Microsoft Foundry offer you deployed.

    The upstream provider's direct API price is not necessarily the Azure price.

    For this issue, compare:

    Azure API usage telemetry

        ↓
    

    Azure meter quantities

        ↓
    

    Microsoft's price for that exact Foundry offer

    rather than:

    Azure API usage

        ↓
    

    DeepSeek direct API price

    Those are different commercial services.

    I would collect this evidence before opening a billing case

    For a small reproducible workload, capture:

    1. Exact model/deployment ID
    2. Provider/offer, especially whether it is FW-DeepSeek-V4-Pro
    3. Azure region/data-zone deployment type
    4. Request timestamps
    5. input_tokens
    6. cached_tokens
    7. output_tokens
    8. Azure Cost Management meter name
    9. Meter quantity
    10. Effective rate

    Then calculate:

    non_cached_input = input_tokens - cached_tokens

    and compare the expected split against the actual Azure meter quantities.

    If, for example, the API reports:

    input_tokens: 100,000

    cached_tokens: 70,000

    then the important billing question is whether Azure records approximately:

    30,000 normal input tokens

    70,000 cached input tokens

    or:

    100,000 normal input tokens

    0 cached input tokens

    The second result, if reproduced consistently on an offer for which Microsoft publishes a Cached Input meter, would be strong evidence to provide to Azure billing/Foundry support.

    One caution about Azure OpenAI caching documentation

    I would not use the Azure OpenAI GPT prompt-caching rules as the authoritative behavior for DeepSeek.

    Azure OpenAI and partner models can expose similar OpenAI-compatible response structures while still having different provider-specific caching implementations and billing behavior.

    For DeepSeek through Fireworks, use the Fireworks-on-Foundry documentation and the pricing contract for that exact Foundry offer.

    Microsoft's Fireworks model documentation, including the cache-affinity guidance, is here:

    https://learn.microsoft.com/en-us/azure/foundry/how-to/fireworks/enable-fireworks-models

    The current Foundry DeepSeek pricing page is here:

    https://azure.microsoft.com/en-us/pricing/details/ai-foundry-models/deepseek/

    For deeper Microsoft Learn material on deploying and working with models in Microsoft Foundry, the Foundry learning material is available here:

    https://learn.microsoft.com/en-us/training/browse/?products=azure-ai-foundry&wt.mc_id=studentamb_521824

    So I would not yet conclude that Azure is ignoring the cache simply because the calculated invoice does not match the API's "cached_tokens" field.

    First establish the exact Foundry offer and inspect the actual Azure meter split.

    But if this is a Fireworks DeepSeek V4 Pro deployment, the API repeatedly reports substantial cached usage, Microsoft's pricing for that offer provides a Cached Input meter, and Azure Cost Management still puts all of those tokens onto the normal input meter, you would have a well-isolated billing/metering discrepancy rather than merely a prompt-caching configuration issue.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.