Azure OpenAI (GPT-5.4 Data Zone Standard): Responses API intermittently errors with message "The AI service failed to generate a response."

Suhas Subramanya 45 Reputation points
2026-06-09T11:05:12.75+00:00

Service

Azure OpenAI — gpt-5.4, Data Zone Standard (PAYG), Responses API (POST /openai/responses), streaming, via the Azure OpenAI .NET SDK. Stateless reasoning use: store=false, reasoning.encrypted_content in include, reasoning.effort low/medium, set max_output_tokens. Generations are large — typically 1–3 minutes end to end. Region: EU data zone (France Central + Sweden Central).

What happens

The request returns 200 OK and streaming starts normally (response.created, then output events). Partway through, a top-level error event is emitted and generation aborts. Our SDK surfaced only a generic "The AI service failed to generate a response."; after enabling raw-payload logging we captured the actual body:

{
  "type": "error",
  "sequence_number": 2,
  "error": {
    "type": "too_many_requests",
    "code": "no_capacity",
    "message": "The system is currently experiencing high demand and cannot process your request. Your request exceeds the maximum usage size allowed during peak load. For improved capacity reliability, consider switching to Provisioned Throughput."
  }
}

Key points

  • No HTTP 429. Status was 200 throughout — too_many_requests appears only as a field inside the streamed error event, not as a response status.
  • Not our rate limit. Failing calls show near-full headroom: x-ratelimit-remaining-requests: 3999, x-ratelimit-remaining-tokens: ≈3.97M. No throttling response precedes the error.
  • Mid-stream. sequence_number: 2 — the request is admitted, streaming begins, then the error arrives mid-generation. A retry of the same request almost always succeeds.

The pattern — this is what we'd like explained

This error was not present before 2026-06-05 (only sporadic, isolated, unnoticeable errors). It then appeared as clustered bursts that escalated ~50× over four days, all concentrated in a consistent ~08:00–10:00 UTC morning band (≈10:00–12:00 CEST), with a hot core around 09:00–09:40 UTC:

Day (UTC) Peak window Volume
06-05 09:15–09:40 (single 20-min burst) ~21
06-08 ~08:30 ramp → sustained to 13:35 (spike 13:05–13:15), tail to 18:39 ~315
06-09 07:45 → 10:10 ~1,050

Then on 2026-06-09 the errors stopped abruptly at 10:10 UTC — a hard cliff to zero, not a taper (the 10:05 5-min bucket had 56 errors; from 10:10 onward, none). Critically, our request volume was still substantial at that moment (comparable to the morning ramp that was erroring heavily) — so this was not a drop in our traffic.

So the signature is: off before the 5th → ramps up sharply and daily → switches off instantly mid-day, independent of our load.

Example occurrence

  • UTC: 2026-06-08T10:30:37Z · Region: Sweden Central
  • apim-request-id: 7987c5e7-4f2e-4b4e-b694-2afdfb125d9f
  • x-request-id: 02b624c4-bd74-4cab-80e2-5d66e1a70a53

Questions

  1. Is this a service-side issue or something on our end? The clean on/off pattern — absent before 06-05, escalating ~50× daily, then instantly zero at ~10:10 UTC on 06-09 while our traffic continued — looks like a capacity event on the service side, not our usage. Can you confirm?
  2. What changed around 06-05, and what cleared at ~10:10 UTC on 06-09? Any capacity, region-balancing, model-version, or quota-policy change to gpt-5.4 Data Zone Standard (EU) in those windows?
  3. Is no_capacity a shared-pool capacity throttle, separate from our TPM/RPM quota? We have full rate-limit headroom and never saw a 429 status — we believe these are two independent admission gates.
  4. Why mid-stream? Is capacity re-checked during generation, making long (1–3 min) reasoning calls more exposed when the pool saturates? Any mitigation for long generations?
  5. Recommended handling: the event carries no Retry-After, so we're using exponential backoff + jitter. Is that correct for no_capacity, and what PTU sizing would eliminate it for our workload?
Azure OpenAI in Foundry Models
0 comments No comments

Answer accepted by question author
Amira Bedhiafi 43,046 Reputation points MVP Volunteer Moderator
2026-06-09T18:50:11.2866667+00:00

Hello Suhas !

Thank you for posting on MS Learn Q&A.

This looks much more like Azure OpenAI shared capacity pressure than an application-side TPM/RPM quota issue.

The important part is the following

"type": "too_many_requests",
"code": "no_capacity",
"message": "The system is currently experiencing high demand..."

For Data Zone Standard or Standard PAYG, capacity is shared and the traffic is routed to available capacity within the deployment boundary but high sustained or bursty usage can see greater latency variability and in PTU is recommended for latency-critical or high-volume workloads requiring predictable capacity.

https://learn.microsoft.com/en-us/azure/foundry/openai/quotas-limits

The fact that your headers still show large remaining TPM/RPM strongly suggests this is not the normal deployment quota gate because if you check the Azure OpenAI docs you will find separate normal quota/rate-limit accounting from system-capacity throttling and the capacity-related 429-style errors are cases where the backend cannot process the request at that time, often transient, and recommend retry/backoff or Provisioned Throughput if persistent.

https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/quota

The 200 OK followed by a streamed error event is also expected behavior for the Responses API because during streaming, the Responses API can emit an error event for 500, 429 and similar failures after streaming has started so clients should detect that event and gracefully stop or restart the stream.

https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/responses

Was this answer helpful?

1 person found this answer helpful.

0 additional answers

Sort by: Newest

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.