Service
Azure OpenAI — gpt-5.4, Data Zone Standard (PAYG), Responses API (POST /openai/responses), streaming, via the Azure OpenAI .NET SDK. Stateless reasoning use: store=false, reasoning.encrypted_content in include, reasoning.effort low/medium, set max_output_tokens. Generations are large — typically 1–3 minutes end to end. Region: EU data zone (France Central + Sweden Central).
What happens
The request returns 200 OK and streaming starts normally (response.created, then output events). Partway through, a top-level error event is emitted and generation aborts. Our SDK surfaced only a generic "The AI service failed to generate a response."; after enabling raw-payload logging we captured the actual body:
{
"type": "error",
"sequence_number": 2,
"error": {
"type": "too_many_requests",
"code": "no_capacity",
"message": "The system is currently experiencing high demand and cannot process your request. Your request exceeds the maximum usage size allowed during peak load. For improved capacity reliability, consider switching to Provisioned Throughput."
}
}
Key points
- No HTTP 429. Status was 200 throughout —
too_many_requests appears only as a field inside the streamed error event, not as a response status.
- Not our rate limit. Failing calls show near-full headroom:
x-ratelimit-remaining-requests: 3999, x-ratelimit-remaining-tokens: ≈3.97M. No throttling response precedes the error.
- Mid-stream.
sequence_number: 2 — the request is admitted, streaming begins, then the error arrives mid-generation. A retry of the same request almost always succeeds.
The pattern — this is what we'd like explained
This error was not present before 2026-06-05 (only sporadic, isolated, unnoticeable errors). It then appeared as clustered bursts that escalated ~50× over four days, all concentrated in a consistent ~08:00–10:00 UTC morning band (≈10:00–12:00 CEST), with a hot core around 09:00–09:40 UTC:
| Day (UTC) |
Peak window |
Volume |
| 06-05 |
09:15–09:40 (single 20-min burst) |
~21 |
| 06-08 |
~08:30 ramp → sustained to 13:35 (spike 13:05–13:15), tail to 18:39 |
~315 |
| 06-09 |
07:45 → 10:10 |
~1,050 |
Then on 2026-06-09 the errors stopped abruptly at 10:10 UTC — a hard cliff to zero, not a taper (the 10:05 5-min bucket had 56 errors; from 10:10 onward, none). Critically, our request volume was still substantial at that moment (comparable to the morning ramp that was erroring heavily) — so this was not a drop in our traffic.
So the signature is: off before the 5th → ramps up sharply and daily → switches off instantly mid-day, independent of our load.
Example occurrence
- UTC: 2026-06-08T10:30:37Z · Region: Sweden Central
-
apim-request-id: 7987c5e7-4f2e-4b4e-b694-2afdfb125d9f
-
x-request-id: 02b624c4-bd74-4cab-80e2-5d66e1a70a53
Questions
- Is this a service-side issue or something on our end? The clean on/off pattern — absent before 06-05, escalating ~50× daily, then instantly zero at ~10:10 UTC on 06-09 while our traffic continued — looks like a capacity event on the service side, not our usage. Can you confirm?
- What changed around 06-05, and what cleared at ~10:10 UTC on 06-09? Any capacity, region-balancing, model-version, or quota-policy change to gpt-5.4 Data Zone Standard (EU) in those windows?
- Is
no_capacity a shared-pool capacity throttle, separate from our TPM/RPM quota? We have full rate-limit headroom and never saw a 429 status — we believe these are two independent admission gates.
- Why mid-stream? Is capacity re-checked during generation, making long (1–3 min) reasoning calls more exposed when the pool saturates? Any mitigation for long generations?
- Recommended handling: the event carries no
Retry-After, so we're using exponential backoff + jitter. Is that correct for no_capacity, and what PTU sizing would eliminate it for our workload?