Azure OpenAI embeddings returning 403 temporarily blocked due to unusual behavior

Jorge Ortiz Flores 20 Reputation points
2026-05-27T19:37:36.85+00:00
  • We have an Azure-hosted app whose non-OpenAI endpoints are healthy:
    • /api/health returns 200
    • /api/openapi.json returns 200
    • Swagger loads
  • Calls that invoke Azure OpenAI embeddings fail with:
    • 403 Forbidden
    • Your resource has been temporarily blocked because we detected unusual behavior.
  • We observed 9 failures to text-embedding-3-large/embeddings in 3 bursts.
  • We have correlation IDs and timestamps available.
  • We need guidance on whether this requires manual unblock, cooldown, quota/rate changes, or resource replacement.
Azure OpenAI in Foundry Models

Answer accepted by question author
Karnam Venkata Rajeswari 5,255 Reputation points Microsoft External Staff Moderator
2026-06-08T17:13:24.1433333+00:00

Hello @Jorge Ortiz Flores ,

Thank you for helping us with the required details to assist you better

The observed 403 errors originate from two different scenarios, and each follows a separate resolution path.

  1. GPT deployments (chat and mini) The failures are not caused by a platform issue. Instead, requests are being blocked at the access layer due to configured authentication and network policies, meaning traffic is rejected before reaching the model backend.
  2. Embedding deployment (text-embedding-3-large) This issue is independent and is caused by a Content Safety enforcement block applied at the resource level. This type of block requires a formal review and cannot be resolved through configuration changes.

GPT models are returning 403 as

The behavior is consistent with security settings applied on the resource:

  1. Authentication mismatch
    1. Key-based authentication is disabled
    2. Application or Azure AI Search embedding flow still using API key
    3. Requests are rejected during access validation
  2. Network restrictions
    1. Public network access disabled while calls are made via public endpoint
    2. Firewall rules blocking incoming traffic
  3. Azure AI Search configuration
    1. Embedding skill still configured with apiKey instead of authIdentity
    2. Key authentication takes precedence when configured

How to unblock GPT resources

  1. Immediate workaround (fast recovery)
    1. Re-enable key-based authentication temporarily
    2. Location: Azure OpenAI resource → Keys and Endpoint → enable local auth This helps restore access quickly while long-term fixes are implemented.

Permanent resolution

  1. Migrate to Microsoft Entra ID authentication
    1. Use managed identity or token-based authentication
    2. Assign required role: Cognitive Services OpenAI User
  2. Update Azure AI Search embedding configuration
    1. Replace apiKey with authIdentity
    2. Enable managed identity on Azure AI Search
    3. Ensure role assignment on Azure OpenAI resource
  3. Validate network access settings
    1. Check if public access is disabled
    2. Verify firewall rules allow traffic from calling service
    3. Validate private endpoint / trusted service configuration

How to resolve embedding (text-embedding-3-large)

This scenario requires review through the official process:

  1. Submit review request: Code of Conduct for Microsoft AI Services | Microsoft Learn
  2. This block is enforced by Content Safety systems and needs manual evaluation.

Please refer the following for further reference:

  1. gpt-5-chat Content filtering for Microsoft Foundry Models (classic) - Microsoft Foundry (classic) portal | Microsoft Learn Azure API Management Troubleshooting Scenario 5 - Request throttling problems and HTTP 403 - Forbidden issues - Azure | Microsoft Learn
  2. gpt-5-mini https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/managed-identity https://learn.microsoft.com/en-us/troubleshoot/azure/api-mgmt/availability/request-throttling-http-403#troubleshooting-steps

To summarize

  • GPT-related errors are caused by authentication and network configuration mismatch, and can be resolved through authentication migration or network corrections.
  • Embedding failure is due to a Content Safety enforcement block, which requires submission for review.

Following these steps should help restore functionality for GPT deployments and guide the next steps for the embedding scenario.

Please let us know the resolution updates

Thank you

Was this answer helpful?

1 person found this answer helpful.
0 comments No comments

2 additional answers

Sort by: Most helpful
  1. Karnam Venkata Rajeswari 5,255 Reputation points Microsoft External Staff Moderator
    2026-05-27T19:51:43.99+00:00

    Hello @Jorge Ortiz Flores ,

    Welcome to Microsoft Q&A .Thank you for reaching out to us.

    Based on the details shared, the observed 403 response for the embeddings endpoint is consistent with a temporary service-side protection behavior, rather than an authentication or permission issue.

    This type of response typically occurs when the service detects unusual or bursty traffic patterns, such as high concurrency or sudden spikes in request volume.

    Key Understanding of the Behavior:

    1. The request is successfully received, but temporarily blocked due to traffic characteristics, not access configuration
    2. Such conditions are often linked to rate and throughput limits at the model/deployment level (Requests per minute and Tokens per minute)
    3. This behavior is automated and transient, not a persistent fault condition

    Regarding the recovery -

    1. The resource typically recovers automatically after a short cooldown period
    2. No manual intervention is required in most cases
    3. Recovery is dependent on traffic normalization rather than configuration changes

    To improve stability and avoid future blocks, the following practices are recommended:

    1. Controlling request patterns
    2. Limit concurrent requests
      1. Avoid sudden bursts; distribute traffic evenly over time
    3. Implementing retry strategy
      1. Use exponential backoff with jitter to handle transient failures
      2. Respect retry delays indicated by service responses where available
    4. Aligning with quota limits
      1. Review model-level limits such as:
      • Requests per minute (RPM)
        • Tokens per minute (TPM)
      1. Adjust batching and payload sizes accordingly
    5. Optimizing workload design
    6. Tune batch sizes instead of sending large spikes
      1. Reduce unnecessary request duplication
    7. Monitoring usage and trends
    8. Track request volume, token usage, and errors using Azure Monitor metrics
      1. Key metrics such as tokens processed and request rates provide early signals for throttling risk

    The following references might be helpful , please check them out

    Please let us know if the response was helpful

     

    Thank you

    Was this answer helpful?

    1 person found this answer helpful.
    0 comments No comments

  2. Alex Burlachenko 25,290 Reputation points MVP Volunteer Moderator
    2026-05-29T09:33:28.5166667+00:00

    hi Jorge Ortiz Flores & thanks for join me here at Q&A portal,

    this is not your app health issue. /api/health, Swagger and non-OpenAI routes working only proves the app is fine. The 403 message means the Azure OpenAI resource itself was temporarily blocked by Microsoft’s safety or abuse detection layer.

    This is different from quota or normal rate limiting. Quota/rate issues usually return 429. Here the message is explicit temporarily blocked because unusual behavior was detected.... Stop retries for now. Repeated bursts against embeddings can keep the block active or make the signal worse. Then open Azure support ticket under Azure OpenAI / Azure AI Foundry Models and request a temporary block review / unusual behavior review. Include

    resource name and resource ID, deployment name text-embedding-3-large, timestamps, correlation IDs, sample request volume pattern, explanation that this was embeddings traffic from your app, not abuse.

    Do not replace the resource unless Microsoft confirms it is required. Resource replacement may not help if the block is at subscription, tenant or risk-policy level.

    For future runs, add exponential backoff, lower concurrency, batching control and request throttling around embeddings https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/abuse-monitoring

    rgds,

    Alex

    &

    If my answer was helpful pls mark it and additional thx if u follow me at Q&A portal
    

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.