Building and customizing solutions using Microsoft 365 Copilot APIs and tools
The behavior described matches known timeout and performance limitations when running prompts and other AI operations, especially during periods of high demand or when executions approach platform limits.
For Copilot Studio prompts (including document-based prompts):
- Execution time and timeouts
- Prompt execution is limited to 100 seconds. If the combined work of retrieving RAG data, analyzing documents, and running the LLM exceeds this, execution can time out and surface as internal or server errors.
- Larger or more complex documents, or prompts that require extensive grounding data, increase processing time and can trigger these limits.
- Load- and resource-related slowdowns
- GPU and model capacity are finite and allocated per region and per model. High demand can cause:
- Execution timeouts
- Inconsistent response times
- Throttling or transient internal errors
- This aligns with the “massive slow down” and new timeouts observed across multiple agents.
- GPU and model capacity are finite and allocated per region and per model. High demand can cause:
- Recommended mitigations before continuing work
- Simplify and narrow the prompt:
- Reduce the amount and size of document data used in the prompt.
- Remove unnecessary instructions or context.
- Use the most efficient model:
- Prefer the Basic or Standard model for simpler tasks and reserve Premium models only when necessary.
- Retry after short intervals:
- Transient backend or capacity issues often resolve after some time; repeated failures in a short window can be due to temporary throttling or regional load.
- If using long-context models or large outputs in related tools (for example, prompt flow or other OpenAI-based flows), consider:
- Shortening prompts and expected responses, or
- Using non-interactive/bulk modes where available, which may have more relaxed timeout behavior.
- Simplify and narrow the prompt:
- When to suspect a platform-side issue
- Multiple independent prompts and agents suddenly start timing out or failing with internal errors.
- Simple tests that previously worked now fail without configuration changes.
- Studio UI operations (such as testing or saving prompts) intermittently fail.
In that situation, the next steps are:
- Pause intensive testing temporarily to avoid repeated failures during a possible transient incident.
- Check service health and status for the relevant region in the admin portal.
- If the issue persists beyond normal transient behavior, open a support case with timestamps and session IDs so the backend team can investigate.
References: