Formerly known as Azure AI Services or Azure Cognitive Services is a unified collection of prebuilt AI capabilities within the Microsoft Foundry platform
Assistant/Agent threads stuck as "QUEUED"
Lots of my threads are stuck as "RunStatus.QUEUED" seemingly at random. They get stuck like this for ~30 mins before they're cancelled. Why might this be happening?
Foundry Tools
-
Anonymous
2025-09-12T21:07:26.09+00:00 Threads getting stuck in
RunStatus.QUEUEDfor approximately 30 minutes before cancellation typically point to underlying issues in resource availability, orchestration logic, or tool dependencies. This status indicates that the thread run has been accepted but is waiting to start, often due to compute resource bottlenecks such as unavailable or saturated CPU/GPU nodes—which delay execution. Azure Machine Learning tracks this behavior through metrics like “Queued Runs” and “Started Runs,” helping identify when resource constraints are the root cause.Another common reason is delayed or missing tool outputs; if a thread depends on external tool responses and they aren’t submitted in time, it may remain queued or transition to
requires_action, eventually leading to cancellation. Additionally, Azure’s cancellation policies or manual triggers may automatically cancel threads that exceed predefined wait thresholds. Misconfigured thread lifecycles—such as malformed or missing run IDs can also prevent threads from progressing beyond the queued state. These issues may stem from bugs in orchestration logic or incorrect API usage. Azure provides diagnostic metrics like “Cancelled Runs,” “Cancel Requested Runs,” “Failed Runs,” and “Finalizing Runs” to help pinpoint patterns and failure points. To resolve this, ensure that compute resources are adequately provisioned and autoscaling is configured, tool outputs are submitted promptly, thread configurations are validated, and monitoring dashboards are actively used to track and analyze thread behavior across lifecycle stages.Hope it helps!
Thank you
-
Williams, Dan • 66 Reputation points
2025-09-12T21:49:56.9+00:00 I am using the Azure AI Agent service.
- Unavailable or saturated CPU/GPU nodes - I cannot control this, can I? This sounds like an Azure issue that needs to be resolved by Microsoft? What can I do about this when I am using the Azure AI Agent service?
- Delayed or missing tool outputs - I am having the same issue on newly created agents without any tools. I am testing these in the playground. I am having the same issue on both gpt-4.1 and 4o, both of which have TPM quotas beyond what I am using. Whenever I've had issues with tool output timing before, I get a clear message in the thread logs. In this case it seems like the thread just never starts.
- These issues may stem from bugs in orchestration logic or incorrect API usage. - I get the same issue from the playground UI and also from the API. When the thread/run expires, the status I get is: RunStatus.EXPIRED
How do you suggest I troubleshoot this when we are using the Azure AI Foundry Agent service?
-
Anonymous
2025-09-17T06:33:57.8266667+00:00 Hello Williams, Dan
Sorry for delayed response!
If you're using the Azure AI Agent service and encountering issues like unavailable or saturated CPU/GPU nodes, delayed or missing tool outputs, and threads expiring with
RunStatus.EXPIRED, these are typically caused by Azure infrastructure constraints or orchestration logic bugs. First, CPU/GPU saturation is managed by Azure and cannot be controlled by users—it results in threads staying inQUEUEDstate until they expire. Second, delayed or missing tool outputs—even on agents without tools—can occur due to orchestration expecting system-level responses, and this affects both Playground and API usage. Third,RunStatus.EXPIREDmeans the thread did not complete within the allowed time, often due to stuck states likeQUEUEDorREQUIRES_ACTION.To troubleshoot, ensure your agent has valid tool bindings, monitor thread lifecycle states, use asynchronous run handling where possible, and test in different Azure regions to rule out regional saturation. If the issue persists, escalate to Azure support with thread IDs and timestamps for deeper investigation. This approach addresses all three of your questions: control over compute saturation, causes of tool output delays, and how to troubleshoot expired runs.
Hope it Helps!
Thank you
-
Anonymous
2025-09-23T02:15:17.4766667+00:00 -
Anonymous
2025-09-29T23:47:54.09+00:00 Hello Williams, Dan
Just following up to see if you had a chance to review the above response.
Thank you!
-
Ali Asgar • 11 Reputation points • Microsoft Employee2025-11-07T23:59:12.6033333+00:00 Same issue as reported here. No resolution from Microsoft Support yet. Can someone please NOT use copilot to just post useless answers and instead help resolve the issue?
Sign in to comment