Getting AIModelStreamingTimeout Conversation in a Copilot Studio agent connected to an MCP server for Qlik Sense BI

Vivek Milind Shimpi 15 Reputation points
2026-07-16T09:54:13.5166667+00:00

We have developed a bot in copilot studio and talks to a MCP for Qliksense and pulls data from our enterprise Qliksense BI.  In a chat session, as the discussion progresses, I have noticed that after 2-3 prompts , just when the answers were getting better, I start getting error  like 'Error code: ContextTokenLimitExceeded Conversation Id: df3c4693-d806-4255-815f-104d4857f4ed Time (UTC): 2026-07-13T10:09:30.818Z.'    Why am i getting it and what do to eliminate it.'

Architecture:

 

Copilot Studio agent → MCP server → Qlik Sense Enterprise BI → AI-generated response.

 

Please confirm:

 

  1. What condition triggers AIModelStreamingTimeout Conversation and ContextTokenLimitExceeded?
  2. Is this caused by AI model response streaming time, MCP tool latency, connector response size, conversation state, or model context size?
  3. What are the applicable timeout limits for AI model streaming in Copilot Studio?
  4. Are MCP tool outputs injected fully into model context before response streaming?
  5. Can platform logs for our Conversation ID show token count, payload size, streaming duration and tool latency?
  6. What is Microsoft’s recommended pattern for BI agents returning large analytical results through MCP?
Microsoft Copilot | Other
0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.