How to interpret content filter offsets in streaming Responses API with async filtering

Mark Gregson 20 Reputation points
2026-05-28T04:31:54.0766667+00:00

I'm using Azure OpenAI Responses API streaming with asynchronous content filtering enabled and I'm trying to understand the semantics of the moderation offset fields (start_offset, check_offset, end_offset) returned in streaming events in the case of response content triggering the content filter.

I'm having trouble understanding the offsets I receive during testing based on the documentation https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/content-streaming#asynchronous-filtering.

Here's an example with a prompt designed so the model produces a few thousand chars of safe text followed by a scene for a novel that triggers the filter.

  • Model: gpt-4o v1 Responses API Async filtering mode
  • The stream terminated due to content filtering at ~5500 characters
  • Violating content appeared around >= 4500 chars
  • Only the final response.incomplete SSE event contained moderation/filter information (I've removed the filter category info for simplicity):
{ "blocked": true, "source_type": "completion", "content_filter_raw": [], "content_filter_results": { ... }, "content_filter_offsets": { "start_offset": 0, "end_offset": 2795, "check_offset": 0 } } 

This example shares some characteristics with all the test runs I've done:

  • Content filter information only appears in the final response.incomplete event, no other annotations events received
  • start_offset and check_offset always zero

start_offset and end_offset don't appear to described the span of the content that triggered the filter, since the violent content doesn't start for another ~1700 chars and because there are > 1000 characters before the stream was stopped.

Perhaps they describe the span of content that has been checked and judged safe but in that case I would expect check_offset to match end_offset.

In summary my questions are:

  1. How do I interpret start_offset, end_offset and check_offset in the context of Responses API streaming?
  2. Is it expected that:
    1. no intermediate annotation events are emitted during streaming?
    2. only the final response.incomplete event contains moderation metadata?
Azure OpenAI in Foundry Models
0 comments No comments

Answer accepted by question author
Alex Burlachenko 25,290 Reputation points MVP Volunteer Moderator
2026-05-28T08:35:03.8866667+00:00

hello Mark Gregson & thanks for join me here at Q&A portal,

offsets are not meant to point to the exact violating substring. They describe the moderation annotation range and moderation progress. In async filtering, the filter signal is delayed and may arrive after unsafe text was already streamed. Microsoft says the signal is guaranteed within about a 1,000-character window of the violating content https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/content-streaming#asynchronous-filtering

start_offset is the beginning of the text range that the annotation applies to. end_offset is the end of that annotation range. check_offset is the character position up to which content has been fully moderated, and it should never decrease. https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/content-streaming#understanding-content_filter_offsets

So in theory, for async streaming u should receive annotation events during the stream. Completion tokens are sent immediately, then annotation messages arrive later and refer back to already-sent text ranges. The docs explicitly say annotations are continuously returned during the stream.

Ur case is odd because only the final response.incomplete event contains moderation metadata, with start_offset: 0, check_offset: 0, and end_offset: 2795, while the suspected violating content appears much later. That does not line up cleanly with the documented Chat Completions examples. Two likely explanations: either the Responses API event schema maps these offsets differently than the older streaming docs, or this is a bug or incomplete implementation in Responses API async filtering.

I would not build production redaction logic assuming those values identify the exact unsafe span. Treat them as coarse moderation ranges only. Buffer recent output client-side, and when a block arrives, redact at least the last ~1,000 characters or the whole untrusted tail since the last confirmed safe check_offset.

For escalation, give Microsoft the exact model, API version, deployment filter config, full SSE trace, and ask specifically: “Do Responses API async filtering offsets follow the same semantics as Chat Completions streaming offsets, and should intermediate annotation events be emitted?”

Bottom line is docs say intermediate annotations should appear and check_offset should advance. If Responses API only emits final metadata with zero offsets, that needs clarification or a service bug fix.

rgds,

Alex

&

If my answer was helpful pls mark it and additional thx if u follow me at Q&A portal

Was this answer helpful?

1 person found this answer helpful.

2 additional answers

Sort by: Most helpful
  1. Amira Bedhiafi 43,046 Reputation points MVP Volunteer Moderator
    2026-05-28T07:29:09.21+00:00

    Hi Mark !

    Thank you for posting on MS Learn Q&A.

    The offsets should not be treated as exact coordinates of the unsafe text since they are moderation batch offsets. start_offset and end_offset identify the approximate text range to which the content filter result applies while check_offset is the moderation progress watermark: the point up to which the service considers text fully moderated. In async filtering, this watermark can lag behind the streamed text because tokens are sent before moderation completes.

    Therefore start_offset=0, end_offset=2795, check_offset=0 does not mean "the violation starts at 0" or "the stream stopped at 2795"

    It means the final blocking moderation result is associated with that moderation window and no earlier fully moderated watermark was advanced in the events you received.

    For the Responses API, I would not assume you will always receive intermediate annotation events. Handle content filter metadata wherever it appears especially on the terminal response.incomplete event and use incomplete_details.reason == "content_filter" plus content_filters[].blocked == true as the stop condition.

    If after accounting for offset basis or character counting, the filter signal is consistently more than the documented 1,000-character async filtering window behind the violating content, I would raise this with Azure support and include the request ID, API version, model deployment, full SSE event sequence and the measured output offsets.

    Was this answer helpful?


  2. Jerald Felix 18,760 Reputation points Volunteer Moderator
    2026-05-28T07:26:19.5066667+00:00

    Hello Mark Gregson,

    Greetings! Thanks for raising this question in Q&A forum.

    This is a really well-observed and detailed question! You've spotted some genuine nuances in how the Asynchronous Filter reports offset metadata. Let me address each of your questions clearly.

    Understanding the three offset fields

    check_offset shows how much text is fully moderated — it acts as an exclusive lower bound on the end_offset values of future annotation events and never decreases. All offsets are character positions with 0 at the beginning of the completion output.

    So to put it simply, here is what each field means in plain terms:

    start_offset — the character position in the completion where this particular moderation batch begins being evaluated. It marks the start of the chunk the filter is examining.

    end_offset — the character position where this moderation batch ends. The filter has evaluated content up to but not necessarily including this point.

    check_offset — the "safe frontier" — everything before this position has been fully checked and cleared. It guarantees that no future annotation will report an end_offset below this value.

    Why you are seeing start_offset and check_offset as zero in the blocking case

    This is the key insight for your scenario. When a policy violation is detected and the stream is terminated, the service emits only a single final response.incomplete event with the moderation metadata. In this terminal event, the content filtering error signal is delayed — if there is a policy violation, it is returned as soon as it is available and the stream stops. The content filtering signal is guaranteed within a ~1,000-character window of the policy-violating content.

    The offsets in this terminal event reflect the moderation window that triggered the block, not the span of safe content before it. In your case, end_offset: 2795 with start_offset: 0 and check_offset: 0 appears to indicate that the filter evaluated a block starting at position 0 of the output and the violation was detected within that evaluation window. The start_offset and check_offset being zero in a terminal block event is a known behavior — they essentially reset or default in the final blocking event rather than carrying forward the running position, which is admittedly not intuitive and is a documentation gap.

    Answering your specific questions

    Regarding your first question in normal non-blocking async streams, you would expect to see intermediate annotation events where start_offset, end_offset, and check_offset all advance progressively, as shown in the real-world example from the SDK: content_filter_offsets={'check_offset': 118, 'start_offset': 118, 'end_offset': 149} — here you can see all three moving forward together as each moderation batch completes. This is the healthy intermediate annotation pattern.

    Regarding your second question on whether it is expected that no intermediate annotation events are emitted and only the final response.incomplete contains metadata yes, this is expected behavior specifically in the blocking/violation case. Annotations and content moderation messages are continuously returned during the stream in the normal flow, but when a policy violation causes the stream to terminate, the signal is returned as soon as it is available and the stream stops immediately there is no opportunity to emit earlier intermediate annotations if the violation is detected in the first moderation pass.

    Practical guidance for your application

    Here is what I recommend for handling this reliably in your code:

    Step 1: Always check for response.incomplete with blocked: true This is your primary signal that the stream was terminated by a content filter. Do not rely on intermediate annotation events being present — they may not appear in the blocking case.

    Step 2: Use end_offset from the terminal event for redaction While start_offset and check_offset being zero is not very useful in the blocking case, end_offset: 2795 tells you approximately up to what character position of the streamed output the filter had evaluated. Use this to decide how much of your buffered output to discard or redact before displaying anything to the user.

    Step 3: Buffer content client-side and redact on block The documentation strongly recommends consuming annotations in your app and implementing client-side content redaction or returning other safety information to the user — since the async filter is a latency-safety tradeoff, your app should always be prepared to retroactively hide content when a blocking event arrives.

    Step 4: Raise a documentation feedback request The behavior of start_offset and check_offset defaulting to zero in the terminal blocking event is not clearly documented. I would encourage you to submit documentation feedback on the content-streaming page using the feedback link at the bottom of the article, as this would benefit other developers. You can also raise a technical support ticket with Azure OpenAI if you need an official confirmation of this offset behavior from the product team.

    If this answer helps you kindly accept the answer which will help others who have similar questions.

    Best Regards,

    Jerald Felix.

    Was this answer helpful?


Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.