An Azure service that provides access to OpenAI’s GPT-3 models with enterprise capabilities.
hello Mark Gregson & thanks for join me here at Q&A portal,
offsets are not meant to point to the exact violating substring. They describe the moderation annotation range and moderation progress. In async filtering, the filter signal is delayed and may arrive after unsafe text was already streamed. Microsoft says the signal is guaranteed within about a 1,000-character window of the violating content https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/content-streaming#asynchronous-filtering
start_offset is the beginning of the text range that the annotation applies to. end_offset is the end of that annotation range. check_offset is the character position up to which content has been fully moderated, and it should never decrease. https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/content-streaming#understanding-content_filter_offsets
So in theory, for async streaming u should receive annotation events during the stream. Completion tokens are sent immediately, then annotation messages arrive later and refer back to already-sent text ranges. The docs explicitly say annotations are continuously returned during the stream.
Ur case is odd because only the final response.incomplete event contains moderation metadata, with start_offset: 0, check_offset: 0, and end_offset: 2795, while the suspected violating content appears much later. That does not line up cleanly with the documented Chat Completions examples. Two likely explanations: either the Responses API event schema maps these offsets differently than the older streaming docs, or this is a bug or incomplete implementation in Responses API async filtering.
I would not build production redaction logic assuming those values identify the exact unsafe span. Treat them as coarse moderation ranges only. Buffer recent output client-side, and when a block arrives, redact at least the last ~1,000 characters or the whole untrusted tail since the last confirmed safe check_offset.
For escalation, give Microsoft the exact model, API version, deployment filter config, full SSE trace, and ask specifically: “Do Responses API async filtering offsets follow the same semantics as Chat Completions streaming offsets, and should intermediate annotation events be emitted?”
Bottom line is docs say intermediate annotations should appear and check_offset should advance. If Responses API only emits final metadata with zero offsets, that needs clarification or a service bug fix.
rgds,
Alex
&
If my answer was helpful pls mark it and additional thx if u follow me at Q&A portal