FEAT: Capture API response stop reason on MessagePiece metadata - #2340
Open
varunj-msft wants to merge 1 commit into
Open
FEAT: Capture API response stop reason on MessagePiece metadata#2340varunj-msft wants to merge 1 commit into
varunj-msft wants to merge 1 commit into
Conversation
Targets already recorded token usage but discarded why generation stopped, and dropped usage entirely on content-filtered responses. Capture the provider's stop reason alongside token usage: finish_reason for Chat Completions, Completions and LiteLLM; status plus incomplete_reason for the Responses API. A base no-op hook on OpenAITarget is called from _handle_content_filter_response, so a filtered response now records the tokens it consumed instead of returning bare metadata. These keys are reserved for the provider. construct_response_from_request merges the request's metadata into every response piece, so all of them are cleared before a capture writes back the subset its own API reports.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Targets already recorded token usage but discarded why generation stopped, and dropped
usage entirely on content-filtered responses. This captures the provider's stop reason
alongside token usage on
MessagePiece.prompt_metadata:finish_reason— Chat Completions, Completions, LiteLLMstatus+incomplete_reason— Responses APIA base no-op hook
OpenAITarget._capture_response_metadatais called from_handle_content_filter_response, so a filtered response now records the tokens it consumed.These keys are reserved for the provider.
construct_response_from_requestmerges therequest's
prompt_metadatainto every response piece, so a caller-suppliedfinish_reasonwould otherwise be indistinguishable from the real one. All reserved keys are cleared from
every piece before each capture writes back what its own API reported;
token_usage_*iscleared by prefix for the same reason.
Two bugs fixed along the way:
OpenAICompletionTargetcaptured neither usage norfinish_reason.n>1, every piece got choice 0'sfinish_reason; each piece now gets its own.Not breaking:
_METADATA_PREFIX→TOKEN_USAGE_METADATA_PREFIXwas private with noexternal callers.
Tests and Documentation
tests/unit: 14,701 passed / 0 failed. Diff coverage 95% (gate is 90%).pre-commit run --all-filesclean, includingty.per-choice
finish_reasonwithn>1, content-filtered responses, "not reported" cases(missing/empty/non-string omitted rather than stored as zeros), and a SQLite round-trip.
finish_reason=stop; a truncated reasoning request →status=incomplete+incomplete_reason=max_output_tokens; and a forged caller-suppliedfinish_reasoncorrectly cleared.