fix: take streaming token counts from the engine - #385
Merged
Merged
Conversation
Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Nothing asked upstream for usage on a stream, so every streaming request fell through to estimating tokens from the message text, which misses the chat template. Measured at 10 prompt tokens against the engine's 35. check_quota and increment_redis_only count against these numbers, and the Insights token totals are built from them.
StreamProcessor already preferred provider usage when it saw any — it never saw any. The request now sets stream_options.include_usage unless the client set stream_options itself, and the usage-only chunk that produces is recorded and
dropped rather than forwarded, since the client did not ask for it.
External providers are unaffected. The injection happens before adapter.transform_request, and AnthropicAdapter and CohereAdapter both rebuild the payload from scratch, so the field is discarded before it reaches those APIs. The Anthropic surface sets stream_options itself, so this leaves that path alone.
Verified on a live deployment — A10G, vLLM 0.22.1 serving Qwen2.5-14B-Instruct-AWQ. A streaming request for "Count from one to twenty." logged prompt_tokens 35, matching what the same prompt reports non-streaming. The response ends at finish_reason "stop" then [DONE], with no usage frame.
closes #381