Skip to content

fix: take streaming token counts from the engine - #385

Merged
Priyanshu-u07 merged 1 commit into
mainfrom
fix/stream-usage-from-upstream
Oct 6, 2026
Merged

Priyanshu-u07 merged 1 commit into
mainfrom
fix/stream-usage-from-upstream

Conversation

@Priyanshu-u07

Copy link
Copy Markdown
Collaborator

Nothing asked upstream for usage on a stream, so every streaming request fell through to estimating tokens from the message text, which misses the chat template. Measured at 10 prompt tokens against the engine's 35. check_quota and increment_redis_only count against these numbers, and the Insights token totals are built from them.

StreamProcessor already preferred provider usage when it saw any — it never saw any. The request now sets stream_options.include_usage unless the client set stream_options itself, and the usage-only chunk that produces is recorded and
dropped rather than forwarded, since the client did not ask for it.

External providers are unaffected. The injection happens before adapter.transform_request, and AnthropicAdapter and CohereAdapter both rebuild the payload from scratch, so the field is discarded before it reaches those APIs. The Anthropic surface sets stream_options itself, so this leaves that path alone.

Verified on a live deployment — A10G, vLLM 0.22.1 serving Qwen2.5-14B-Instruct-AWQ. A streaming request for "Count from one to twenty." logged prompt_tokens 35, matching what the same prompt reports non-streaming. The response ends at finish_reason "stop" then [DONE], with no usage frame.

closes #381

Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
@Priyanshu-u07
Priyanshu-u07 merged commit b88afe6 into main Oct 6, 2026
1 of 3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Streaming requests log token counts the engine did not report

1 participant