[Bugfix][Rust Frontend] Fail token-less streams in vllm-bench openai-chat - #59736
Open
ZhenchengLin wants to merge 2 commits into
Open
ZhenchengLin wants to merge 2 commits into
ZhenchengLin wants to merge 2 commits into
Conversation
…chat The openai-chat backend set `success = true` whenever the stream ended cleanly, even if no token arrived. Such requests were counted as successes with zero TTFT and E2EL, and the `1 + itl.len()` output-length fallback added a phantom output token for each, inflating the success count and goodput and pulling Mean TTFT/E2EL down. Gate `success` on `first_token_received` and fail the request with the same error as the openai (completions) backend and Python `vllm bench serve`. Fixes vllm-project#59722 Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: zhenchenglin <zhl132@ucsd.edu>
…mpty-stream Signed-off-by: zhenchenglin <zhl132@ucsd.edu>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Fixes #59722.
The Rust
vllm-benchopenai-chatbackend setoutput.success = truewhenever the stream ended cleanly, even if no token ever arrived. Theopenai(completions) backend and Pythonvllm bench serveonly count a request as successful once a first token is received.A token-less 200 response was therefore counted as a success with
ttft,latencyandoutput_tokensall at their defaults. As @aron-intframe pointed out in the issue, the impact goes beyond Mean TTFT:0for those requests.1 + output.itl.len()output-length fallback inmetrics/calculator.rscounts each of them as one phantom output token, inflating total output tokens and output throughput.0.This PR gates
successon the existingfirst_token_receivedflag and fails the request with the same error message as the completions backend and Python, so the two Rust backends behave the same.Behavior change (multi-turn).
multi_turn.rsbranches onoutput.success. Today a token-less turn is appended to the conversation history as an empty assistant message and the run continues; after this change that turn is reported as failed, like any other failed turn. This matches how the completions backend and Python already treat it.Not a duplicate
Claimed in #59722. I checked open PRs that touch
rust/src/bench/src/backends/openai_chat.rs:Test Plan
New test
openai_chat::tests::test_tokenless_stream_is_failed: the existing local SSE test server streams zero token chunks followed by the usage chunk and[DONE]. The test assertssuccess == falseand that the error reports the missing first chunk.cargo test -p vllm-bench --lib cargo clippy -p vllm-bench --all-targets --locked -- -D warnings cargo fmt --all -- --checkEnd-to-end: a mock OpenAI-compatible server that returns 200 with only the usage chunk and
[DONE]for every 8th request,--dataset-name random --random-input-len 100 --random-output-len 8 --num-prompts 32 --max-concurrency 8.Test Result
cargo test -p vllm-bench --lib: 169 passed. Clippy and fmt are clean.assertion failed: !output.success).End-to-end against the mock server:
vllm bench serve(openai-chat)vllm-bench(openai), referencevllm-bench(openai-chat), beforevllm-bench(openai-chat), afterThe before row shows the four phantom output tokens (228 vs 224) and the TTFT skew. After the fix, the chat backend matches the completions backend and Python.
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.AI assistance: prepared with Claude Code; I reviewed every changed line and ran the tests above locally.