Skip to content

[Bugfix][Rust Frontend] Fail token-less streams in vllm-bench openai-chat - #59736

Open
ZhenchengLin wants to merge 2 commits into
vllm-project:mainfrom
ZhenchengLin:fix/rust-bench-chat-empty-stream
Open

ZhenchengLin wants to merge 2 commits into
vllm-project:mainfrom
ZhenchengLin:fix/rust-bench-chat-empty-stream

Conversation

@ZhenchengLin

Copy link
Copy Markdown
Contributor

Purpose

Fixes #59722.

The Rust vllm-bench openai-chat backend set output.success = true whenever the stream ended cleanly, even if no token ever arrived. The openai (completions) backend and Python vllm bench serve only count a request as successful once a first token is received.

A token-less 200 response was therefore counted as a success with ttft, latency and output_tokens all at their defaults. As @aron-intframe pointed out in the issue, the impact goes beyond Mean TTFT:

  • Mean TTFT and Mean E2EL are pulled down, because both are 0 for those requests.
  • The 1 + output.itl.len() output-length fallback in metrics/calculator.rs counts each of them as one phantom output token, inflating total output tokens and output throughput.
  • They pass every goodput SLO, since TTFT/TPOT/E2EL are all 0.

This PR gates success on the existing first_token_received flag and fails the request with the same error message as the completions backend and Python, so the two Rust backends behave the same.

Behavior change (multi-turn). multi_turn.rs branches on output.success. Today a token-less turn is appended to the conversation history as an empty assistant message and the run continues; after this change that turn is reported as failed, like any other failed turn. This matches how the completions backend and Python already treat it.

Not a duplicate

Claimed in #59722. I checked open PRs that touch rust/src/bench/src/backends/openai_chat.rs:

Test Plan

New test openai_chat::tests::test_tokenless_stream_is_failed: the existing local SSE test server streams zero token chunks followed by the usage chunk and [DONE]. The test asserts success == false and that the error reports the missing first chunk.

cargo test -p vllm-bench --lib
cargo clippy -p vllm-bench --all-targets --locked -- -D warnings
cargo fmt --all -- --check

End-to-end: a mock OpenAI-compatible server that returns 200 with only the usage chunk and [DONE] for every 8th request, --dataset-name random --random-input-len 100 --random-output-len 8 --num-prompts 32 --max-concurrency 8.

Test Result

  • cargo test -p vllm-bench --lib: 169 passed. Clippy and fmt are clean.
  • With the fix reverted, the new test fails (assertion failed: !output.success).

End-to-end against the mock server:

client Successful Failed Total generated tokens Mean TTFT
Python vllm bench serve (openai-chat) 28 4 224 212.28 ms
Rust vllm-bench (openai), reference 28 4 224 203.13 ms
Rust vllm-bench (openai-chat), before 32 0 228 178.32 ms
Rust vllm-bench (openai-chat), after 28 4 224 202.93 ms

The before row shows the four phantom output tokens (228 vs 224) and the TTFT skew. After the fix, the chat backend matches the completions backend and Python.


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

AI assistance: prepared with Claude Code; I reviewed every changed line and ran the tests above locally.

ZhenchengLin and others added 2 commits October 1, 2026 22:37
…chat

The openai-chat backend set `success = true` whenever the stream ended
cleanly, even if no token arrived. Such requests were counted as
successes with zero TTFT and E2EL, and the `1 + itl.len()` output-length
fallback added a phantom output token for each, inflating the success
count and goodput and pulling Mean TTFT/E2EL down.

Gate `success` on `first_token_received` and fail the request with the
same error as the openai (completions) backend and Python
`vllm bench serve`.

Fixes vllm-project#59722

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: zhenchenglin <zhl132@ucsd.edu>
…mpty-stream

Signed-off-by: zhenchenglin <zhl132@ucsd.edu>
@ZhenchengLin
ZhenchengLin requested a review from esmeetu as a code owner October 2, 2026 05:38

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added rust bug Something isn't working labels Oct 2, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working rust

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Rust vllm-bench openai-chat counts token-less streams as successful

1 participant