fix(serve): return the requested number of chat top_logprobs - #5008
Open
MohammadHijjawi97 wants to merge 1 commit into
Open
MohammadHijjawi97 wants to merge 1 commit into
MohammadHijjawi97 wants to merge 1 commit into
Conversation
The engines return the model top-k logprobs plus the sampled token when it falls outside the top-k. The chat completions builder always dropped the sampled token from `top_logprobs`, so a request for `top_logprobs=k` got only k-1 candidates whenever the sampled token was in the top-k, and an empty list for greedy decoding with `top_logprobs=1`. With `top_logprobs=0` it could still return a candidate. Keep the sampled token when it is part of the top-k, drop it only when it is the extra appended row, and cap the list at `top_logprobs`, matching the `/generate` endpoint and the OpenAI API.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
/v1/chat/completionsreturns the wrong number oftop_logprobscandidates.For each output position, both engines return the model top-k logprobs, plus the sampled token when it falls outside the top-k (PyTorch
compute_logprobs, TurboMind_get_logprobs_impl)._create_chat_completion_logprobsalways removed the sampled token fromtop_logprobs, so:top_logprobs=k, only k-1 candidates are returned whenever the sampled token is in the top-k (the same symptom as [Bug] set logprobs = true and top_logprobs = 5 in restful server. The number of top logrobs is 4 which is unexpected. #1548);top_logprobs=1,top_logprobsis always[];logprobs=trueandtop_logprobsunset/0, a candidate can still be returned, because the engine is asked for 1 logprob.The OpenAI API lists the k most likely tokens, including the sampled one when it is among them. The
/generateendpoint (_create_top_logprobs) already does this.Modification
_create_chat_completion_logprobstakestop_logprobs. It sorts candidates by logprob, drops the sampled token only when it is the extra row appended outside the top-k, and caps the list attop_logprobs. The sampled token's owntoken/bytes/logprobare unchanged.request.top_logprobs or 0.tests/test_lmdeploy/serve/openai/chat_completions/test_logprobs.py. It covers the sampled token inside the top-k (k=1 and k=3), the sampled token outside the top-k, andtop_logprobs=0.pytest tests/test_lmdeploy/serve/openai/chat_completions/passes (51 tests), and pre-commit is clean on the changed files.BC-breaking (Optional)
No API change. Responses now contain
top_logprobsentries as requested; the new parameter defaults to 0.Checklist