Skip to content

No eos_token_id resolution from tokenizer_config.json: requests carry no eos and generation never stops #3364

Description

@ghazni101

Row: MODEL-QWEN35-EXL3

Local-Issue: ISSUE-LOCAL-01M3RY20H90NMK37V74EH34Z77
Kind: bug

Reproduced live on the vllm-qwen35-9b-exl3 container (Qwen3.5-9B-EXL3-4.00bpw, ROCm gfx1101, 2026-09-30). The checkpoint's config.json carries NO eos_token_id and ships no generation_config.json, and its tokenizer.json post_processor is a bare ByteLevel with no TemplateProcessing template, so Tokenizer::FromHfJson leaves eos_id_ = -1 (ExtractBosEos only reads post_processor). InputProcessor then finds no eos anywhere (input_processor.cpp:42-67) and every request goes out with eos_token_id unset: the model emits <|im_end|> (id 248046) at token 10 of a no-think reply and the engine ignores it, generating fake 'user\n\nassistant\n' turns to max_tokens (finish_reason=length). Upstream resolves the same id from the tokenizer itself: HF eos_token_id comes from tokenizer_config.json's eos_token ('<|im_end|>'), which FromHfJson never reads. User-visible symptom: thinking looks inconsistent — every answer ends by hallucinating more think blocks because <|im_end|> never terminates the turn.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions