Row: MODEL-QWEN35-EXL3
Local-Issue: ISSUE-LOCAL-01M3RY20H90NMK37V74EH34Z77
Kind: bug
Reproduced live on the vllm-qwen35-9b-exl3 container (Qwen3.5-9B-EXL3-4.00bpw, ROCm gfx1101, 2026-09-30). The checkpoint's config.json carries NO eos_token_id and ships no generation_config.json, and its tokenizer.json post_processor is a bare ByteLevel with no TemplateProcessing template, so Tokenizer::FromHfJson leaves eos_id_ = -1 (ExtractBosEos only reads post_processor). InputProcessor then finds no eos anywhere (input_processor.cpp:42-67) and every request goes out with eos_token_id unset: the model emits <|im_end|> (id 248046) at token 10 of a no-think reply and the engine ignores it, generating fake 'user\n\nassistant\n' turns to max_tokens (finish_reason=length). Upstream resolves the same id from the tokenizer itself: HF eos_token_id comes from tokenizer_config.json's eos_token ('<|im_end|>'), which FromHfJson never reads. User-visible symptom: thinking looks inconsistent — every answer ends by hallucinating more think blocks because <|im_end|> never terminates the turn.
Row: MODEL-QWEN35-EXL3
Local-Issue: ISSUE-LOCAL-01M3RY20H90NMK37V74EH34Z77
Kind: bug
Reproduced live on the vllm-qwen35-9b-exl3 container (Qwen3.5-9B-EXL3-4.00bpw, ROCm gfx1101, 2026-09-30). The checkpoint's config.json carries NO eos_token_id and ships no generation_config.json, and its tokenizer.json post_processor is a bare ByteLevel with no TemplateProcessing template, so Tokenizer::FromHfJson leaves eos_id_ = -1 (ExtractBosEos only reads post_processor). InputProcessor then finds no eos anywhere (input_processor.cpp:42-67) and every request goes out with eos_token_id unset: the model emits <|im_end|> (id 248046) at token 10 of a no-think reply and the engine ignores it, generating fake 'user\n\nassistant\n' turns to max_tokens (finish_reason=length). Upstream resolves the same id from the tokenizer itself: HF eos_token_id comes from tokenizer_config.json's eos_token ('<|im_end|>'), which FromHfJson never reads. User-visible symptom: thinking looks inconsistent — every answer ends by hallucinating more think blocks because <|im_end|> never terminates the turn.