Skip to content

convert: ZAYA special tokens are CONTROL, so chat templates tokenize - #27

Merged
bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/zaya-special-tokens
Sep 27, 2026
Merged

bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/zaya-special-tokens

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

Bug: in every ZAYA GGUF our converter wrote, <|im_start|> (id 105) and <eos> (id 1) had token type NORMAL, although tokenizer.json marks both special: true (LlamaHfVocab left them normal; both are also in the base vocab). llama.cpp matches special-token text in a prompt only for CONTROL and USER_DEFINED tokens. So every chat turn's <|im_start|> reached the model as seven text tokens instead of token 105:

llama (before): [2, 236820, 236909, 548, 236779, 3041, 111038, 2364]   <|im_start|>user
transformers:   [2, 105, 2364]

Chat answers were off-template: "What is 12 + 30?" gave "22". The teacher-forced checks never saw this, because they feed token IDs directly.

Fix: every token tokenizer.json lists as special becomes CONTROL, for ZAYA and ZAYA1-VL.

Checked:

  • Converting ZAYA1-8B with this commit gives the published GGUF's 1,283 tensors unchanged. The token types match the published file with tokens 1 and 105 set to CONTROL. The only other differences are general.* model-card fields (base model, licence, name).
  • llama-server now tokenizes <|im_start|>user as [2, 105, 2364].
  • With thinking off, on questions where Q4_K_M and BF16 both land on the same token, our F16 answers match Zyphra's model in transformers (BF16): "7 + 8" → "25" (the model's own answer), "100 + 23" → "123", "17 × 23" → "391".

The published GGUFs on Hugging Face get the same two-byte metadata fix in place, so they needn't be reconverted.

🤖 Generated with Claude Code

LlamaHfVocab left <|im_start|> (105) and <eos> (1) NORMAL although
tokenizer.json marks them special. llama.cpp only matches special-token
text for CONTROL and USER_DEFINED tokens, so every chat turn's
<|im_start|> reached the model as seven text tokens and ZAYA1-8B answered
off-template (12 + 30 -> 22). Tokens tokenizer.json lists as special are
now CONTROL. Converting ZAYA1-8B now gives the published GGUF's 1283
tensors unchanged and token types with 1 and 105 CONTROL.
@bong-water-water-bong
bong-water-water-bong merged commit fa226f9 into 1bit/hrx-vulkan-patched Sep 27, 2026
1 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant