Skip to content

Pin llama.cpp fa226f9: ZAYA chat templates tokenize; HRX all-NaN router row is loud (engine#123) - #175

Merged
bong-water-water-bong merged 1 commit into
mainfrom
pin/hrx-nan-loud
Sep 27, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
pin/hrx-nan-loud

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves third_party/llama.cpp from 895d63f to fa226f9, which brings in two fork PRs.

llama.cpp #27: ZAYA chat templates tokenize.

  • Every ZAYA GGUF we converted stored <|im_start|> (105) and <eos> (1) as ordinary tokens. llama.cpp only matches control and user-defined tokens in prompt text, so each chat turn's <|im_start|> reached the model as seven text tokens, and ZAYA1-8B chat answers went off-template ("12 + 30" gave "22").
  • The converter now marks every token tokenizer.json calls special as a control token. The ZAYA GGUFs on our Hugging Face org were fixed in place (2 bytes of metadata each; tensors unchanged): 8B, VL-8B, base, reasoning-base, with the 74B finishing now.
  • Checked: <|im_start|>user now tokenizes as [2, 105, 2364], as transformers does. With thinking off, our F16 matches Zyphra's BF16 model on the reference prompts: "7 + 8" gives "25", which is the model's own answer, not ours.

llama.cpp #26: an all-NaN MoE router row is loud instead of silently decoding a wrong expert (engine#123 follow-up).

  • Checked: perplexity on HRX0 is identical to before: Qwen3-Coder-30B 17.2526, ZAYA1-8B 25.2474 (8 chunks). Chat answers are right on HRX0 and Vulkan0 with no false aborts.

ZAYA1-8B end to end through 1bit serve at this pin:

  • serve_e2e passes on vulkan, hrx and auto.
  • Streaming, 4 concurrent requests and a 5,400-token prompt all work on Vulkan and HRX.

Also here: the registry is regenerated for the pin (counts unchanged). docs/hrx.md covers #26, and docs/vulkan.md warns that ZAYA GGUFs converted before today need re-downloading.

🤖 Generated with Claude Code

…HRX makes an all-NaN MoE router row loud (llama.cpp #26, engine#123)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 27, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 7c37201

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ No major issues detected

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant