docs(labs): Qwen3.8-Flash-Next on Gufo on one Strix Halo - #1966
Merged
Defilan merged 2 commits intoOct 2, 2026
Merged
Conversation
Add a lab build page for the Strix Halo endpoint after its move from llama.cpp Vulkan to Gufo v0.5.0 at 262K context: the bake-off against llama.cpp tuned four ways, the quality gate, the thinking-on prompt cache A/B and 3 hour soak behind the gufo-org/gufo#335 fix, the exact-restore check, memory, and the generic-runtime staging gap fixed by defilantech#1962. Shows the deployed PVC-mount InferenceService and the stageModel form for LLMKube 0.10.2 and later. Add the page to the Lab builds nav and index, and point the llama.cpp Vulkan page at it as the current engine on the box. Signed-off-by: Christopher Maher <chris@mahercode.io>
The v0.3.0 soak record has its cumulative counts: 0 stalls, 0 errors, 0 empty replies and 0 repetitive-tail replies over 102 turns. Signed-off-by: Christopher Maher <chris@mahercode.io>
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds a lab build page,
docs/site/labs/qwen38-flash-next-gufo-on-strix-halo.md, for the Strix Halo endpoint after it moved from llama.cpp Vulkan to Gufo v0.5.0 at the full 262,144-token context. The page follows the structure of the existing Strix page: why this combination, the runtime (imagellmkube-gufo-rocm-gfx1151@sha256:3055ba7d..., Gufo v0.5.0 at23cacbb9), hardware, weights, the InferenceService in two forms (the deployed read-only PVC mount on 0.10.1, and thestageModelform for 0.10.2 and later), measured results with the client path and prompt shape for each, what went wrong, reproducing, and related links.Also adds the page to the Lab builds section of
docs/site/_meta/nav.yaml(after the existing Strix entry) and to the Builds table indocs/site/labs/index.md, and puts a short "superseded on this box" note at the top ofqwen38-flash-next-mtp-on-strix-halo.mdpointing to the new page. The old page is otherwise unchanged.Why
The Strix endpoint changed engine, quant and context, and the existing lab page now describes a build that is no longer serving. The new page records the bake-off numbers (cold 250K prefill 101 to 1,184 tok/s, 41 min to 3.5 min to first token; decode about 7% slower single-stream), the quality gate, the thinking-on prompt-cache A/B and 3 hour soak behind the Gufo team's fix for gufo-org/gufo#335 (median TTFT at 180K to 241K tokens 199.2 s to 4.1 s), the exact-restore check, and the LLMKube generic-runtime staging gap that #1961 / #1962 closed.
Refs #1961
How
Docs only. Every number comes from the bake-off's recorded results and the applied manifests, restated with its method (median, best-of, single run, harness settings). Internal hostnames, addresses and private repository paths are scrubbed; the S3 source in the
stageModelexample uses a generic bucket path and a placeholder Secret name. ThestageModelexample uses the attested v0.5.0 digest. It requires LLMKube 0.10.2 (release PR #1959, not yet released), and the page says so.Checks run locally:
nav.yamlparses as YAML; every YAML block on the new page parses (1 and 3 documents); internal/docs/...links resolve to existing nav slugs; no em dashes or " :: " in the changed files. The repo has no markdown lint or link-check target, so none was run. The manifests were not dry-run against a cluster.Checklist
make testpasses locally: n/a, docs only, no Go changesmake lintpasses locally: n/a, docs only, no Go changesgit commit -s) per DCOAssisted-by: Claude Code (Claude Opus 5.5)