Skip to content

docs(labs): Qwen3.8-Flash-Next on Gufo on one Strix Halo - #1966

Merged
Defilan merged 2 commits into
defilantech:mainfrom
Defilan:docs/lab-qwen38-gufo-strix-halo
Oct 2, 2026
Merged

Defilan merged 2 commits into
defilantech:mainfrom
Defilan:docs/lab-qwen38-gufo-strix-halo

Conversation

@Defilan

@Defilan Defilan commented Oct 2, 2026

Copy link
Copy Markdown
Member

What

Adds a lab build page, docs/site/labs/qwen38-flash-next-gufo-on-strix-halo.md, for the Strix Halo endpoint after it moved from llama.cpp Vulkan to Gufo v0.5.0 at the full 262,144-token context. The page follows the structure of the existing Strix page: why this combination, the runtime (image llmkube-gufo-rocm-gfx1151@sha256:3055ba7d..., Gufo v0.5.0 at 23cacbb9), hardware, weights, the InferenceService in two forms (the deployed read-only PVC mount on 0.10.1, and the stageModel form for 0.10.2 and later), measured results with the client path and prompt shape for each, what went wrong, reproducing, and related links.

Also adds the page to the Lab builds section of docs/site/_meta/nav.yaml (after the existing Strix entry) and to the Builds table in docs/site/labs/index.md, and puts a short "superseded on this box" note at the top of qwen38-flash-next-mtp-on-strix-halo.md pointing to the new page. The old page is otherwise unchanged.

Why

The Strix endpoint changed engine, quant and context, and the existing lab page now describes a build that is no longer serving. The new page records the bake-off numbers (cold 250K prefill 101 to 1,184 tok/s, 41 min to 3.5 min to first token; decode about 7% slower single-stream), the quality gate, the thinking-on prompt-cache A/B and 3 hour soak behind the Gufo team's fix for gufo-org/gufo#335 (median TTFT at 180K to 241K tokens 199.2 s to 4.1 s), the exact-restore check, and the LLMKube generic-runtime staging gap that #1961 / #1962 closed.

Refs #1961

How

Docs only. Every number comes from the bake-off's recorded results and the applied manifests, restated with its method (median, best-of, single run, harness settings). Internal hostnames, addresses and private repository paths are scrubbed; the S3 source in the stageModel example uses a generic bucket path and a placeholder Secret name. The stageModel example uses the attested v0.5.0 digest. It requires LLMKube 0.10.2 (release PR #1959, not yet released), and the page says so.

Checks run locally: nav.yaml parses as YAML; every YAML block on the new page parses (1 and 3 documents); internal /docs/... links resolve to existing nav slugs; no em dashes or " :: " in the changed files. The repo has no markdown lint or link-check target, so none was run. The manifests were not dry-run against a cluster.

Checklist

  • Tests added/updated: n/a, docs only
  • make test passes locally: n/a, docs only, no Go changes
  • make lint passes locally: n/a, docs only, no Go changes
  • Commit messages follow conventional commits
  • All commits are signed off (git commit -s) per DCO
  • AI assistance (if any) is disclosed above, per CONTRIBUTING.md
  • Documentation updated (if user-facing change)

Assisted-by: Claude Code (Claude Opus 5.5)

Add a lab build page for the Strix Halo endpoint after its move from
llama.cpp Vulkan to Gufo v0.5.0 at 262K context: the bake-off against
llama.cpp tuned four ways, the quality gate, the thinking-on prompt
cache A/B and 3 hour soak behind the gufo-org/gufo#335 fix, the
exact-restore check, memory, and the generic-runtime staging gap fixed
by defilantech#1962. Shows the deployed PVC-mount InferenceService and the
stageModel form for LLMKube 0.10.2 and later.

Add the page to the Lab builds nav and index, and point the llama.cpp
Vulkan page at it as the current engine on the box.

Signed-off-by: Christopher Maher <chris@mahercode.io>
The v0.3.0 soak record has its cumulative counts: 0 stalls, 0 errors,
0 empty replies and 0 repetitive-tail replies over 102 turns.

Signed-off-by: Christopher Maher <chris@mahercode.io>
@Defilan
Defilan merged commit 0a66ad6 into defilantech:main Oct 2, 2026
9 checks passed
@codecov

codecov Bot commented Oct 2, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant