Run a Qwen model on a DGX Spark and connect to it from macOS using
Agent Canvas
(native, no Docker).
Push notifications of agent activity to your phone via
ntfy are included (ntfy/ + agent_canvas_native/README.md
→ Notifications).
The DGX side offers three serving options, all on port 8000 (run one at a time):
qwen38-27b-nvfp4— the current NVFP4 quantization from NVIDIA (default),qwen38-27b-bf16— the original BF16 Qwen checkpoint, andflash_ultrafast— the Qwen3.8 Flash DGX UltraFast v16b recipe.
See dgx_spark_host/ → The three stacks.
Host a Qwen model via vLLM on the DGX Spark machine — one self-contained stack
per model, in its own folder (own image, compose, entrypoint and defaults; they
share no code or parameters). The two 27B stacks build their image from a
Dockerfile; flash_ultrafast/ pins a prebuilt patched image produced by its
setup-upstream.sh, so it has no Dockerfile. All bind port 8000; run only ONE
at a time.
| Stack | Serves | How it's run | Served alias |
|---|---|---|---|
qwen38-27b-nvfp4/ |
nvidia/Qwen3.8-27B-NVFP4 (NVFP4+FP8, ~22 GB) — the current default |
cd dgx_spark_host/qwen38-27b-nvfp4 && docker compose up --build |
qwen-local |
qwen38-27b-bf16/ |
Qwen/Qwen3.8-27B (official BF16, ~55 GB) |
cd dgx_spark_host/qwen38-27b-bf16 && docker compose up --build |
qwen-local |
flash_ultrafast/ |
Qwen3.8 Flash DGX UltraFast v16b recipe | cd dgx_spark_host/flash_ultrafast && ./setup-upstream.sh && docker compose up --build |
qwen-local |
Run the current default:
cd dgx_spark_host/qwen38-27b-nvfp4
docker compose -f compose.yml up --buildThe API is available at http://localhost:8000/v1. For shared networks, bind
to 127.0.0.1 in the stack's compose.yml.
Each stack's README.md documents its own environment variables (defaults live
in its Dockerfile for the two 27B stacks, and in compose.yml/entrypoint.sh
for flash_ultrafast; override via .env or compose.yml) and tuning
notes (DGX Spark, GB10, 128 GB unified memory, LLM-only box). Both 27B
checkpoints ship a 1-layer MTP head, so MTP speculative decoding is on by
default for them; both are native vision-language models with image inputs
enabled via --limit-mm-per-prompt '{"image":4}'.
- What it is: the dime-online/qwen3.8-Flash-DGX-UltraFast
v16b recipe — a patched vLLM image (CUDA 13.0, custom low-latency
GEMM/Mamba/PLE/MTP kernels) serving the W4A16/FP8 AutoRound-hybrid
Qwen3.8-Flash-Nextcheckpoint. The PLE table is memory-mapped from storage, which leaves room for a 16 GB KV pool at the full 262,144-token context. A dense T80 MTP drafter (depth 3, block rejection) provides the speed — upstream reports 74 tok/s single stream and 212 tok/s aggregate at 8 streams on one GB10 (re-verify on your hardware). - Why fully isolated: it uses a different (patched) image, extra
downloads (~130 GB), and a drafter build — so it has its own folder with its
own compose/entrypoint/pinned env, rather than sharing parameters with the
two 27B stacks. It is served on the same port 8000 with the same alias,
qwen-localas the other two — clients need no changes. - One-time setup (on the Spark):
./dgx_spark_host/flash_ultrafast/setup-upstream.sh— clones the upstream Apache-2.0 repo, downloads the pinned checkpoint + PLE table, builds the patched image and the T80 drafter, and installs the draft vocabulary. - Run:
cd dgx_spark_host/flash_ultrafast && docker compose -f compose.yml up --build. Stop the other stacks first — the three stacks share port 8000.
Full details, provenance, the upstream's claims, and the overridable
parameters are in dgx_spark_host/flash_ultrafast/README.md.
Run Agent Canvas natively: UI + agent-server + automation server + ingress run as local processes via Node.js and uv, with no Docker. Agents run as your user on the local filesystem (there is no container sandbox).
./agent_canvas_native/install.sh # one-time (upgrades on re-run)
./agent_canvas_native/run.sh- Address
http://localhost:8020(avoids 8000, the vLLM tunnel). - LLM profile: Settings → LLM, provider OpenAI-compatible, base
http://localhost:8000/v1, keylocal-dgx-key, modelqwen-local(the alias all three DGX stacks serve — the client configuration never changes when switching stacks). - Optional ntfy push notifications for the phone: with
NTFY_ENABLED=trueinagent_canvas_native/.env,run.shalso starts the notifier daemon (and the per-PC ntfy server fromntfy/) that pings you when your agent finishes, needs input, or errors. Seeagent_canvas_native/README.md→ Notifications (ntfy).
| Variable | Default | Description |
|---|---|---|
AGENT_CANVAS_PORT |
8020 |
Ingress (UI + proxied API) port |
AGENT_CANVAS_STATE |
./openhands-state |
Where agent-server keeps per-conversation runtime state (conversations, workspaces, terminal history, logs). API key + LLM profile live in ~/.openhands. |
VLLM_BASE_URL |
http://localhost:8000/v1 |
Endpoint run.sh checks before launch (the value for the LLM profile) |
VLLM_API_KEY |
local-dgx-key |
API key for the vLLM preflight check |
Needs Node.js ≥ 22.12 and uv. Ports, overrides, and troubleshooting are
documented in agent_canvas_native/README.md.
Per-PC ntfy push server (single Docker container, port
2020). Pairs with the Agent Canvas notifier
(agent_canvas_native/ntfy_notifier.py) to send push notifications to your
phone when an agent finishes a turn, needs input, or hits an error. Designed
for multiple PCs on Tailscale, each running its own server + notifier:
the notifier publishes to 127.0.0.1:2020, the phone subscribes over
Tailscale at http://<pc>.tail:2020/<topic>.
cd ntfy
cp example.env .env # set NTFY_BASE_URL=http://<this-pc>.tail:2020
docker compose up -d # healthcheck via docker compose psSecurity model (unguessable topic = credential on Tailscale), optional
account auth, phone setup (Android instant delivery / iOS relay) and the
Firebase/custom-APK caveat are documented in
ntfy/README.md. Enable end-to-end notifications via
NTFY_ENABLED=true in agent_canvas_native/.env (see
agent_canvas_native/README.md → Notifications (ntfy)).