Skip to content

Latest commit

 

History

95 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Local LLM Agent

Run a Qwen model on a DGX Spark and connect to it from macOS using Agent Canvas (native, no Docker). Push notifications of agent activity to your phone via ntfy are included (ntfy/ + agent_canvas_native/README.md → Notifications).

The DGX side offers three serving options, all on port 8000 (run one at a time):

  • qwen38-27b-nvfp4 — the current NVFP4 quantization from NVIDIA (default),
  • qwen38-27b-bf16 — the original BF16 Qwen checkpoint, and
  • flash_ultrafast — the Qwen3.8 Flash DGX UltraFast v16b recipe.

See dgx_spark_host/ → The three stacks.

dgx_spark_host/

Host a Qwen model via vLLM on the DGX Spark machine — one self-contained stack per model, in its own folder (own image, compose, entrypoint and defaults; they share no code or parameters). The two 27B stacks build their image from a Dockerfile; flash_ultrafast/ pins a prebuilt patched image produced by its setup-upstream.sh, so it has no Dockerfile. All bind port 8000; run only ONE at a time.

The three stacks

Stack Serves How it's run Served alias
qwen38-27b-nvfp4/ nvidia/Qwen3.8-27B-NVFP4 (NVFP4+FP8, ~22 GB) — the current default cd dgx_spark_host/qwen38-27b-nvfp4 && docker compose up --build qwen-local
qwen38-27b-bf16/ Qwen/Qwen3.8-27B (official BF16, ~55 GB) cd dgx_spark_host/qwen38-27b-bf16 && docker compose up --build qwen-local
flash_ultrafast/ Qwen3.8 Flash DGX UltraFast v16b recipe cd dgx_spark_host/flash_ultrafast && ./setup-upstream.sh && docker compose up --build qwen-local

Run the current default:

cd dgx_spark_host/qwen38-27b-nvfp4
docker compose -f compose.yml up --build

The API is available at http://localhost:8000/v1. For shared networks, bind to 127.0.0.1 in the stack's compose.yml.

Each stack's README.md documents its own environment variables (defaults live in its Dockerfile for the two 27B stacks, and in compose.yml/entrypoint.sh for flash_ultrafast; override via .env or compose.yml) and tuning notes (DGX Spark, GB10, 128 GB unified memory, LLM-only box). Both 27B checkpoints ship a 1-layer MTP head, so MTP speculative decoding is on by default for them; both are native vision-language models with image inputs enabled via --limit-mm-per-prompt '{"image":4}'.

flash_ultrafast/ — the Qwen3.8 Flash DGX UltraFast option

  • What it is: the dime-online/qwen3.8-Flash-DGX-UltraFast v16b recipe — a patched vLLM image (CUDA 13.0, custom low-latency GEMM/Mamba/PLE/MTP kernels) serving the W4A16/FP8 AutoRound-hybrid Qwen3.8-Flash-Next checkpoint. The PLE table is memory-mapped from storage, which leaves room for a 16 GB KV pool at the full 262,144-token context. A dense T80 MTP drafter (depth 3, block rejection) provides the speed — upstream reports 74 tok/s single stream and 212 tok/s aggregate at 8 streams on one GB10 (re-verify on your hardware).
  • Why fully isolated: it uses a different (patched) image, extra downloads (~130 GB), and a drafter build — so it has its own folder with its own compose/entrypoint/pinned env, rather than sharing parameters with the two 27B stacks. It is served on the same port 8000 with the same alias, qwen-local as the other two — clients need no changes.
  • One-time setup (on the Spark): ./dgx_spark_host/flash_ultrafast/setup-upstream.sh — clones the upstream Apache-2.0 repo, downloads the pinned checkpoint + PLE table, builds the patched image and the T80 drafter, and installs the draft vocabulary.
  • Run: cd dgx_spark_host/flash_ultrafast && docker compose -f compose.yml up --build. Stop the other stacks first — the three stacks share port 8000.

Full details, provenance, the upstream's claims, and the overridable parameters are in dgx_spark_host/flash_ultrafast/README.md.

agent_canvas_native/

Run Agent Canvas natively: UI + agent-server + automation server + ingress run as local processes via Node.js and uv, with no Docker. Agents run as your user on the local filesystem (there is no container sandbox).

Run

./agent_canvas_native/install.sh   # one-time (upgrades on re-run)
./agent_canvas_native/run.sh
  • Address http://localhost:8020 (avoids 8000, the vLLM tunnel).
  • LLM profile: Settings → LLM, provider OpenAI-compatible, base http://localhost:8000/v1, key local-dgx-key, model qwen-local (the alias all three DGX stacks serve — the client configuration never changes when switching stacks).
  • Optional ntfy push notifications for the phone: with NTFY_ENABLED=true in agent_canvas_native/.env, run.sh also starts the notifier daemon (and the per-PC ntfy server from ntfy/) that pings you when your agent finishes, needs input, or errors. See agent_canvas_native/README.md → Notifications (ntfy).
Variable Default Description
AGENT_CANVAS_PORT 8020 Ingress (UI + proxied API) port
AGENT_CANVAS_STATE ./openhands-state Where agent-server keeps per-conversation runtime state (conversations, workspaces, terminal history, logs). API key + LLM profile live in ~/.openhands.
VLLM_BASE_URL http://localhost:8000/v1 Endpoint run.sh checks before launch (the value for the LLM profile)
VLLM_API_KEY local-dgx-key API key for the vLLM preflight check

Needs Node.js ≥ 22.12 and uv. Ports, overrides, and troubleshooting are documented in agent_canvas_native/README.md.

ntfy/

Per-PC ntfy push server (single Docker container, port 2020). Pairs with the Agent Canvas notifier (agent_canvas_native/ntfy_notifier.py) to send push notifications to your phone when an agent finishes a turn, needs input, or hits an error. Designed for multiple PCs on Tailscale, each running its own server + notifier: the notifier publishes to 127.0.0.1:2020, the phone subscribes over Tailscale at http://<pc>.tail:2020/<topic>.

cd ntfy
cp example.env .env     # set NTFY_BASE_URL=http://<this-pc>.tail:2020
docker compose up -d    # healthcheck via docker compose ps

Security model (unguessable topic = credential on Tailscale), optional account auth, phone setup (Android instant delivery / iOS relay) and the Firebase/custom-APK caveat are documented in ntfy/README.md. Enable end-to-end notifications via NTFY_ENABLED=true in agent_canvas_native/.env (see agent_canvas_native/README.md → Notifications (ntfy)).

About

DGX Spark + OpenHands setup

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Used by

Contributors

Languages