diff --git a/docs/upstream-changes.md b/docs/upstream-changes.md index 4e875d19..7d3515ff 100644 --- a/docs/upstream-changes.md +++ b/docs/upstream-changes.md @@ -4,6 +4,14 @@ Track announced llama.cpp changes that may require coordinated Llama-GUI updates ## Pending +### Direct reads for lazy embedding tables (`--lazy-mode on-direct`) + +- **Upstream:** [ggml-org/llama.cpp#28136](https://github.com/ggml-org/llama.cpp/pull/28136). +- **Status:** Open, not merged (checked 2026-09-14). Current upstream [`common/arg.cpp`](https://github.com/ggml-org/llama.cpp/blob/master/common/arg.cpp) accepts only `auto`, `on`, and `off` for `--lazy-mode` / `-lzm`; it rejects `on-direct`. Using the proposed value currently requires a custom build containing the PR. +- **Behavior:** Reads needed rows from eligible per-layer embedding tables explicitly instead of relying on memory-mapped page faults. This aims to improve prompt processing when the tables are not already cached; performance depends on hardware and workload. +- **Local handling:** Added `on-direct` to the existing Configure → Advanced **Lazy Mode** dropdown (`tensor_read_lazy` in `ui/js/flags/definitions-server.js`) on 2026-09-14 at the maintainer's request, ahead of upstream merge. The option is labelled experimental and requires a custom build; Auto remains the default. Help text explains that current upstream builds reject the value and the PR currently falls back to lazy mmap reads on Windows. The existing `--lazy-mode` flag remains available for upstream-supported values. +- **Recheck after merge:** Verify the final enum values, platform support, load-mode requirements, and first supported release, then update the experimental label and help text. Confirm support using the installed binary's `--help` and account for older builds that reject the value. + ### Legacy load flags removed in favor of --load-mode - - - Done - **Upstream:** b10875, Sep 9 (PR #28334): `--mmap` / `--no-mmap`, `--mlock`, and `-dio` / `-ndio` / `--direct-io` / `--no-direct-io` removed from `llama-cli` / `llama-server` in favor of `--load-mode`. Installed b10917 `--help` output confirms `llama-bench` / `llama-perplexity` also only advertise `--load-mode`. diff --git a/ui/js/flags/definitions-server.js b/ui/js/flags/definitions-server.js index 8dd9770e..d554307f 100644 --- a/ui/js/flags/definitions-server.js +++ b/ui/js/flags/definitions-server.js @@ -362,12 +362,13 @@ const FLAG_DEFINITIONS_SERVER = [ type: "enum", label: "Lazy Mode", short_desc: "Read eligible model tensors from disk on demand to reduce resident RAM use.", - desc: "Controls on-demand reading for tensors marked as eligible by the model architecture, such as large per-layer embedding tables. Requires mmap. Auto uses llama.cpp's default behavior (lazy only above 4 GiB); On trades some inference speed for lower resident RAM; Off keeps eligible tensors resident.", + desc: "Controls on-demand reading for tensors marked as eligible by the model architecture, such as large per-layer embedding tables. Auto uses llama.cpp's default behavior (lazy only above 4 GiB); On reads eligible tensors on demand using mmap; Off keeps them resident. On Direct reads per-layer embedding rows explicitly and may improve uncached prompt processing. Experimental: requires a custom build with llama.cpp PR #28136; current upstream builds reject on-direct. The PR currently falls back to lazy mmap reads on Windows.", tool: "both", default: "", options: [ { value: "", label: "Auto (Recommended)" }, { value: "on", label: "On (Lower RAM)" }, + { value: "on-direct", label: "On Direct (Experimental, Custom Build)" }, { value: "off", label: "Off (Keep Resident)" }, ], },