Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
71 commits
Select commit Hold shift + click to select a range
fcaac3d
expert heatmap: decay-tracked usage counters per layer
miltos22 Aug 2, 2026
6108023
expert heatmap: add top-S ranking, log top-8 per layer
miltos22 Aug 2, 2026
21537c1
expert heatmap: add --expert-hot-s flag for GPU hot store slot count
miltos22 Aug 2, 2026
b27d2f5
expert heatmap: fix GPU tensor readback via ggml_set_output and synch…
miltos22 Aug 2, 2026
413e65d
expert heatmap: move readback logic into heatmap module
miltos22 Aug 2, 2026
3fd5778
Merge remote-tracking branch 'origin/master'
miltos22 Aug 2, 2026
23ff80e
expert hotstore: add per-layer expert slot sizing
miltos22 Aug 2, 2026
289c41e
expert hotstore: allocate GPU hot store buffers for S slots
miltos22 Aug 2, 2026
284215c
expert hotstore: reduced cross contamination when expert args off
miltos22 Aug 2, 2026
8e95570
expert heatmap: count updates in tokens not layers
miltos22 Aug 2, 2026
08e657d
expert heatmap: fix decay and log trigger
miltos22 Aug 2, 2026
1aefbbe
expert hotstore: copy top-S experts to GPU after first ubatch
miltos22 Aug 3, 2026
cee1740
expert hotstore: re-sync hot slots on cadence (stable slots)
miltos22 Aug 3, 2026
a3c4425
expert hotstore: add sentinel slot for zero-contribution routing
miltos22 Aug 3, 2026
9040902
expert hotstore: per-layer LUTs and masks for in-graph routing
miltos22 Aug 3, 2026
fb0b039
Merge branch 'ggml-org:master' into master
miltos22 Aug 3, 2026
8d1a68e
Merge remote-tracking branch 'origin/master'
miltos22 Aug 3, 2026
3c9616b
expert tier: hook GPU hot store into graph with MUL_MAT_ID_COLD cold op
miltos22 Aug 3, 2026
d1cdf1b
ggml-cpu: move MUL_MAT_ID_COLD kernel to its own file
miltos22 Aug 3, 2026
8687ee2
common: added -las short flag for expert hot store slots (alias of --…
miltos22 Aug 3, 2026
6e75753
Merge branch 'ggml-org:master' into master
miltos22 Aug 3, 2026
5c5020d
expert hotstore: freeze swapping during multi-slot batches
miltos22 Aug 3, 2026
1a850ec
llama : gate expert hot store to CUDA only, with force override
miltos22 Aug 4, 2026
b18c426
common: expert hot store manual slots activate --cmoe
miltos22 Aug 4, 2026
f201ef5
common: autofit expert hot store slots via --expert-hot-s -1
miltos22 Aug 4, 2026
ca624bc
expert hot store: hysteresis gate for slot swaps
miltos22 Aug 4, 2026
fac401b
Merge branch 'ggml-org:master' into master
miltos22 Aug 4, 2026
0b7687a
merge upstream master into local fork
miltos22 Aug 4, 2026
7633d5d
common: fix expert flag help text (decay default, autofit -1)
miltos22 Aug 4, 2026
25558af
common: rename -las short flag for expert hot store slots to -ehs
miltos22 Aug 4, 2026
0170536
ggml: fix windows build of the cold op (missing stdatomic include)
miltos22 Aug 4, 2026
336beab
Merge branch 'ggml-org:master' into master
miltos22 Aug 4, 2026
919168e
ggml: use MSVC Interlocked fallback for atomics in the cold op
miltos22 Aug 4, 2026
d21fcbd
Merge branch 'ggml-org:master' into master
miltos22 Aug 5, 2026
e26868d
Adjusted sentinel autofit to be more conservative though properly acc…
miltos22 Aug 4, 2026
08562e5
llama : added -ecf/--expert-cache-force to enable hot store on non-cuda
miltos22 Aug 5, 2026
d3cba24
After mobilinkd's recomendation, skip expert cache upkeep during pref…
miltos22 Aug 5, 2026
4ce955a
llama : rename -ecf to --ecf and document expert cache flags
miltos22 Aug 5, 2026
3897637
Merge branch 'ggml-org:master' into master
miltos22 Aug 6, 2026
d7345f7
expert tier: per-tensor hot store with tunable cadence and i32 cpu mask
miltos22 Aug 6, 2026
581cbb0
wip: multi-gpu hot store with sentinel mask
miltos22 Aug 6, 2026
e0efcb3
expert tier: add per-slice cpu pool for exps weights
miltos22 Aug 6, 2026
e3004a9
expert tier: cold op reads slices from the per-slice pool
miltos22 Aug 6, 2026
09b1e45
expert tier: count+rank batching and native memory moves
miltos22 Aug 6, 2026
396412c
expert tier: hash-verify gpu move-ins with deferred cpu release
miltos22 Aug 6, 2026
20f1564
expert tier: shrink move-in hash sample to 1024 bytes
miltos22 Aug 6, 2026
6b75226
expert tier: verified swap handshake, start-up full sync, no-evict
miltos22 Aug 7, 2026
cc0de31
expert tier: fused cold path with no sentinel mask
miltos22 Aug 9, 2026
8e3e8e1
expert tier: zero the store buffer at preload
miltos22 Aug 9, 2026
d294939
expert tier: drop dead heatmap paths and debug hash hook
miltos22 Aug 9, 2026
ab875d2
common: warn when -ehs -1 cannot autofit expert slots
miltos22 Aug 9, 2026
1550f11
expert tier: rework mmap page hints into a pin manager
miltos22 Aug 9, 2026
ef1af4b
expert tier: pin flag, drop cache force, log heatmap at generation end
miltos22 Aug 9, 2026
a435e6c
expert tier: add --expert-copy, fix adaptive cadence
miltos22 Aug 9, 2026
27a4e50
expert tier: d2d promote/demote, heatmap sidecar, frozen store
miltos22 Aug 9, 2026
bfe48ad
expert tier: refine d2d swap gating and dwell units
miltos22 Aug 9, 2026
8e34071
expert tier: autofit copy/move mode, d2d scratch vram offset
miltos22 Aug 9, 2026
be76d2d
expert tier: copy mode for mmap, host-pool d2d, autofit abort warning
miltos22 Aug 9, 2026
5097df0
expert tier: warm top 70% of cold experts by default
miltos22 Aug 9, 2026
7c08c59
sync with upstream
miltos22 Aug 9, 2026
75afc97
expert tier: pin 30 default, sidecar no-freeze, ram-pressure eviction
miltos22 Aug 9, 2026
f2b29ef
ggml: bump rpc patch version, fix msvc prefetch in cold op
miltos22 Aug 9, 2026
813353f
expert tier: windows portability for hotstore and preload
miltos22 Aug 9, 2026
efa09e5
server: document expert tier flags
miltos22 Aug 9, 2026
5e5a498
expert tier: fix werror issues, revert expert-gpu list, cli docs
miltos22 Aug 9, 2026
48f107d
expert tier: fix windows and macos portability
miltos22 Aug 9, 2026
ea63c7f
fix unused param and export cold slice fn
miltos22 Aug 9, 2026
b0baf88
fix cross platform compiler errors
miltos22 Aug 9, 2026
86beff0
export preload setters for common lib
miltos22 Aug 10, 2026
118554a
expert tier: rename mode flag to expert-move-mode
miltos22 Aug 10, 2026
4401075
Merge branch 'master' of https://github.com/ggml-org/llama.cpp
miltos22 Aug 10, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion common/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ configure_file(${TEMPLATE_FILE} ${OUTPUT_FILE})
set(TARGET llama-common-base)
add_library(${TARGET} STATIC ${OUTPUT_FILE})

target_include_directories(${TARGET} PUBLIC .)
target_include_directories(${TARGET} PUBLIC . ../src)

if (BUILD_SHARED_LIBS)
set_target_properties(${TARGET} PROPERTIES POSITION_INDEPENDENT_CODE ON)
Expand Down
100 changes: 100 additions & 0 deletions common/arg.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@
#include "download.h"
#include "json-schema-to-grammar.h"
#include "llama.h"
#include "llama-expert-preload.h"
#include "log.h"
#include "sampling.h"
#include "speculative.h"
Expand Down Expand Up @@ -35,6 +36,7 @@
#include <regex>
#include <set>
#include <string>
#include <cstring>
#include <thread> // for hardware_concurrency
#include <vector>

Expand Down Expand Up @@ -884,6 +886,22 @@ static bool common_params_parse_ex(int argc, char ** argv, common_params_context
params.cors_origins = "localhost";
}

// manual hot store slots need all MoE weights in the CPU (host pointers);
// auto-activate -cmoe unless the user already did (or wants autofit slots)
if (params.expert_hot_s > 0) {
bool has_cmoe = false;
for (const auto & o : params.tensor_buft_overrides) {
if (o.pattern != nullptr && strcmp(o.pattern, LLM_FFN_EXPS_REGEX) == 0) {
has_cmoe = true;
break;
}
}
if (!has_cmoe) {
params.tensor_buft_overrides.push_back(llm_ffn_exps_cpu_override());
LOG_WRN("manually selecting --expert-hot-s slots activates --cmoe (all MoE weights kept in the CPU)\n");
}
}

// pad tensor_buft_overrides for llama_params_fit:
const size_t ntbo = llama_max_tensor_buft_overrides();
while (params.tensor_buft_overrides.size() < ntbo) {
Expand Down Expand Up @@ -2679,6 +2697,88 @@ common_params_context common_params_parser_init(common_params & params, llama_ex
}
}
).set_env("LLAMA_ARG_N_CPU_MOE"));
add_opt(common_arg(
{"--expert-heat-decay"}, "F",
"expert heatmap decay rate per update (default: 0.999)",
[](common_params & params, const std::string & value) {
params.expert_heat_decay = std::stof(value);
}
).set_env("LLAMA_ARG_EXPERT_HEAT_DECAY"));
add_opt(common_arg(
{"--expert-heat-log-period"}, "N",
"print the expert heatmap at generation end (default: 0, 0 = off)",
[](common_params & params, int value) {
params.expert_heat_log_period = value;
}
).set_env("LLAMA_ARG_EXPERT_HEAT_LOG_PERIOD"));
add_opt(common_arg(
{"--expert-sync-period"}, "N",
"expert hot store re-sync cadence in tokens (default: 1)",
[](common_params & params, int value) {
params.expert_sync_period = value;
}
).set_env("LLAMA_ARG_EXPERT_SYNC_PERIOD"));
add_opt(common_arg(
{"--expert-hyst"}, "F",
"expert hot store hysteresis ratio (default: 1.3, 0 = off)",
[](common_params & params, const std::string & value) {
params.expert_hyst = std::stof(value);
}
).set_env("LLAMA_ARG_EXPERT_HYST"));
add_opt(common_arg(
{"--expert-dwell"}, "N",
"expert hot store minimum dwell updates before swap (default: 0 = off)",
[](common_params & params, int value) {
params.expert_dwell = value;
}
).set_env("LLAMA_ARG_EXPERT_DWELL"));
add_opt(common_arg(
{"-ehs", "--expert-hot-s"}, "N",
"-1 = autofit slots from free VRAM, 0 = disabled, N = manual top-N slots",
[](common_params & params, int value) {
params.expert_hot_s = value;
llama_expert_preload::set_slots(value);
}
).set_env("LLAMA_ARG_EXPERT_HOT_S"));
add_opt(common_arg(
{"--expert-pin"}, "N",
"fraction (percent) of cold experts to keep pinned in RAM via madvise, "
"0 = off, -1 = auto (hot store sets 40, else 0)",
[](common_params & params, int value) {
params.expert_pin_pct = value;
}
).set_env("LLAMA_ARG_EXPERT_PIN"));
add_opt(common_arg(
{"--expert-no-evict"},
{},
"never evict experts from the hot store (fill-only, no move-back)",
[](common_params &, bool value) {
llama_expert_preload::set_no_evict(value);
}
));
add_opt(common_arg(
{"--expert-move-mode"}, "N",
"expert store mode: 0 = auto, 1 = copy (keep RAM copy), 2 = move "
"(free RAM after verified transfer)",
[](common_params & params, int value) {
params.expert_move_mode = value;
}
).set_env("LLAMA_ARG_EXPERT_MOVE_MODE"));
add_opt(common_arg(
{"--expert-sidecar"},
{},
"load the expert heatmap sidecar (<model>.tier) at start, save it at exit",
[](common_params & params, bool value) {
params.expert_sidecar = value;
}
).set_env("LLAMA_ARG_EXPERT_SIDECAR"));
add_opt(common_arg(
{"--expert-gpu"}, "N",
"put the expert store on this GPU index (default: -1 = all GPUs)",
[](common_params & params, int value) {
params.expert_gpu = value;
}
).set_env("LLAMA_ARG_EXPERT_GPU"));
GGML_ASSERT(params.n_gpu_layers < 0); // string_format would need to be extended for a default >= 0
add_opt(common_arg(
{"-ngl", "--gpu-layers", "--n-gpu-layers"}, "N",
Expand Down
52 changes: 50 additions & 2 deletions common/common.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -1239,12 +1239,48 @@ common_init_result::common_init_result(common_params & params, bool model_only)
if (params.fit_params) {
COM_TRC("%s", "fitting params to device memory ...\n");
COM_TRC("%s", "(for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)\n");
common_fit_params(params.model.path.c_str(), &mparams, &cparams,
int n_expert_hot_s = params.expert_hot_s;
int * p_expert_hot_s = params.expert_hot_s == -1 ? &n_expert_hot_s : nullptr;
const common_params_fit_status fit_status = common_fit_params(params.model.path.c_str(), &mparams, &cparams,
params.tensor_split,
params.tensor_buft_overrides.data(),
params.fit_params_target.data(),
params.fit_params_min_ctx,
params.verbosity >= LOG_LEVEL_DEBUG ? GGML_LOG_LEVEL_DEBUG : GGML_LOG_LEVEL_ERROR);
params.verbosity >= LOG_LEVEL_DEBUG ? GGML_LOG_LEVEL_DEBUG : GGML_LOG_LEVEL_ERROR,
p_expert_hot_s);
if (params.expert_hot_s == -1) {
// -1 = autofit slots from what the fit leaves on GPU; send all experts
// to CPU so the hot store copy reads host pointers (<=> -cmoe).
params.expert_hot_s = n_expert_hot_s > 0 ? n_expert_hot_s : 0;
cparams.expert_hot_s = params.expert_hot_s;
if (params.expert_hot_s > 0) {
for (auto & o : params.tensor_buft_overrides) {
if (o.pattern == nullptr) {
o.buft = ggml_backend_cpu_buffer_type();
o.pattern = LLM_FFN_EXPS_REGEX;
break;
}
}
} else if (fit_status == COMMON_PARAMS_FIT_STATUS_FAILURE) {
LOG_WRN("%s: --expert-hot-s -1 autofit aborted (explicit -ngl/-ncmoe or fit error); expert cache is OFF\n",
__func__);
} else {
LOG_WRN("%s: --expert-hot-s -1 autofit found no free VRAM for expert slots; expert cache is OFF\n",
__func__);
}
}
} else if (params.expert_hot_s == -1) {
// autofit only runs inside --fit; without it -1 is meaningless
params.expert_hot_s = 0;
cparams.expert_hot_s = params.expert_hot_s;
LOG_WRN("%s: --expert-hot-s -1 requires --fit (disabled by -fit off or explicit -ngl/-ncmoe); expert cache is OFF\n",
__func__);
}

// --expert-pin -1 auto: 40 with the hot store, 0 without
if (params.expert_pin_pct == -1) {
params.expert_pin_pct = params.expert_hot_s != 0 ? 40 : 0;
cparams.expert_pin_pct = params.expert_pin_pct;
}

llama_model * model = llama_model_load_from_file(params.model.path.c_str(), mparams);
Expand Down Expand Up @@ -1667,6 +1703,18 @@ struct llama_context_params common_context_params_to_llama(const common_params &
cparams.type_k = params.cache_type_k;
cparams.type_v = params.cache_type_v;

cparams.expert_heat_decay = params.expert_heat_decay;
cparams.expert_heat_log_period = params.expert_heat_log_period;
cparams.expert_hot_s = params.expert_hot_s;
cparams.expert_sync_period = params.expert_sync_period;
cparams.expert_hyst = params.expert_hyst;
cparams.expert_dwell = params.expert_dwell;
cparams.expert_pin_pct = params.expert_pin_pct;
cparams.expert_move_mode = params.expert_move_mode;
cparams.expert_sidecar = params.expert_sidecar;
cparams.expert_gpu = params.expert_gpu;
cparams.model_path = params.model.path.c_str();

return cparams;
}

Expand Down
11 changes: 11 additions & 0 deletions common/common.h
Original file line number Diff line number Diff line change
Expand Up @@ -523,6 +523,17 @@ struct common_params {
int32_t verbosity = 3; // LOG_LEVEL_INFO
int32_t control_vector_layer_start = -1; // layer range for control vector
int32_t control_vector_layer_end = -1; // layer range for control vector

float expert_heat_decay = 0.999f; // multiplicative decay per update
int expert_heat_log_period = 0; // print heatmap at generation end (0 = off)
int expert_hot_s = 0; // top-S expert slots (0 = disabled)
int expert_sync_period = 1; // hot store re-sync cadence in tokens
float expert_hyst = 1.3f; // hysteresis ratio: only swap when cold >= hyst x hot
int expert_dwell = 0; // minimum updates a resident slot must keep before a swap
int expert_pin_pct = -1; // percent of cold experts to keep pinned (madvise); -1 = auto
int expert_move_mode = 0; // expert store mode: 0 = auto, 1 = copy, 2 = move
bool expert_sidecar = false; // load/save the expert heatmap sidecar (<model>.tier)
int expert_gpu = -1; // expert store GPU index (-1 = all GPUs)
bool offline = false;

int32_t ppl_stride = 0; // stride for perplexity calculations. If left at 0, the pre-existing approach will be used.
Expand Down
38 changes: 35 additions & 3 deletions common/fit.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -178,7 +178,7 @@ common_device_memory_data_vec common_get_device_memory_data(
static void common_params_fit_impl(
const char * path_model, struct llama_model_params * mparams, struct llama_context_params * cparams,
float * tensor_split, struct llama_model_tensor_buft_override * tensor_buft_overrides,
size_t * margins_s, uint32_t n_ctx_min, enum ggml_log_level log_level) {
size_t * margins_s, uint32_t n_ctx_min, enum ggml_log_level log_level, int * n_expert_hot_s) {
if (mparams->split_mode == LLAMA_SPLIT_MODE_TENSOR) {
throw common_params_fit_exception("llama_params_fit is not implemented for SPLIT_MODE_TENSOR, abort");
}
Expand Down Expand Up @@ -228,6 +228,8 @@ static void common_params_fit_impl(
int64_t sum_projected_free = 0;
int64_t sum_projected_used = 0;
int64_t sum_projected_model = 0;
int64_t total_moe_bytes = 0; // MoE expert tensor bytes (for slot autofit)
int64_t dense_model_gpu = 0; // dense-only model bytes on GPU (for slot autofit)
std::vector<int64_t> projected_free_per_device;
projected_free_per_device.reserve(nd);

Expand Down Expand Up @@ -541,6 +543,8 @@ static void common_params_fit_impl(
for (size_t id = 0; id < nd; id++) {
global_surplus_cpu_moe += dmds_cpu_moe[id].free;
global_surplus_cpu_moe -= int64_t(dmds_cpu_moe[id].mb.total()) + margins[id];
total_moe_bytes += int64_t(dmds_full[id].mb.model) - int64_t(dmds_cpu_moe[id].mb.model);
dense_model_gpu += int64_t(dmds_cpu_moe[id].mb.model);
}

if (global_surplus_cpu_moe > 0) {
Expand Down Expand Up @@ -641,6 +645,10 @@ static void common_params_fit_impl(
}
if (hp_nex == 0 || global_surplus_cpu_moe <= 0) {
set_ngl_tensor_split_tbo(ngl_per_device, overflow_bufts, *mparams);
if (n_expert_hot_s) {
// all MoE stays on CPU (no surplus), so no GPU hot slots fit
*n_expert_hot_s = 0;
}
return;
}

Expand Down Expand Up @@ -786,6 +794,29 @@ static void common_params_fit_impl(
}

set_ngl_tensor_split_tbo(ngl_per_device, overflow_bufts, *mparams);

// step 5: autofit the expert hot store slots when --expert-hot-s -1 is set.
// the fit above left some MoE bytes on GPU (final_gpu_model - dense_model_gpu);
// s = experts-per-layer that fit, minus one plane for the sentinel slot.
if (n_expert_hot_s && total_moe_bytes > 0) {
const dmds_t dmds_final = common_get_device_memory_data_impl(
path_model, mparams, cparams, devs, hp_ngl, hp_nct, hp_nex, log_level);
int64_t final_gpu_model = 0;
bool is_vulkan = false;
for (size_t id = 0; id < nd; id++) {
final_gpu_model += dmds_final[id].mb.model;
if (dev_names[id].find("Vulkan") != std::string::npos) {
is_vulkan = true;
}
}
// Vulkan reserves per-layer descriptor/pool memory the fit does not
// account for; subtract an 8 MiB per offloaded layer estimate so S
// does not overshoot and OOM at graph capture.
const int64_t vulkan_padding = is_vulkan ? int64_t(hp_ngl) * 8 * MiB : 0;
const int64_t moe_on_gpu = final_gpu_model - dense_model_gpu - vulkan_padding;
const int64_t s = moe_on_gpu > 0 ? int64_t(hp_nex) * moe_on_gpu / total_moe_bytes : 0;
*n_expert_hot_s = s > 1 ? (int) (s - 1) : 0;
}
}

enum common_params_fit_status common_fit_params(
Expand All @@ -796,11 +827,12 @@ enum common_params_fit_status common_fit_params(
llama_model_tensor_buft_override * tensor_buft_overrides,
size_t * margins,
uint32_t n_ctx_min,
ggml_log_level log_level) {
ggml_log_level log_level,
int * n_expert_hot_s) {
const int64_t t0_us = llama_time_us();
common_params_fit_status status = COMMON_PARAMS_FIT_STATUS_SUCCESS;
try {
common_params_fit_impl(path_model, mparams, cparams, tensor_split, tensor_buft_overrides, margins, n_ctx_min, log_level);
common_params_fit_impl(path_model, mparams, cparams, tensor_split, tensor_buft_overrides, margins, n_ctx_min, log_level, n_expert_hot_s);
LOG_TRC("%s: successfully fit params to free device memory\n", __func__);
} catch (const common_params_fit_exception & e) {
LOG_WRN("%s: failed to fit params to free device memory: %s\n", __func__, e.what());
Expand Down
3 changes: 2 additions & 1 deletion common/fit.h
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,8 @@ common_params_fit_status common_fit_params(
llama_model_tensor_buft_override * tensor_buft_overrides, // writable buffer for overrides, needs at least llama_max_tensor_buft_overrides elements
size_t * margins, // margins of memory to leave per device in bytes
uint32_t n_ctx_min, // minimum context size to set when trying to reduce memory use
ggml_log_level log_level); // minimum log level to print during fitting, lower levels go to debug log
ggml_log_level log_level, // minimum log level to print during fitting, lower levels go to debug log
int * n_expert_hot_s = nullptr);// out: fitted expert hot store slots, untouched if not applicable

// print estimated memory to stdout
void common_fit_print(
Expand Down
42 changes: 42 additions & 0 deletions counter.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
# counter

Times the user's input produced a materially better decision than my default.
Update only when the user asks.

## Count: 17

## Examples (2026-08-06 session)

1. **Rotation + cooldown=0 test** - I concluded the corruption was AMD-specific; the user's test ("removing the cooldown should instantly corrupt the rtx?") proved it is Vulkan-wide and duplicate-id driven.

2. **"It's not the model"** - I attributed a failure to model randomness; the user's 200-run knowledge + the CUDA IQ2 test proved it is the tier/Vulkan.

3. **Sentinel + mask must stay** - I claimed copy-on-read eliminates them; the user asked "is the sentinel not still necessary?" and was right (alignment + Vulkan safety).

4. **-no-cnv invalidates corruption tests** - EOS "failures" were ambiguous without conversation mode; I had judged corruption from run counts.

5. **--fit-target 64 was missing** - the correct fit flag changed the measured config.

6. **-ehs -1 autofit** - the valid config revealed the tier is ~61 tok/s (faster than my invalid S=96 numbers).

7. **Native+lazy+madvise instead of the custom pool** - "use llama's native rampool and send a release... load them in vram from disk" replaced two committed pool phases with a simpler, better design.

8. **"RAM allocation is not actually instant"** - caught that the pool's thousands of per-slice mallocs are slow vs one native allocation.

9. **Hash is of the memory bytes, not the output** - "I said generate a hash that can only be generated from the memory" - avoided a float-tolerance mess.

10. **"Try 1024"** - shrinking the hash sample from 16KB to 1KB recovered ~4 tok/s.

11. **Copy-at-init doubles VRAM** - "will we lose the ability to use the 3gb for actual slots?" - caught the transient double-buffer that would OOM an 8GB card.

12. **Uniform first-S startup** - the user chose it over my heatmap-seeded idea.

13. **32+8 memory-fits constraint** - "we crash and oom if the model cant fit in the ram + gpu" - shaped the startup as memory distribution, not just warming.

14. **"Are you sure there is no other way?"** - led to lifting the gate and discovering the real n_tokens>1 blocker was a mask shape assert, not the predicted kernel crash.

15. **Kernel speed loss unacceptable** - pushed to the count+rank kernel fix (v3 reference) over batch-split.

16. **Deferred release** - madvise after verification, less CPU overhead.

17. **Streaming to GPU instead of second disk read** - "move the layers into the gpu in 128mb chunks" - better than my re-read fallback.
4 changes: 2 additions & 2 deletions ggml/include/ggml-rpc.h
Original file line number Diff line number Diff line change
Expand Up @@ -8,10 +8,10 @@ extern "C" {

#define RPC_PROTO_MAJOR_VERSION 5
#define RPC_PROTO_MINOR_VERSION 0
#define RPC_PROTO_PATCH_VERSION 0
#define RPC_PROTO_PATCH_VERSION 2

#ifdef __cplusplus
static_assert(GGML_OP_COUNT == 101, "GGML_OP_COUNT has changed - update RPC_PROTO_PATCH_VERSION");
static_assert(GGML_OP_COUNT == 103, "GGML_OP_COUNT has changed - update RPC_PROTO_PATCH_VERSION");
#endif

#define GGML_RPC_MAX_SERVERS 16
Expand Down
Loading