feat: device_fit governor calc + resident-override wire + K3 fork bump - #2109
Conversation
…nto canary Advances the vendored llama.cpp fork 30 commits (clean FF over canary's stale 66594cc3f): container-serve resident-override (LLAMA_RESIDENT_OVERRIDE), the rung-2 ResidencyCache plan-file consumer, the score-hint/generation-bias actuator, PagerCaptureEvent emit, fit-device --reserve-gb. Makes canary USE the K3 misfit-serving stack (measured 0.33 tok/s WASTE-parity on a 32GB card). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…M5's replay
Live GGML_MOE_TRACE_FILE slice (12B records: u64 tkey + u32 e) + the reverse
table so BanditPlanController recovers (layer,expert): tkey=FNV-1a of
blk.{layer}.ffn_{gate,up,down}_exps.weight, e=within-layer expert idx, expert
identity=(layer,e) deduped across the 3 matrices. RUN-1 static-pin datum input.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
The actual std-only Rust prototypes written against live K3 traces this session: trace_replay (recency beats LFU 3-4x), predictor (offline learned-decay +5pts held-out), online_predictor (bandit 49.8 vs 47.8 best-fixed on non-stationary), self_optimize (joint speed×quality). These are the faithful-port source for the learned policy behind TierPolicy (continuum-core expert_tier_policy.rs, #276). Numbers are properties of these exact constants + reward math — reproduce before improving. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
The major GPU speedup: promote hot experts to persistent VRAM so decode's hot path is GPU-native (zero fetch, zero copy). 3 increments (copy-skip -> VRAM hot cache -> pipeline), the 32GB rate-distortion constraint (imatrix-enabled resident shrink frees VRAM for the hot set), measured per-increment via k3-bench. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…vs 2 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…DeviceUploadFetcher Mechanism is the existing (buft,fetcher)-generic ResidencyCache; a VRAM cache = same class + device buft + host->device fetcher. 3 small parameterized pieces (DeviceUploadFetcher, instantiate w/ GGML_MOE_VRAM_CACHE_GB, seam hook). Stats only tune params -> zero mechanism rework. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…rnor seam) Diagnoses the hardcoded-cache overcommit that collapsed K3 fetch bandwidth (40GB pinned + mmap = 95.9GB on 63GB -> pagefile thrash -> 205 MB/s -> 0.027 tok/s) and lays out the clean architecture: governor owns the residency budget net of the model's mmap footprint, plan-file is the one wire, ResidencyCache is pure mechanism. Governor-interface sections marked [M5 OWNS] for her to edit. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…f-mmap is explicit arithmetic; plan_file.budget_bytes is the lease wire) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo
…path Records the BigMama measurements feeding M5's #287 derivation (non-cache ~56GB, per-token working set 5.5GB, governed budget ~6GB, fetch recovers to 2.5GB/s at fit), the now-complete C++ cache mechanism (enable-from-plan, grow, shrink), and the three-piece graduated path to serving/load kimi-k3 (catalog row + serving-lane MoE launch + #287) replacing the rigged .bat. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
start-server.sh now provisions the full Windows CUDA build env before the cargo builds (the cargo/nvcc path had none, unlike the vcvars-wrapped llama cmake): import MSVC via vswhere->VS2022-14.4x + a .bat env dump (cl.exe for nvcc), pin CMAKE to the manifest install, force CMAKE_GENERATOR=Ninja (the VS18-2026 auto- pick is undefined in cmake 3.30), add the Windows SDK bin (mt.exe/rc.exe), select a complete CUDA toolkit + CUDA_PATH (a provisioning split left cuda-env with 0 import libs vs cuda-13.2's 12), and RUSTFLAGS -L for pocket-tts (which emits no link-search) while re-carrying +crt-static so the /MT GPU stack still links. Portability: expert_container.rs + commands/capacity.rs used Unix-only std::os::unix::fs::FileExt::read_exact_at. Add crate::platform_io::pread_exact (unix read_exact_at / windows seek_read loop) - one place for positioned reads. Build validated (npm start exit 0, continuum-core lib clean). A separate runtime hot-loop on the #2088 core at startup is tracked apart from this build fix. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
Pure calc the governor uses to fit a streaming-MoE's RESIDENT (non-expert) tier to a device VRAM budget and reconcile it with the expert tier on ONE budget — fixing the double-count where the expert pager was handed the full VRAM ceiling while resident silently ate most of it. Partition (in order): compute reserve -> resident (Native | device-fit Override | Unfittable) -> sufficient-context KV -> everything left = hot-expert VRAM budget (maximized: more on-GPU experts, fewer streams). Context is derived + clamped, never hand-picked. Artifact resolver injected (no hardcoded paths). Standalone-validated 7/7; M5 wires it into the daemon spawn path + launch (ServingTarget.resident_override) per the K3 sprint split. Refs #29 #31 #36. Arch-confirmed on real K3 UD-IQ2 (93 blk/896 exp/top-16). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
Wire foundation for the governor's device_fit plan: ServingTarget carries resident_override: Option<PathBuf>, and the launcher exports it as LLAMA_RESIDENT_OVERRIDE so llama.cpp sources the precision-shrunk RESIDENT (non-expert) tensors from the device-fit GGUF (all offloaded to GPU) while the primary streams experts. All builders updated; defaults None (resident serves as-shipped, no behavior change) until compute_resident_override + the resolve-or-generate resolver (#35) land next. In-crate validated. Refs #29 #36. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
M5's per-layer KV accessors (n_head_kv_il + n_embd_head_{k,v}_il, continuum #238)
+ graph reconciliation. The K3 engine now builds against these — enables the
device_fit resident-override serve + honest per-layer K3 KV sizing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
|
M5 review (formal approve blocked — same GitHub identity): APPROVED. Merge when your compute step is ready, or merge this foundation now — it's inert ( What's right, notably: Two notes for the
🤖 Generated with Claude Code |
# Conflicts: # tools/scripts/start-server.sh
The governor now DECIDES the resident source per serve: compute_resident_override derives resident_bytes (weights - expert_bytes_total) vs the governed VRAM ceiling via capacity::device_fit, and sets ServingTarget.resident_override. A dense/small model fits native (None); a >VRAM-resident MoE (K3) resolves a cached device-fit override that fits, else Unfittable → route to grid / generate (#35), glass-boxed. resolve_device_fit_override (model_registry::artifacts): looks up a per-user device-fit cache convention (<storage_root>/device-fit/<id>/) + a resident-bytes sidecar; returns the override only when its resident fits the usable budget. No hardcoded paths; generation/HF discovery is #35. The resident-fit decision turns only on resident_bytes vs budget — per-layer KV (#2107 ModelCapabilities) drives the context/expert split elsewhere, so KV is not consulted here. Refs #29 #35 #36. In-crate validated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
|
Review of ec1b541 (compute_resident_override) — approve, with two binding follow-ups before the K3 serve consumes downstream outputs: ✅ The resident-fit verdict is correct and honest: Native/Override/Unfittable turns only on resident_bytes vs (budget − reserve), trust-but-verify rejects oversized overrides, Unfittable is loud.
Neither blocks the LiveUploadPager run — go. 🎯 |
…naged like VRAM/RAM Joel: 'like vram and memory, this contention has to be managed between cold storage and nvme.' Design: NVMe is a governed HOT-SERVING tier (a ResourcePool, same TrackedDir + evict_at_least machinery as CargoTargetPool), whose eviction = MIGRATE frozen/duplicate artifacts to the Cold drive, not a manual rm. Serving asks ensure_hot_resident(model); composes with device_fit's Unfittable one tier down (VRAM). Corrects the DriveRole bug: Cold (HDD) is FROZEN storage, never the per-token streaming tier (HDD = unservable). Dissolves today's K3 container disk fight: the C: IQ2 is a verified duplicate of the D: copy -> governor migrates it off NVMe -> container fits, no human deletes anything. Refs #12 #36. Design for M5's system_resources lane. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
… for the storage tier The gate the NVMe serving-tier eviction (#302) consults before dropping a frozen GGUF: is an IDENTICAL twin already on cold storage? is_structural_twin (pure) = same shard count + per-shard name + size, zero-byte shards never match. scan_shards + find_cold_twin are the thin fs layer. Never drop an NVMe artifact without a VERIFIED cold twin (dropping 662GB on a path guess is the failure this guards). Standalone-validated 5/5. Composes with device_fit + M5's NvmeServingTierPool. Refs #12 #36. Design: STORAGE-SERVING-TIER-GOVERNOR.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…rve wired Both halves of the DirContainerFetcher wire (BigMama fetcher + moe_pick_fetcher branch 175ac9d6a; M5 caller-side encode + record_bytes reader c6469d5). Serving now reads the aligned per-layer container (GGML_MOE_CONTAINER) instead of the scattered raw GGUF — the honest ~2.6GB/s path. Retires the built-not-wired ContainerFetcher. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
# Conflicts: # core/continuum-core/src/modules/serving_daemon.rs
…coverage measurement Merge fix: vision_sidecar's ServingTarget was missing resident_override (added by #29). Plus a measurement test that drains the real K3 routed-access fixture through the #282 predictive instrument and prints repeat_recall / predicted_delta / schedulable_coverage — the go/no-go for the LiveUploadPager predictive pipeline (H2D/token = (1 - coverage) x ~11GB). Prints, never asserts (real routing sample). NOTE: can't run on windows-msvc (pre-existing cargo-test Unix-socket block, ipc/mod.rs); runs on M5's Mac. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…, --once, --budget-slots) The moe-pager-driver gains an offline replay mode so any completed GGML_MOE_TRACE_FILE can be scored on any box (windows-msvc clean by crate constraint), not just tailed live next to a serve: - --synth-layers N: synthesize the tkey->layer map from layer count alone (TkeyTable::for_layers, the same zero-config seam MoeTraceTail owns) instead of requiring an operator tkey-to-layer-matrix.json. - --once: exit when the trace stops growing (EOF) and print a SUMMARY line with mean DECODE-token serving hit = warm schedulable coverage. - --budget-slots N: override the predictor residency budget (default auto = first token x1.5) to measure the coverage-vs-free-VRAM curve (the device-fit tradeoff). Measured on BigMama run2.trace (302 warm decode tokens): bandit coverage 13.8% @250 slots -> 51.3% @2000 -> 65.7% @4024, beating naive last-N recency by +7-9pts at matched VRAM. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…ver for K3 VDD offline measurement (cooccur-ceiling bin) on the real warm serve trace (run2.trace, 122 held-out decode tokens, 11102 layer-steps): cross_layer_cooccur_hit 0.159 (adjacent-layer noisy-OR) recency_same_layer_hit 0.403 (last token, same layer) structure beyond recency -0.244 cooccur_recall_on_recency_misses 0.112 (11824/106058) Adjacent-layer co-occurrence predicts <half what plain recency does, and recovers only 11% of the experts recency misses (~base rate). K3 expert routing has no exploitable cross-layer structure — the CrossLayerExpert- Predictor prefetch lever is not worth wiring (saves the ggml pass-id capture slice). Recency-family residency (the bandit EMA curve) is THE signal; the only lever that lifts K3 is freeing VRAM (device-fit shrink) so residency coverage can reach the measured 51%. Caveat: adjacent-layer, one workload trace. Wider-predecessor noisy-OR would regress toward the frequency baseline (which underperforms recency), so a large lift is unlikely — but not measured here. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…ache half (#23) Pins the fork at the DeviceUploadFetcher wiring (my half of the LiveUpload- Pager H2D-kill). Off unless GGML_MOE_VRAM_CACHE_GB / plan device_budget_bytes enables it; host serving path byte-for-byte unchanged. M5's expert-loop D2D half lands next on the same seam. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
… + tier manifest (#40) Enables the device-fit division: produce a small resident override per precision tier + a (tier_label, resident_bytes) sidecar the governor reads to co-optimize the VRAM split. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
#3) The second control rung above the pager's DecayBandit. The pager decides WHICH experts stay resident (reward=hit-rate, cheap, online). This decides HOW TO DIVIDE the card — resident (non-expert) weights vs expert cache — to MAXIMIZE tok/s. That reward (actual tok/s) is EXPENSIVE (a serve), so naive online RL flails; the fix is SIM-WARM-START: predict tok/s per division OFFLINE from the measured coverage curve, then a slow bandit refines each arm from real measured tok/s. - CoverageModel: piecewise-linear coverage(slots) over MEASURED points (k3_measured() = the trace-replay curve); saturates, never extrapolates up. - predict_tok_s: coverage -> (1-coverage)*experts/token*expert_bytes H2D -> t_token -> tok/s. Higher coverage -> less H2D -> faster (the load-bearing property, tested). - feasible_divisions: tier catalog (from --resident-only manifests) x HardwareBudget -> cache budget/slots per tier; drops VRAM-overflow tiers. - DivisionBandit: warm_start from the predictor; observe(tier, measured_tok_s) overrides the prior on first serve then EMAs — the expensive reward spent only on the arm actually run. Policy lives here (windows-clean, 4 tests pass); serving_daemon actuates it (M5's #2: discover manifests, feed catalog+budget+live tok/s, apply the chosen {resident_tier, device_budget_bytes} to the plan). Fractal control law: pager (experts<->hit-rate) -> this (VRAM split<->tok/s) -> grid. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…e (clamp to free VRAM + null-buffer guard, #23) Testing convicted the segfault as VRAM oversubscription (K3 33GB resident + env cache on a 32GB card, cudaMalloc lazy-VMM deferred fault). Fix: clamp device budget to measured free VRAM (mine) + M5's D2D null-buffer guard. Device cache now disables safely where there's no room (K3) and works where there is (V4-Flash); can't crash from any budget source. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…lash curve Feeds the real BigMama RTX 5090 --n-cpu-moe sweep (DeepSeek-V4-Flash UD-IQ2_M) into DivisionBandit: 0 resident=1.39, 8 resident=1.69, 14 resident=1.68 tok/s. Asserts the bandit converges on the SATURATION KNEE (8 layers), not max residency — 8->14 layers buys nothing at +11GB VRAM. Encodes the measured finding that the governor must learn 'minimal static residency + max device cache', the freed VRAM belonging to the recency cache (#43), not to over-pinned static layers. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
Measured V4-Flash device-cache coverage curve (5090, GGML_MOE_VRAM_CACHE_GB sweep): 6GB/992slots=1.80, 12GB/1985=3.10, 22GB/3630=2.96 tok/s, all 100% hit. tok/s is NON-MONOTONIC in budget: undersized churns, 12GB is the plateau knee, 22GB is no better (100% hit but O(slots) reserve_slot eviction scan). predict_tok_s's monotonic prior would pick 22GB; only the measured reward lands on 12GB — which frees ~20GB of a 32GB card for co-resident lanes. Pins the invariant that the governor must not oversize the cache and starve other models. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…sync restore + enum fix Includes the prefetch host_visible guard (THE #43 crash fix, validated 3.05 tok/s V4-Flash device cache on the 5090), M5's async cpy_tensor_async restore, and the moe-pack quant-enum fix. A fresh parent build now includes the un-crashable device cache instead of the pre-fix pin. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
# Conflicts: # core/continuum-core/src/capacity/expert_container.rs # core/continuum-core/src/commands/capacity.rs
|
Merged via admin override. The one red check was a KNOWN-FLAKY timing test — |
The current windows-cuda-serving lane (supersedes the stale PR #2056, which predates the canary K3-expert modules).
Composes with M5's #2107 (per-layer KV) + #2108 (uniform_offload_required GDN guard). Arch-confirmed on real K3.
🤖 Generated with Claude Code