Add TML Inkling architecture - #25731
danielhanchen wants to merge 27 commits into
Conversation
|
Hi @danielhanchen, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
d5f48c4 to
1cb0374
Compare
50cde5d to
a015409
Compare
Merge upstream PR ggml-org#25731: TML Inkling architecture (+ upstream master sync)
4d8798a to
f47a88c
Compare
Hybrid attention model: 55 sliding-window plus 11 global layers, banded content-dependent relative position bias instead of RoPE, per-layer short convolution state, fine-grained MoE (256 experts top-6 plus 2 shared), attention log-scaling past 128K, 1M context. Includes the GGML_OP_FLASH_ATTN_EXT_BANDED operator (CPU and CUDA, fused into the MMA flash attention kernel with an fp16 accumulator overflow guard), HF to GGUF conversion, chat template with typed content block parsing (interleaved thinking, narration and tool calls), mmproj vision and audio support, and backend op tests at production shapes.
f47a88c to
ce16fff
Compare
|
First: thank you for day-0 support, and it works. We ran inkling UD-Q4_K_XL (587GB) fully resident across two cloud VMs (4x A100 80GB each, 8 total) split over the RPC backend on ordinary Ethernet, and output is coherent and correct (it first-shot passed our hardest structured Terraform generation task). Text only, audio/vision untested here. We also profiled it, and the data below may be useful for the "might need to redesign some parts later" list. Decode runs at roughly 25 percent of what the hardware gives comparable MoEs, and the evidence points at per-token graph rebuild and launch overhead, not compute. Numbers (all this PR at ce16fff, same boxes, sustained single-stream, ctx 8192)
Decode is flat from empty context to 2.2k depth (6.2 to 6.6), so it is not KV/attention scaling. Single box, inkling runs at ~80 percent of the comparable resident MoE. Over RPC it drops to ~27 percent. Where the time goesPer token at 6.45 t/s = 155 ms. During decode:
perf record on llama-server during decode gives a maximally flat profile: 1,783 distinct symbols, top entry 3.5 percent. Named entries in the top 20: cudaStreamSynchronize / ggml_backend_cuda_synchronize, pthread_mutex_lock under cuLaunchKernel, ggml_backend_sched_split_graph, ggml_gallocr_alloc_graph, and ggml-rpc add_tensor + get_alloc_size (the whole graph re-serialized to the worker every token, which is the 5.5 MB). So the shape looks like: a ~9k-node graph rebuilt, re-split, re-allocated, re-serialized and launched node by node every token while the GPUs idle. Our guess is the banded rel-bias / shortconv path defeats the llama-side graph reuse, and the node count does the rest. If graph reuse engages for this arch, the RPC number should improve several-fold on its own. Happy to re-run any branch of this PR on the 2-node A100 cluster or the single box and report the same measurements. Full logs, perf.data and exact commands available on request. |
# Conflicts: # ggml/src/ggml-cuda/mmq.cuh # src/llama-model-saver.cpp # src/llama-vocab.h
…#39) Both pins pointed at commits that no longer merge onto the selected upstream base, so the nightly full release build failed while resolving the PR set. ggml-org#24523 was pinned to 66f43aa6, which was not even the PR head at the time (the resolver logged that the head had moved to 0b78558a). That commit conflicts in common/chat.cpp against both b10107 and b10133. The PR has since been rebased onto current master, so the pin now points at its new head a58a7fa6, which merges cleanly. ggml-org#25731 was pinned to ce16fff, which conflicts on its own in ggml/src/ggml-cuda/mmq.cuh, src/llama-model-saver.cpp and src/llama-vocab.h. This never surfaced in the log because the resolver stops at the first failing entry. Its current head 453c438 merges cleanly by itself. Note that ggml-org#24523 and ggml-org#25731 still conflict with each other in common/chat.cpp and src/llama-arch.h: both append a new llm_arch enum value immediately before LLM_ARCH_UNKNOWN and both add a chat parser in the same region. Reordering does not help, since whichever entry is applied second hits the same conflict. Resolving that needs a decision about which of the two to carry, so it is left alone here. Requires a base of b10133 or newer: b10107 predates the rename of common_chat_params::thinking_end_tag to thinking_end_tags, which the rebased ggml-org#24523 depends on. Co-authored-by: Daniel Han <unslothai@gmail.com>
# Conflicts: # ggml/include/ggml-rpc.h # ggml/src/ggml-backend-meta.cpp # ggml/src/ggml-cuda/ggml-cuda.cu # src/llama-model-saver.cpp # tests/test-llama-archs.cpp
Prebuilt: repin ggml-org#24423 and ggml-org#25731 onto carry branches
# Conflicts: # src/llama-model.cpp
# Conflicts: # ggml/src/ggml-rpc/ggml-rpc.cpp
Resolves the two conflicts left by upstream ggml-org#27764 (chat parsers split into common/parsers) and ggml-org#26675 (ggml_prec rework): - common_chat_params_init_inkling moves to common/parsers/inkling.cpp, registered in parsers.h and sources.cmake; chat.cpp keeps only the template detection branch. - GGML_PREC_F32_PEDANTIC joins the new ranked ggml_prec scale at 5, below GGML_PREC_F32, since it is the stricter contract; ggml_flash_attn_ext_banded keeps its declaration next to the flash attention API.
…C_F32_PEDANTIC Every caller of the deprecated ggml_mul_mat_set_prec / ggml_flash_attn_ext_set_prec now uses ggml_prec_set_acc (the macOS prebuilt leg builds with LLAMA_FATAL_WARNINGS, so the deprecation would be fatal there), and ggml_prec_set_acc accepts GGML_OP_FLASH_ATTN_EXT_BANDED. With the enum ranked, the CUDA cuBLAS compute-type gate and the CPU, spacemit and Vulkan flash-attention accumulator checks compare by rank instead of equality with GGML_PREC_F32, so a pedantic request never falls through to a lower-precision path. The CUDA mul_mat_id slice carries the acc and src precision slots, and ggml_cuda_mul_mat_id_needs_sync applies the same pedantic gate as ggml_cuda_mul_mat_id so a pedantic F32 expert matmul cannot trip its assertion.
…master (#209) The nightly on b10865 (run 34287079907) stopped in resolve: the pinned commit no longer merged after upstream ggml-org#27764 moved the chat parsers into common/parsers and ggml-org#26675 reworked ggml_prec. The PR branch now carries a merge of upstream master with both conflicts resolved, so the mix merges again on b10865 and b10870 with only the usual additive add/add merges.
|
Heads up, #29042 makes the saver write the SWA pattern, so once it lands this architecture no longer needs to be excluded from llama_model_saver_supports_arch and can get the test-llama-archs roundtrip. |
|
Just a heads-up: I am currently working through my backlog where this has showed up but I categorically refuse to review the CUDA changes like. As laid out in the contributing guidelines, the initial PR should only contain the baseline CPU implementation and the equivalent CUDA changes should be in follow-up PRs. At most I am willing to provide general advice regarding how the implementation should look like. |
Adds support for the Inkling architecture, a Python safetensors-to-GGUF converter, the graph build, and the kernel changes needed for correct and deterministic inference.
int64_ton some ops since large MoEs would go out of indexI tried to keep changes unbreaking - might need to redesign some parts later.
Used AI for kernels, but hand verified and checked everything carefully.