Skip to content

Add TML Inkling architecture - #25731

Draft
danielhanchen wants to merge 27 commits into
ggml-org:masterfrom
danielhanchen:add-inkling
Draft

danielhanchen wants to merge 27 commits into
ggml-org:masterfrom
danielhanchen:add-inkling

Conversation

@danielhanchen

@danielhanchen danielhanchen commented Jul 15, 2026 •

Copy link
Copy Markdown
Contributor

Adds support for the Inkling architecture, a Python safetensors-to-GGUF converter, the graph build, and the kernel changes needed for correct and deterministic inference.

  1. Had to use int64_t on some ops since large MoEs would go out of index
  2. Made a Flash Attention banded attention kernel
  3. Checked logits, outputs and long context + needle in a haystack and others
  4. And more - will enumerate later
  5. Has audio + vision support
  6. Testing: Runs GGUFs made at https://huggingface.co/unsloth/inkling-GGUF multimodal well

I tried to keep changes unbreaking - might need to redesign some parts later.

Used AI for kernels, but hand verified and checked everything carefully.

@github-actions github-actions Bot added model Model specific testing Everything test related ggml changes relating to the ggml tensor library for machine learning mtmd Related to multimodal functionality (video/image/audio) CUDA Related to the CUDA backend conversion labels Jul 15, 2026
oobabooga added a commit to unslothai/llama.cpp that referenced this pull request Jul 15, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Jul 15, 2026

Copy link
Copy Markdown

Hi @danielhanchen, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 3 open PRs.

  • Multiple backend changes in one PR: When adding support for a new model or feature, focus on CPU support only in the initial PR. Add support for other backends like CUDA in follow-up PRs. If you have a good reason to modify multiple backends in one PR, please explain it.

  • Large PR: Large changes require prior discussion (e.g. an issue or RFC) and maintainers may not be able to review this PR as-is. Consider splitting it into smaller, focused PRs.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

oobabooga added a commit to unslothai/llama.cpp that referenced this pull request Jul 15, 2026
@danielhanchen
danielhanchen force-pushed the add-inkling branch 3 times, most recently from d5f48c4 to 1cb0374 Compare July 15, 2026 19:31
Comment thread models/templates/Inkling.jinja Outdated
@danielhanchen
danielhanchen force-pushed the add-inkling branch 4 times, most recently from 50cde5d to a015409 Compare July 16, 2026 11:00
Vect0rM added a commit to AtomicBot-ai/atomic-llama-cpp-turboquant that referenced this pull request Jul 17, 2026
Merge upstream PR ggml-org#25731: TML Inkling architecture (+ upstream master sync)
@danielhanchen
danielhanchen force-pushed the add-inkling branch 2 times, most recently from 4d8798a to f47a88c Compare July 18, 2026 06:26
Hybrid attention model: 55 sliding-window plus 11 global layers, banded
content-dependent relative position bias instead of RoPE, per-layer short
convolution state, fine-grained MoE (256 experts top-6 plus 2 shared),
attention log-scaling past 128K, 1M context.

Includes the GGML_OP_FLASH_ATTN_EXT_BANDED operator (CPU and CUDA, fused
into the MMA flash attention kernel with an fp16 accumulator overflow
guard), HF to GGUF conversion, chat template with typed content block
parsing (interleaved thinking, narration and tool calls), mmproj vision
and audio support, and backend op tests at production shapes.
@jaholmesuk

jaholmesuk commented Jul 19, 2026 •

Copy link
Copy Markdown

First: thank you for day-0 support, and it works. We ran inkling UD-Q4_K_XL (587GB) fully resident across two cloud VMs (4x A100 80GB each, 8 total) split over the RPC backend on ordinary Ethernet, and output is coherent and correct (it first-shot passed our hardest structured Terraform generation task). Text only, audio/vision untested here.

We also profiled it, and the data below may be useful for the "might need to redesign some parts later" list. Decode runs at roughly 25 percent of what the hardware gives comparable MoEs, and the evidence points at per-token graph rebuild and launch overhead, not compute.

Numbers (all this PR at ce16fff, same boxes, sustained single-stream, ctx 8192)

Config Decode t/s Note
Qwen3-235B UD-Q4, 4x A100 single box 53.2 matches master (53.7), so this PR does not regress existing archs
Qwen3-235B UD-Q4, 8x A100 over RPC 44.5 matches master (45.9), so RPC is healthy
GLM-5.2-744B UD-Q4, 8x A100 over RPC (master, same boxes, previous day) 24.0 the expected class for a ~40B-active MoE at Q4
inkling UD-Q4_K_XL, 8x A100 over RPC, fully resident 6.45 41B active, expected ~GLM-class
inkling UD-Q2_K_XL, 4x A100 single box, fully resident 20.65 isolates RPC from the equation

Decode is flat from empty context to 2.2k depth (6.2 to 6.6), so it is not KV/attention scaling. Single box, inkling runs at ~80 percent of the comparable resident MoE. Over RPC it drops to ~27 percent.

Where the time goes

Per token at 6.45 t/s = 155 ms. During decode:

  • All 8 GPUs sit at 0 to 6 percent utilization.
  • llama-server burns ~35 percent of one host core.
  • The host pushes ~5.5 MB per token to the RPC worker (36 MB/s at 6.5 t/s). Qwen-235B on the same pair moves ~0.75 MB per token.
  • The server log reports graphs reused = 0 on every request, on both topologies.
  • sched_reserve reports graph nodes = 8985, graph splits = 23, plus a 35 MiB CPU compute buffer (a few ops fall to CPU).
  • GGML_CUDA_DISABLE_GRAPHS=1 changes nothing (20.74 vs 20.65 single box), so CUDA graph capture is not engaging either way and every node pays a discrete launch.

perf record on llama-server during decode gives a maximally flat profile: 1,783 distinct symbols, top entry 3.5 percent. Named entries in the top 20: cudaStreamSynchronize / ggml_backend_cuda_synchronize, pthread_mutex_lock under cuLaunchKernel, ggml_backend_sched_split_graph, ggml_gallocr_alloc_graph, and ggml-rpc add_tensor + get_alloc_size (the whole graph re-serialized to the worker every token, which is the 5.5 MB).

So the shape looks like: a ~9k-node graph rebuilt, re-split, re-allocated, re-serialized and launched node by node every token while the GPUs idle. Our guess is the banded rel-bias / shortconv path defeats the llama-side graph reuse, and the node count does the rest. If graph reuse engages for this arch, the RPC number should improve several-fold on its own.

Happy to re-run any branch of this PR on the 2-node A100 cluster or the single box and report the same measurements. Full logs, perf.data and exact commands available on request.

oobabooga pushed a commit to oobabooga/llama.cpp that referenced this pull request Jul 21, 2026
The add-inkling branch was force-pushed, so the old pin
a015409 is no longer a commit of the PR and the nightly Resolve
tag step refuses it. Repin to the current head ce16fff.
# Conflicts:
#	ggml/src/ggml-cuda/mmq.cuh
#	src/llama-model-saver.cpp
#	src/llama-vocab.h
danielhanchen added a commit to unslothai/llama.cpp that referenced this pull request Jul 26, 2026
…#39)

Both pins pointed at commits that no longer merge onto the selected
upstream base, so the nightly full release build failed while resolving
the PR set.

ggml-org#24523 was pinned to 66f43aa6, which was not even the
PR head at the time (the resolver logged that the head had moved to
0b78558a). That commit conflicts in common/chat.cpp against both b10107
and b10133. The PR has since been rebased onto current master, so the
pin now points at its new head a58a7fa6, which merges cleanly.

ggml-org#25731 was pinned to ce16fff, which conflicts on its
own in ggml/src/ggml-cuda/mmq.cuh, src/llama-model-saver.cpp and
src/llama-vocab.h. This never surfaced in the log because the resolver
stops at the first failing entry. Its current head 453c438 merges
cleanly by itself.

Note that ggml-org#24523 and ggml-org#25731 still conflict with each other in
common/chat.cpp and src/llama-arch.h: both append a new llm_arch enum
value immediately before LLM_ARCH_UNKNOWN and both add a chat parser in
the same region. Reordering does not help, since whichever entry is
applied second hits the same conflict. Resolving that needs a decision
about which of the two to carry, so it is left alone here.

Requires a base of b10133 or newer: b10107 predates the rename of
common_chat_params::thinking_end_tag to thinking_end_tags, which the
rebased ggml-org#24523 depends on.

Co-authored-by: Daniel Han <unslothai@gmail.com>
@lee-b lee-b mentioned this pull request Aug 19, 2026
4 tasks done
# Conflicts:
#	ggml/include/ggml-rpc.h
#	ggml/src/ggml-backend-meta.cpp
#	ggml/src/ggml-cuda/ggml-cuda.cu
#	src/llama-model-saver.cpp
#	tests/test-llama-archs.cpp
danielhanchen added a commit to unslothai/llama.cpp that referenced this pull request Aug 26, 2026
Resolves the two conflicts left by upstream ggml-org#27764 (chat parsers split into
common/parsers) and ggml-org#26675 (ggml_prec rework):

- common_chat_params_init_inkling moves to common/parsers/inkling.cpp,
  registered in parsers.h and sources.cmake; chat.cpp keeps only the
  template detection branch.
- GGML_PREC_F32_PEDANTIC joins the new ranked ggml_prec scale at 5, below
  GGML_PREC_F32, since it is the stricter contract; ggml_flash_attn_ext_banded
  keeps its declaration next to the flash attention API.
…C_F32_PEDANTIC

Every caller of the deprecated ggml_mul_mat_set_prec /
ggml_flash_attn_ext_set_prec now uses ggml_prec_set_acc (the macOS prebuilt
leg builds with LLAMA_FATAL_WARNINGS, so the deprecation would be fatal
there), and ggml_prec_set_acc accepts GGML_OP_FLASH_ATTN_EXT_BANDED.

With the enum ranked, the CUDA cuBLAS compute-type gate and the CPU, spacemit
and Vulkan flash-attention accumulator checks compare by rank instead of
equality with GGML_PREC_F32, so a pedantic request never falls through to a
lower-precision path. The CUDA mul_mat_id slice carries the acc and src
precision slots, and ggml_cuda_mul_mat_id_needs_sync applies the same
pedantic gate as ggml_cuda_mul_mat_id so a pedantic F32 expert matmul cannot
trip its assertion.
@github-actions github-actions Bot added the Vulkan Issues specific to the Vulkan backend label Sep 9, 2026
danielhanchen added a commit to unslothai/llama.cpp that referenced this pull request Sep 9, 2026
…master (#209)

The nightly on b10865 (run 34287079907) stopped in resolve: the pinned
commit no longer merged after upstream ggml-org#27764 moved the chat parsers into
common/parsers and ggml-org#26675 reworked ggml_prec. The PR branch now carries a
merge of upstream master with both conflicts resolved, so the mix merges
again on b10865 and b10870 with only the usual additive add/add merges.
iodeh referenced this pull request in iodeh/unsloth-llama.cpp Sep 14, 2026
@ServeurpersoCom

Copy link
Copy Markdown
Contributor

Heads up, #29042 makes the saver write the SWA pattern, so once it lands this architecture no longer needs to be excluded from llama_model_saver_supports_arch and can get the test-llama-archs roundtrip.

@JohannesGaessler

Copy link
Copy Markdown
Contributor

Just a heads-up: I am currently working through my backlog where this has showed up but I categorically refuse to review the CUDA changes like. As laid out in the contributing guidelines, the initial PR should only contain the baseline CPU implementation and the equivalent CUDA changes should be in follow-up PRs. At most I am willing to provide general advice regarding how the implementation should look like.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning model Model specific mtmd Related to multimodal functionality (video/image/audio) testing Everything test related Vulkan Issues specific to the Vulkan backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.