From 1c3c3d4dc1846def941ff498ef485fbf78029b3d Mon Sep 17 00:00:00 2001 From: gthulhu-work Date: Mon, 7 Sep 2026 12:24:28 +0800 Subject: [PATCH 1/5] docs: add node-level scheduling policy guide --- docs/node-policies.md | 211 ++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 211 insertions(+) create mode 100644 docs/node-policies.md diff --git a/docs/node-policies.md b/docs/node-policies.md new file mode 100644 index 0000000..f9827fe --- /dev/null +++ b/docs/node-policies.md @@ -0,0 +1,211 @@ +# Node-Level Scheduling Policies + +Gthulhu can apply scheduling policies directly to Linux tasks on a selected node, without requiring the workload to be represented by a Kubernetes Pod. + +This is useful for workloads such as: + +- standalone or Docker-based vLLM servers; +- systemd services and host daemons; +- local inference runtimes on DGX Spark / GB10 systems; +- other processes that need workload-aware CPU scheduling but are not managed by Kubernetes. + +A node-level policy selects one or more nodes, matches Linux tasks by their `comm` name, and assigns a Gthulhu scheduling strategy to the matched tasks. + +!!! tip "Try the live mock" + The Web GUI mock includes a Node Policies page at [gthulhu.github.io/Gthulhu/#/node-policies](https://gthulhu.github.io/Gthulhu/#/node-policies). + +## Node Policies vs Kubernetes Scheduling Policies + +The two policy paths solve different selection problems: + +| Policy type | Selects workload by | Best fit | +| --- | --- | --- | +| Kubernetes scheduling policy | Namespace, labels, Pod/container identity | Workloads managed by Kubernetes | +| Node-level scheduling policy | Node identity plus Linux task name | Host processes, Docker workloads, standalone vLLM, system services | + +Both paths eventually resolve workload intent into node-local Linux task scheduling. The main difference is how the target task is discovered. + +## Prerequisites + +To change CPU scheduling behavior, the target node must run the Gthulhu scheduler path on Linux 6.12+ with `sched_ext` support enabled. + +The Manager and the node-local Decision Maker must also be connected so the Manager can generate a node scheduling intent and deliver it to the selected node. + +See [Installation](installation.md) and [Loading sched_ext schedulers](scx-loader.md) for scheduler setup. + +## Creating a Node Policy from the Web GUI + +Open **Node Policies** from the Web GUI and click **New Node Policy**. + +The form contains the following fields: + +- **Node Names**: exact node names, separated by commas. For a standalone inference host, this is usually the simplest selector. +- **Node Label Selectors**: select nodes by labels when multiple registered nodes share the same role or hardware class. +- **DRA Selectors**: optional device-based selection for Kubernetes/DRA-aware environments. Standalone vLLM users can normally leave this empty. +- **Command Regex**: regular expression matched against the Linux task name (`comm`). +- **Priority**: controls whether the matched task receives Gthulhu's boosting behavior. +- **Execution Time**: custom scheduling time slice, in nanoseconds. + +After the policy is saved, the Manager generates one or more **Node Scheduling Intents** for the nodes selected by the policy. + +## Task Matching Is TID-Aware + +Linux schedules tasks/threads, not only process leaders. Gthulhu therefore scans thread entries under: + +```text +/proc//task//comm +``` + +and applies `Command Regex` to each thread name independently. + +For example, a vLLM process may look like this: + +```text +/proc/3785998/comm = python3.12 +/proc/3785998/task/3786004/comm = EngineCore_DP0 +/proc/3785998/task/3786005/comm = EngineCore_DP1 +``` + +A policy using: + +```regex +^EngineCore(_DP[0-9]+)?$ +``` + +can therefore target the latency-sensitive vLLM worker threads even when the process leader itself is only named `python3.12`. + +The resolved strategy is keyed by the matched **TID**. In user-space scheduling, an exact TID strategy is checked first and a TGID strategy is used as a fallback. + +!!! note + Match the value in `/proc/.../comm`, not the full command line from `ps` or `/proc/.../cmdline`. + +## Priority and Time-Slice Semantics + +Current Gthulhu semantics are: + +- `Priority > 0`: the strategy is **boosting**. The matching task receives priority treatment in the user-space scheduler. +- `Priority == 0`: the strategy is **non-boosting**. In user-space mode, it can still apply a custom time slice without jumping the run queue. +- `Execution Time`: the custom time slice in nanoseconds, where supported by the active scheduler path. + +For example: + +```text +2,000,000 ns = 2 ms +20,000,000 ns = 20 ms +``` + +Do not treat the numeric priority value as a general-purpose Linux `nice` level. In the current Gthulhu user-space policy path, the important distinction is whether the priority value is greater than zero. + +!!! warning "Kernel-mode slice-only behavior" + Kernel mode currently does not have a separate normal-priority custom-slice state. A `Priority == 0` slice-only policy therefore has no scheduling effect in kernel mode. Use the user-space Gthulhu scheduler when you need slice-only behavior. + +## Example: Standalone vLLM on DGX Spark + +Assume vLLM is running directly on a host named `dgx-spark-gb10`, outside Kubernetes. + +First, inspect the actual task names: + +```bash +ps -eT -o pid,tid,comm,args | grep -E 'vllm|EngineCore' +``` + +You can also verify a specific process directly: + +```bash +for task in /proc//task/*; do + printf '%s ' "${task##*/}" + cat "$task/comm" +done +``` + +Then create a node policy such as: + +| Field | Example | +| --- | --- | +| Node Names | `dgx-spark-gb10` | +| Command Regex | `^EngineCore(_DP[0-9]+)?$` | +| Priority | `1` | +| Execution Time | `2000000` | + +This example boosts the matching EngineCore tasks and gives them a 2 ms custom slice. + +The values above are an example, not a universal vLLM tuning recommendation. Measure throughput, TTFT/TPOT, scheduler wait time, and system contention before deciding on production values. + +### Why this matters for vLLM + +Kubernetes is not involved in this setup, but Linux still decides when vLLM's CPU-side threads run. Under CPU contention, the CPU-side engine, request-processing, tokenizer, networking, or feeder threads can be delayed even if the GPU itself is available. + +Node policies let Gthulhu protect selected host threads directly instead of requiring users to wrap the inference server in Kubernetes only to gain access to workload-aware scheduling. + +## Creating the Same Policy through the REST API + +The Manager exposes node-policy APIs under `/api/v1`. + +Example: + +```bash +curl -X POST http://:8080/api/v1/node-scheduling-policies \ + -H 'Authorization: Bearer ' \ + -H 'Content-Type: application/json' \ + -d '{ + "nodeNames": ["dgx-spark-gb10"], + "commandRegex": "^EngineCore(_DP[0-9]+)?$", + "priority": 1, + "executionTime": 2000000 + }' +``` + +Useful read endpoints include: + +```text +GET /api/v1/node-scheduling-policies/self +GET /api/v1/node-scheduling-intents/self +``` + +The policy API also supports `nodeSelectors` and `draSelectors` for environments where node selection should be driven by labels or device inventory rather than an exact node name. + +## Verifying the Policy + +After saving a policy, check the **Node Scheduling Intents** table in the Web GUI. + +Intent states are: + +- **Pending**: created but not yet delivered; +- **Sent**: sent toward the target node; +- **Applied**: accepted by the node-side path; +- **Failed**: policy delivery or application failed. + +For a vLLM policy, also verify that the regex actually matches the intended TIDs. A policy can be delivered successfully but still be ineffective if the task names do not match what the workload is currently using. + +## Troubleshooting + +### No node intent is generated + +Check that: + +- the node name exactly matches the node registered with Gthulhu; +- the node label selectors resolve at least one node; +- the Decision Maker is registered and reachable. + +### Intent is Applied but vLLM behavior does not change + +Check that: + +- `Command Regex` matches `/proc//task//comm`; +- the Gthulhu scheduler is enabled on that node; +- the workload is actually CPU-contention-sensitive; +- the policy is using the expected user-space or kernel scheduler mode; +- `Priority == 0` is not being used as a slice-only policy in kernel mode. + +### The policy matches too many tasks + +Use the narrowest possible regex. Node-level policies can target arbitrary host processes, so a broad expression such as `.*` may affect unrelated services on the machine. + +Start with one identifiable vLLM worker thread family, verify the result, and expand the policy only when necessary. + +## Related Documentation + +- [How It Works](how-it-works.md) — TID-aware matching, TID-first lookup, priority semantics, and scheduler internals. +- [Loading sched_ext schedulers](scx-loader.md) — configure the node scheduler runtime. +- [Configuring Scheduling Policies via Web GUI](gui.md) — Kubernetes-oriented scheduling policies. +- [Keeping vLLM Fast Under CPU Pressure](https://vllm-project-github-9ojnpu1yg-simon-mos-projects.vercel.app/2026/08/07/vllm-gthulhu.html) — an example of Gthulhu applied to vLLM under CPU contention. From 2fcd6e81abbcca1e549603c148f27b506dc0623e Mon Sep 17 00:00:00 2001 From: gthulhu-work Date: Mon, 7 Sep 2026 12:25:10 +0800 Subject: [PATCH 2/5] docs: add Chinese node-level scheduling policy guide --- docs/node-policies.zh.md | 211 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 211 insertions(+) create mode 100644 docs/node-policies.zh.md diff --git a/docs/node-policies.zh.md b/docs/node-policies.zh.md new file mode 100644 index 0000000..a334a58 --- /dev/null +++ b/docs/node-policies.zh.md @@ -0,0 +1,211 @@ +# Node-Level Scheduling Policies + +Gthulhu 可以直接對指定節點上的 Linux task 套用排程策略,而不要求該 workload 必須先被 Kubernetes Pod 管理。 + +這特別適合: + +- 直接在主機或 Docker 中執行的 vLLM; +- systemd service 與 host daemon; +- DGX Spark / GB10 上的本地 inference runtime; +- 其他不在 Kubernetes 中、但仍希望使用 workload-aware CPU scheduling 的程序。 + +Node-level policy 會先選出一個或多個節點,再依 Linux task 的 `comm` 名稱比對程序或執行緒,最後把 Gthulhu scheduling strategy 套用到實際匹配到的 task。 + +!!! tip "試用 live mock" + Web GUI mock 已包含 Node Policies 頁面:[gthulhu.github.io/Gthulhu/#/node-policies](https://gthulhu.github.io/Gthulhu/#/node-policies)。 + +## Node Policies 與 Kubernetes Scheduling Policies 的差異 + +兩種 policy 主要差在 workload 的選取方式: + +| Policy 類型 | Workload 選取方式 | 適用場景 | +| --- | --- | --- | +| Kubernetes scheduling policy | Namespace、labels、Pod/container identity | Kubernetes 管理的 workload | +| Node-level scheduling policy | Node identity + Linux task name | Host process、Docker workload、standalone vLLM、system service | + +兩條路徑最後都會把 workload intent 轉成 node-local Linux task scheduling;差異在於目標 task 是怎麼被找到的。 + +## 前提條件 + +若要實際改變 CPU 排程行為,目標節點需要在 Linux 6.12+ 上啟用 `sched_ext`,並執行 Gthulhu scheduler path。 + +Manager 與該節點的 Decision Maker 也需要正常連線,Manager 才能產生 node scheduling intent 並傳送到目標節點。 + +Scheduler 設定可參考[安裝](installation.zh.md)與[載入 sched_ext 排程器](scx-loader.zh.md)。 + +## 透過 Web GUI 建立 Node Policy + +在 Web GUI 中開啟 **Node Policies**,點選 **New Node Policy**。 + +表單包含: + +- **Node Names**:以逗號分隔的精確 node name。對 standalone inference host 而言通常是最直接的選法。 +- **Node Label Selectors**:當多個已註冊節點具有相同 role 或 hardware class 時,可依 label 選取。 +- **DRA Selectors**:Kubernetes / DRA 環境可用的 device-based selection。Standalone vLLM 通常可以留空。 +- **Command Regex**:用正規表示式比對 Linux task name,也就是 `comm`。 +- **Priority**:控制匹配到的 task 是否套用 Gthulhu 的 boosting 行為。 +- **Execution Time**:自訂 scheduling time slice,單位為 nanoseconds。 + +儲存 policy 後,Manager 會為符合條件的節點產生一筆或多筆 **Node Scheduling Intents**。 + +## Task Matching 是 TID-aware 的 + +Linux scheduler 排的是 task/thread,不只是 process leader。Gthulhu 因此會掃描: + +```text +/proc//task//comm +``` + +並對每一個 thread 的名稱獨立套用 `Command Regex`。 + +例如,一個 vLLM process 可能長這樣: + +```text +/proc/3785998/comm = python3.12 +/proc/3785998/task/3786004/comm = EngineCore_DP0 +/proc/3785998/task/3786005/comm = EngineCore_DP1 +``` + +此時可以使用: + +```regex +^EngineCore(_DP[0-9]+)?$ +``` + +直接選到 latency-sensitive 的 vLLM worker threads,即使 process leader 的名稱只是 `python3.12`。 + +解析後的 strategy 會以匹配到的 **TID** 為 key。在 user-space scheduling 中,Gthulhu 會先找 exact TID strategy;找不到時才 fallback 到 TGID strategy。 + +!!! note + 請比對 `/proc/.../comm` 的值,不要直接拿 `ps` 顯示的完整 command line 或 `/proc/.../cmdline` 當作 regex 依據。 + +## Priority 與 Time Slice 語意 + +目前 Gthulhu 的語意是: + +- `Priority > 0`:屬於 **boosting strategy**,匹配到的 task 在 user-space scheduler 中會獲得優先處理。 +- `Priority == 0`:屬於 **non-boosting strategy**。在 user-space mode 下仍可設定 custom time slice,但不會因 policy 而跳到 run queue 前面。 +- `Execution Time`:custom time slice,單位為 nanoseconds;實際效果取決於目前使用的 scheduler path。 + +例如: + +```text +2,000,000 ns = 2 ms +20,000,000 ns = 20 ms +``` + +不要把 `Priority` 的數值直接理解成 Linux `nice` level。目前 user-space policy path 最重要的差異是 priority 是否大於 0。 + +!!! warning "Kernel mode 的 slice-only 行為" + Kernel mode 目前沒有獨立的「normal priority + custom slice」狀態,因此 `Priority == 0` 的 slice-only policy 在 kernel mode 下不會產生排程效果。如果需要 slice-only 行為,請使用 Gthulhu user-space scheduler。 + +## 範例:DGX Spark 上的 Standalone vLLM + +假設 vLLM 直接跑在名為 `dgx-spark-gb10` 的主機上,沒有放進 Kubernetes。 + +先確認實際 task name: + +```bash +ps -eT -o pid,tid,comm,args | grep -E 'vllm|EngineCore' +``` + +也可以直接檢查特定 process: + +```bash +for task in /proc//task/*; do + printf '%s ' "${task##*/}" + cat "$task/comm" +done +``` + +接著建立例如以下 node policy: + +| 欄位 | 範例 | +| --- | --- | +| Node Names | `dgx-spark-gb10` | +| Command Regex | `^EngineCore(_DP[0-9]+)?$` | +| Priority | `1` | +| Execution Time | `2000000` | + +這個範例會 boost 匹配到的 EngineCore task,並套用 2 ms custom slice。 + +以上數值只是示例,不是所有 vLLM workload 都適用的固定最佳值。正式使用前應量測 throughput、TTFT/TPOT、scheduler wait time 與整體 CPU contention,再決定實際策略。 + +### 為什麼這對 vLLM 很重要 + +即使完全沒有 Kubernetes,Linux 仍然決定 vLLM CPU-side threads 什麼時候能取得 CPU。當主機有 CPU contention 時,engine、request-processing、tokenizer、networking 或 feeder thread 都可能被延遲,即使 GPU 本身仍有可用資源。 + +Node policy 讓 Gthulhu 可以直接保護特定 host thread,不需要為了取得 workload-aware scheduling 能力,先把 inference server 包進 Kubernetes。 + +## 透過 REST API 建立相同 Policy + +Manager 提供 `/api/v1` 下的 node-policy API。 + +例如: + +```bash +curl -X POST http://:8080/api/v1/node-scheduling-policies \ + -H 'Authorization: Bearer ' \ + -H 'Content-Type: application/json' \ + -d '{ + "nodeNames": ["dgx-spark-gb10"], + "commandRegex": "^EngineCore(_DP[0-9]+)?$", + "priority": 1, + "executionTime": 2000000 + }' +``` + +常用的查詢 endpoint: + +```text +GET /api/v1/node-scheduling-policies/self +GET /api/v1/node-scheduling-intents/self +``` + +Policy API 也支援 `nodeSelectors` 與 `draSelectors`,適合需要依 node label 或 device inventory 選節點,而不是直接指定 node name 的環境。 + +## 驗證 Policy 是否生效 + +儲存 policy 後,可在 Web GUI 的 **Node Scheduling Intents** 表格查看狀態。 + +Intent state 包含: + +- **Pending**:已建立,但尚未送出; +- **Sent**:已送往目標節點; +- **Applied**:node-side path 已接受; +- **Failed**:policy 傳送或套用失敗。 + +對 vLLM policy 而言,也應確認 regex 真的匹配到預期的 TID。Policy 即使成功送到 node,如果 task name 沒有匹配,仍可能看不到預期的排程效果。 + +## Troubleshooting + +### 沒有產生 Node Intent + +請確認: + +- node name 與 Gthulhu 註冊到的名稱完全一致; +- node label selector 至少能解析到一個節點; +- Decision Maker 已註冊且可連線。 + +### Intent 顯示 Applied,但 vLLM 行為沒有變化 + +請確認: + +- `Command Regex` 是否真的匹配 `/proc//task//comm`; +- 該節點是否已啟用 Gthulhu scheduler; +- workload 是否真的受到 CPU contention 影響; +- 目前使用的是預期的 user-space / kernel scheduler mode; +- kernel mode 下是否誤用了 `Priority == 0` 的 slice-only policy。 + +### Policy 匹配到太多 task + +請盡可能縮小 regex 範圍。Node-level policy 可以直接影響任意 host process,因此像 `.*` 這種過寬的規則可能連其他系統服務一起匹配。 + +建議先從一組明確可辨識的 vLLM worker thread 開始,驗證結果後再逐步擴大範圍。 + +## 相關文件 + +- [運作原理](how-it-works.zh.md) — TID-aware matching、TID-first lookup、priority semantics 與 scheduler internals。 +- [載入 sched_ext 排程器](scx-loader.zh.md) — 設定 node scheduler runtime。 +- [透過 Web GUI 設定 Scheduling Policies](gui.zh.md) — Kubernetes-oriented scheduling policies。 +- [Keeping vLLM Fast Under CPU Pressure](https://vllm-project-github-9ojnpu1yg-simon-mos-projects.vercel.app/2026/08/07/vllm-gthulhu.html) — Gthulhu 在 vLLM CPU contention 情境下的實驗案例。 From 2b18bae0aac9add9b4f17edd663db029f17a9205 Mon Sep 17 00:00:00 2001 From: gthulhu-work Date: Mon, 7 Sep 2026 12:25:27 +0800 Subject: [PATCH 3/5] docs: add node policies to navigation --- mkdocs.yml | 2 ++ 1 file changed, 2 insertions(+) diff --git a/mkdocs.yml b/mkdocs.yml index 62c3dbd..902523b 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -10,6 +10,7 @@ nav: - Installation: installation.md - K8s Deployment: k8s.md - Configuring the scheduling policies: gui.md + - Node-level scheduling policies: node-policies.md - Loading sched_ext schedulers: scx-loader.md - Pod-Level Scheduling Metrics: pod-metrics.md - Working with multi-node cluster: multi-node.md @@ -82,6 +83,7 @@ plugins: Installation: 安裝 K8s Deployment: K8s 部署 Configuring the scheduling policies: 設定排程策略 + Node-level scheduling policies: Node-level 排程策略 Loading sched_ext schedulers: 載入 sched_ext 排程器 Pod-Level Scheduling Metrics: Pod 級排程指標 Working with multi-node cluster: 多節點叢集 From 84a62fb506eec60aab56a1efc6535da615909c7d Mon Sep 17 00:00:00 2001 From: gthulhu-work Date: Mon, 7 Sep 2026 12:25:41 +0800 Subject: [PATCH 4/5] docs: link Kubernetes policies to node policies --- docs/gui.md | 3 +++ 1 file changed, 3 insertions(+) diff --git a/docs/gui.md b/docs/gui.md index 7f9c70f..9527d4d 100644 --- a/docs/gui.md +++ b/docs/gui.md @@ -4,6 +4,9 @@ Gthulhu provides a Web GUI that allows users to conveniently configure scheduling policies. +!!! tip "Scheduling a host process or standalone vLLM?" + This page describes the Kubernetes-oriented scheduling policy flow. If the workload runs directly on a node, in Docker, or outside Kubernetes, use [Node-Level Scheduling Policies](node-policies.md) instead. + > **Note** > If you have deployed Gthulhu on a Kubernetes cluster, please first use the `kubectl port-forward svc/gthulhu-manager 8080:8080` command to forward the local port 8080 to the Gthulhu Manager service, then visit `http://localhost:8080` in your browser to access the Web GUI. From 43a99b1ae23c9de4df82a844b7ac6ca75cc1d3be Mon Sep 17 00:00:00 2001 From: gthulhu-work Date: Mon, 7 Sep 2026 12:25:58 +0800 Subject: [PATCH 5/5] docs: link Kubernetes policies to node policies in Chinese --- docs/gui.zh.md | 3 +++ 1 file changed, 3 insertions(+) diff --git a/docs/gui.zh.md b/docs/gui.zh.md index e71c5d5..0266532 100644 --- a/docs/gui.zh.md +++ b/docs/gui.zh.md @@ -4,6 +4,9 @@ Gthulhu 提供了一個 Web GUI,讓使用者可以方便地設定 scheduling policies。 +!!! tip "要排程 host process 或 standalone vLLM?" + 本頁介紹的是 Kubernetes-oriented scheduling policy flow。如果 workload 直接跑在 node、Docker,或完全不在 Kubernetes 中,請改用 [Node-Level Scheduling Policies](node-policies.zh.md)。 + > !注意 > 如果你將 Gthulhu 部署於 Kubernetes 叢集,請先使用 `kubectl port-forward svc/gthulhu-manager 8080:8080` 命令將本地的 8080 端口轉發到 Gthulhu Manager 的服務上,然後在瀏覽器中訪問 `http://localhost:8080` 即可使用 Web GUI。