From 05fc746b6ac0c9d6b36f1c32282ffc0233e4d9f9 Mon Sep 17 00:00:00 2001 From: maishivamhoo123 Date: Wed, 5 Aug 2026 03:21:51 +0000 Subject: [PATCH 1/3] docs(tutorials): add KAI Scheduler + HAMi-core lab on nvml-mock Signed-off-by: maishivamhoo123 --- .../current/labs/kai-scheduler-hami.md | 356 ++++++++++++++++++ sidebars-tutorials.js | 5 + tutorials/labs/kai-scheduler-hami.md | 355 +++++++++++++++++ 3 files changed, 716 insertions(+) create mode 100644 i18n/zh/docusaurus-plugin-content-docs-tutorials/current/labs/kai-scheduler-hami.md create mode 100644 tutorials/labs/kai-scheduler-hami.md diff --git a/i18n/zh/docusaurus-plugin-content-docs-tutorials/current/labs/kai-scheduler-hami.md b/i18n/zh/docusaurus-plugin-content-docs-tutorials/current/labs/kai-scheduler-hami.md new file mode 100644 index 000000000..6cf803aeb --- /dev/null +++ b/i18n/zh/docusaurus-plugin-content-docs-tutorials/current/labs/kai-scheduler-hami.md @@ -0,0 +1,356 @@ +--- +title: "实验 11: 在假 GPU 上运行 KAI Scheduler + HAMi-core" +description: "在无需真实 GPU 的情况下,安装 NVIDIA KAI Scheduler 与 HAMi-core 隔离,并验证调度与注入控制面。" +sidebar_label: "实验 11: KAI + HAMi (nvml-mock)" +lab: + level: Advanced + duration: 约 50 分钟 + environment: Linux/macOS 笔记本,使用 kind + nvml-mock · 无需真实 GPU + cost: free + authors: + - maishivamhoo123 + verified: "2026-08-05" +tags: + - kai-scheduler + - hami-core +toc_max_heading_level: 2 +--- + +本实验在 [实验 5](./nvml-mock.md) 的 **nvml-mock** 环境之上——即本地 **kind** 集群中的 8 块假 A100 GPU——安装 **NVIDIA KAI Scheduler**(启用 `hamicore` 插件)以及 [kai-resource-isolator](https://github.com/Project-HAMi/KAI-resource-isolator),然后走通整条控制面:整卡调度、按显存分数的分配核算,以及 HAMi-core(`libvgpu.so`)的注入。KAI 负责调度与 GPU 共享,HAMi-core 仅用于显存隔离,这正是两个项目在生产环境中的分工方式。由于假 GPU 的 Pod 内部没有 CUDA/NVML 运行时,本实验验证调度与注入控制面,并有意在 KAI 的预留 Pod(reservation Pod)边界处停止(步骤 8)。 + +:::note + +nvml-mock 提供 GPU 发现能力和节点级 NVML,因此 KAI 调度、显存核算以及隔离器的注入都能正常工作。但它**不会**把 `libnvidia-ml.so` 注入到任意 Pod 中,所以一个真正**运行中**的共享 Pod、`nvidia-smi` 的显存切分显示,以及 `cudaMalloc` 的强制限制,仍然需要真实 GPU(或完整的 NVIDIA 容器工具链 / CDI)。步骤 8 会精确指出这条边界在哪里。 + +::: + +## 你将学到什么 + +- NVIDIA device-plugin 如何在 nvml-mock 上通告整卡资源(`nvidia.com/gpu: 8`) +- KAI 为何需要 `nvidia.com/gpu.memory` 标签,以及为何该标签必须在 KAI 读取节点之前就存在 +- KAI 队列、`hamicore` 插件与隔离器 webhook 如何协同工作 +- GPU **共享**在哪一步开始需要真实(或由工具链注入的)NVML,以及原因 + +## 实验概览 + +```mermaid +%% title: 实验步骤 +flowchart LR + Step1["步骤 1
kind 集群"] --> Step2["步骤 2
nvml-mock"] + Step2 --> Step3["步骤 3
device-plugin
+ gpu.memory"] + Step3 --> Step4["步骤 4
KAI + 队列"] + Step4 --> Step5["步骤 5
隔离器
+ RuntimeClass"] + Step5 --> Step6["步骤 6
整卡 Pod"] + Step6 --> Step7["步骤 7
共享 Pod
+ 注入"] + Step7 --> Step8["步骤 8
预留 Pod
边界"] +``` + +## 前提条件 + +- 一台运行着 **Docker** 的 Linux 或 macOS 笔记本,至少空闲 4 核 CPU / 8 GB 内存 +- `kind` v0.20+、`kubectl` v1.31+、`helm` 3.x、`git`、`go`(安装命令见 [实验 5](./nvml-mock.md)) +- **KAI Scheduler ≥ v0.17.0**(`hamicore` 插件和 `1.1.0-chart` 隔离器所要求) +- 可访问 GitHub、GHCR、Docker Hub 以及 NVCR/NGC(`nvcr.io`,用于 device-plugin 镜像) + +## 步骤 1: 创建 kind 集群 + +启动一个单节点集群,并把节点名保存到变量中,供后续步骤复用。 + +```bash +kind create cluster --name kai-hami-test +NODE_NAME=$(kubectl get nodes -o jsonpath='{.items[0].metadata.name}') +echo "NODE_NAME=${NODE_NAME}" +``` + +```plaintext +NODE_NAME=kai-hami-test-control-plane +``` + +## 步骤 2: 部署 nvml-mock + +nvml-mock 提供假的 `libnvidia-ml.so`、虚拟设备和 PCI 拓扑,使节点报告 8 块 A100 GPU。这与实验 5 使用的是同一个模拟器。 + +```bash +git clone https://github.com/NVIDIA/k8s-test-infra.git +cd k8s-test-infra +docker build -t nvml-mock:local -f deployments/nvml-mock/Dockerfile . +kind load docker-image nvml-mock:local --name kai-hami-test + +helm install nvml-mock oci://ghcr.io/nvidia/k8s-test-infra/chart/nvml-mock \ + --set image.repository=nvml-mock --set image.tag=local \ + --wait --timeout 120s + +kubectl get node ${NODE_NAME} \ + -o custom-columns=NAME:.metadata.name,GPU_PRESENT:.metadata.labels.nvidia\\.com/gpu\\.present +``` + +```plaintext +NAME GPU_PRESENT +kai-hami-test-control-plane true +``` + +> `GPU_PRESENT=true` 表示该节点已成为 GPU 节点。如果该列为空,说明 nvml-mock 的 Pod 尚未启动完成——稍等片刻后重新运行最后一条命令。 + +## 步骤 3: 安装 device-plugin 并发布 GPU 显存 + +KAI 依据 `nvidia.com/gpu` 进行调度,因此需要安装 NVIDIA device-plugin。请使用 nvml-mock 自带的清单——它会以 hostPath 方式挂载 `/var/lib/nvml-mock`,并传入正确的 driver-root 和 NVML 发现参数,而普通 Helm chart 不会这样做(那会导致 `ERROR_LIBRARY_NOT_FOUND`)。随后发布每卡显存,KAI 需要它来把 `gpu-memory` 请求换算成分数。 + +```bash +kubectl apply -f https://raw.githubusercontent.com/NVIDIA/k8s-test-infra/main/tests/e2e/device-plugin-mock.yaml +kubectl -n kube-system wait --for=condition=ready \ + pod -l name=nvidia-device-plugin-mock --timeout=120s + +kubectl get node ${NODE_NAME} -o jsonpath='{.status.allocatable.nvidia\.com/gpu}{"\n"}' + +kubectl label node ${NODE_NAME} \ + nvidia.com/gpu.memory=40960 \ + nvidia.com/gpu.product=NVIDIA-A100-SXM4-40GB \ + nvidia.com/gpu.count=8 --overwrite +``` + +```plaintext +8 +node/kai-hami-test-control-plane labeled +``` + +> `8`(而非 `80`)确认 KAI 拿到的是 8 块整卡,由 KAI 自己去做分数级共享。请务必在安装 KAI **之前**打上 `nvidia.com/gpu.memory` 标签:KAI 在首次注册节点时会缓存每卡显存,若标签是之后才添加的,KAI 会把显存记为 0,导致每个共享 Pod 都停留在 `Pending` 状态并报 `didn't have enough resources: GPU memory`,直到重启调度器为止。 + +## 步骤 4: 安装 KAI Scheduler 并创建队列 + +`global.gpuSharing=true` 启用 GPU 共享;`binder.plugins.hamicore.enabled=true` 让 KAI 向共享容器注入 `CUDA_DEVICE_MEMORY_LIMIT`。在 Pod 的 `kai.scheduler/queue` 标签所指向的队列存在之前,KAI 不会调度该 Pod。 + +```bash +helm install kai-scheduler oci://ghcr.io/kai-scheduler/kai-scheduler/kai-scheduler \ + --namespace kai-scheduler --create-namespace \ + --set global.gpuSharing=true \ + --set binder.plugins.hamicore.enabled=true \ + --version v0.17.0 + +# 在创建队列之前,先等待 KAI(尤其是 admission webhook)就绪, +# 否则证书尚未就绪时 Queue 的创建可能会被拒绝。 +kubectl -n kai-scheduler wait --for=condition=available --timeout=180s deploy --all + +kubectl apply -f - <<'EOF' +apiVersion: scheduling.run.ai/v2 +kind: Queue +metadata: + name: default +spec: + resources: + cpu: { quota: -1, limit: -1, overQuotaWeight: 1 } + memory: { quota: -1, limit: -1, overQuotaWeight: 1 } + gpu: { quota: -1, limit: -1, overQuotaWeight: 1 } +--- +apiVersion: scheduling.run.ai/v2 +kind: Queue +metadata: + name: default-queue +spec: + parentQueue: default + resources: + cpu: { quota: -1, limit: -1, overQuotaWeight: 1 } + memory: { quota: -1, limit: -1, overQuotaWeight: 1 } + gpu: { quota: -1, limit: -1, overQuotaWeight: 1 } +EOF + +kubectl get pods -n kai-scheduler +kubectl get queues +``` + +```plaintext +NAME READY STATUS RESTARTS AGE +admission-759b9bb99c-4wx9q 1/1 Running 0 2m +binder-69bf5f648-k572n 1/1 Running 0 2m +kai-operator-997c6886c-dthws 1/1 Running 0 2m +kai-scheduler-default-5dfbc85f96-6kp9v 1/1 Running 0 2m +pod-grouper-68f4fb47-5q99f 1/1 Running 0 2m +podgroup-controller-5947b5b4dd-f66pj 1/1 Running 0 2m +queue-controller-6cc8c844c8-sdl67 1/1 Running 0 2m + +NAME PRIORITY PARENT CHILDREN DISPLAYNAME +default ["default-queue"] +default-parent-queue +default-queue default +``` + +> 所有 KAI Pod 大约在两分钟内进入 `Running`(admission/webhook 的证书最后就绪)。`default-parent-queue` 由 operator 自动创建;你的 Pod 使用 `default-queue`。 + +## 步骤 5: 部署隔离器与 RuntimeClass 垫片 + +隔离器会把 HAMi-core 分发到节点,并运行一个 webhook,向共享 Pod 注入 `libvgpu.so` 和 `ld.so.preload`。该 webhook 还会添加 `runtimeClassName: nvidia`;而 kind 集群没有这个运行时,因此需要把 `nvidia` 映射到 `runc`,否则创建 Pod 时会被拒绝并报 `RuntimeClass "nvidia" not found`。 + +```bash +helm install kai-resource-isolator oci://docker.io/projecthami/kai-resource-isolator \ + --namespace kai-resource-isolator --create-namespace \ + --set monitor.enabled=true --version 1.1.0-chart + +kubectl apply -f - <<'EOF' +apiVersion: node.k8s.io/v1 +kind: RuntimeClass +metadata: + name: nvidia +handler: runc +EOF + +kubectl get pods -n kai-resource-isolator +kubectl logs -n kai-resource-isolator deploy/kai-resource-isolator-webhook | tail -1 +``` + +```plaintext +NAME READY STATUS RESTARTS AGE +kai-resource-isolator-libsync-w7vms 1/1 Running 0 40s +kai-resource-isolator-monitor-wmh4f 1/1 Running 0 40s +kai-resource-isolator-webhook-776dd4c45c-nk6sn 1/1 Running 0 40s + +2026/08/05 02:27:08 webhook starting listen=:8443 containerVgpuMount=/usr/local/vgpu annotationKeys=gpu-fraction|gpu-memory +``` + +> `annotationKeys=gpu-fraction|gpu-memory` 确认该 webhook 会对携带任一注解的 Pod 进行改写。`monitor` Pod 在 mock 上探测 NVML 时可能短暂显示 `0/1`,这不影响本实验的其余部分。 + +## 步骤 6: 用整卡 Pod 验证基础调度 + +整卡请求(`nvidia.com/gpu: 1`,不带共享注解)不会触发预留 Pod,因此能干净地运行起来,从而证明 KAI 的调度、绑定以及 RuntimeClass 在 mock 上都工作正常。 + +```bash +kubectl apply -f - <<'EOF' +apiVersion: v1 +kind: Pod +metadata: + name: kai-whole-gpu + labels: + kai.scheduler/queue: default-queue +spec: + schedulerName: kai-scheduler + runtimeClassName: nvidia + containers: + - name: app + image: busybox + command: ["sleep","3600"] + resources: + limits: + nvidia.com/gpu: 1 +EOF + +kubectl get pod kai-whole-gpu -o wide -w +``` + +```plaintext +NAME READY STATUS RESTARTS AGE IP NODE +kai-whole-gpu 1/1 Running 0 10s 10.244.0.22 kai-hami-test-control-plane +``` + +> `Running` 端到端确认了基础 GPU 通路。保留该 Pod 不删——步骤 7 会用它来展示 KAI 的队列核算。 + +## 步骤 7: 调度共享 GPU Pod 并检查注入 + +`gpu-memory` 注解使其成为共享 Pod,且不带 `nvidia.com/gpu` 请求——由 KAI 预留分数。`20480` MiB 是半块 A100。`CUDA_DISABLE_CONTROL=true` 可避免 HAMi-core 因缺少 CUDA 驱动而中止。 + +```bash +kubectl apply -f - <<'EOF' +apiVersion: v1 +kind: Pod +metadata: + name: gpu-sharing-with-isolation + labels: + kai.scheduler/queue: default-queue + annotations: + gpu-memory: "20480" +spec: + schedulerName: kai-scheduler + containers: + - name: gpu-workload + image: busybox + command: ["sleep", "3600"] + env: + - name: CUDA_DISABLE_CONTROL + value: "true" +EOF + +kubectl logs -n kai-scheduler deploy/kai-scheduler-default --tail=40 \ + | grep -iE 'resource division result for queue ' + +kubectl get pod gpu-sharing-with-isolation -o yaml \ + | grep -iE 'runtimeClassName|ld.so.preload|vgpu' +``` + +```plaintext +Resource division result for queue : ... GPU: requested: <1.51>, allocated: <1.51>, fairShare: <1.51> ... + + - name: CONTAINER_VGPU_MOUNT + value: /usr/local/vgpu + - mountPath: /usr/local/vgpu + name: kai-resource-isolator-vgpu + - mountPath: /etc/ld.so.preload + name: kai-resource-isolator-vgpu + subPath: ld.so.preload + - mountPath: /usr/local/vgpu/containers + - mountPath: /tmp/vgpulock + name: kai-resource-isolator-vgpulock + runtimeClassName: nvidia + path: /usr/local/vgpu + name: kai-resource-isolator-vgpu + path: /usr/local/vgpu/containers + path: /tmp/vgpulock + name: kai-resource-isolator-vgpulock +``` + +> KAI 读取了 GPU 显存并分配了分数:`1.51` 是队列总量——即步骤 6 的整卡 Pod(`1.0`)加上这个共享 Pod 的切片。该切片为整卡 `40960` MiB 中的 `20480` MiB,约为一半;KAI 会按两位小数精度把它换算为 GPU 分数(向上取整,因此是 `0.51` 而非 `0.50`)。第二段是隔离器的注入:`/etc/ld.so.preload` 和 `/usr/local/vgpu`(HAMi-core),以及步骤 5 所处理的 `runtimeClassName: nvidia`。这两者在 mock 上都已完全验证。 + +## 步骤 8: 观察预留 Pod 边界 + +共享 Pod 会保持 `Pending`——这是假 GPU 的诚实边界。为了共享一张卡,KAI 会启动一个预留 Pod,它在**容器内部**调用 NVML 来占用设备,而这个调用会失败。 + +```bash +kubectl describe pod gpu-sharing-with-isolation | sed -n '/Events/,$p' | tail -3 + +# KAI 每次绑定尝试都会重建预留 Pod,其名称会变化。 +# 抓取当前存在的那一个并查看它的日志。 +RPOD=$(kubectl get pods -n kai-resource-reservation \ + -o jsonpath='{.items[0].metadata.name}') +kubectl logs -n kai-resource-reservation "$RPOD" --all-containers +``` + +```plaintext + Warning BindingError ... binder Failed to bind pod default/gpu-sharing-with-isolation ...: + failed to reserve GPUs ...: failed waiting for GPU reservation pod to allocate: + kai-resource-reservation/gpu-reservation-fc9757493a36c25e + +INFO Looking for GPU device id for pod {"name": "gpu-reservation-5df74f36086ed6c5"} +Error while running the app: unable to initialize NVML: ERROR_LIBRARY_NOT_FOUND +``` + +> KAI 每次绑定尝试都会重建预留 Pod,因此上面 `BindingError` 中的名称和你这里读到的名称会不同——这是预期行为,每次尝试都以同样的方式失败。预留 Pod 拿到了 GPU 设备节点,但没有拿到驱动**库**——把 nvml-mock 的 `libnvidia-ml.so` 注入 Pod 是 NVIDIA 容器工具链 / CDI 的职责,而普通 kind 集群并未配置它。于是 NVML 初始化失败,绑定始终无法完成。跨越这条边界需要真实 GPU,或针对 mock 驱动根目录配置完整的工具链 / CDI 栈——两者都超出了笔记本实验的范围。 + +## 清理 + +```bash +kubectl delete pod gpu-sharing-with-isolation kai-whole-gpu --ignore-not-found +kubectl delete queue default-queue default --ignore-not-found +kubectl delete runtimeclass nvidia --ignore-not-found + +helm uninstall kai-resource-isolator -n kai-resource-isolator +helm uninstall kai-scheduler -n kai-scheduler +kubectl delete -f https://raw.githubusercontent.com/NVIDIA/k8s-test-infra/main/tests/e2e/device-plugin-mock.yaml --ignore-not-found +helm uninstall nvml-mock + +kind delete cluster --name kai-hami-test +``` + +## 本实验验证了什么 + +| 声明 | 证据 | +| ----------------------- | -------------------------------------------------------- | +| mock 通告整卡 GPU | 可分配 `nvidia.com/gpu: 8`(步骤 3) | +| KAI 安装并运行 | 所有 `kai-scheduler` Pod 均为 `Running`(步骤 4) | +| 队列对调度生效 | 创建了 `default` / `default-queue`(步骤 4) | +| 隔离器 webhook 已激活 | `annotationKeys=gpu-fraction\|gpu-memory`(步骤 5) | +| KAI 基础调度正常 | `kai-whole-gpu` 进入 `Running`(步骤 6) | +| KAI 读取每卡显存 | 分数级 `allocated: <1.51>`(步骤 7) | +| 隔离器注入 HAMi-core | Pod 上出现 `ld.so.preload` + `/usr/local/vgpu`(步骤 7) | +| 边界:共享 Pod 无法运行 | 预留 Pod 报 `NVML: ERROR_LIBRARY_NOT_FOUND`(步骤 8) | + +## 下一步 + +- 在真实 GPU 节点上运行步骤 7 的清单,即可看到预留 Pod 成功、共享 Pod 强制执行其上限——参见 [如何将 KAI Scheduler 与 HAMi 一起使用](https://project-hami.io/docs/next/userguide/kai-scheduler/how-to-use-kai-scheduler)。 +- 与 [实验 5: nvml-mock](./nvml-mock.md) 对比:那里 HAMi 自带的调度器和 device-plugin 会把每块 GPU 切成 10 份——这条路径无需预留 Pod,可以在 mock 上完整运行。 +- 阅读 [kai-resource-isolator 仓库](https://github.com/Project-HAMi/KAI-resource-isolator) 了解隔离器内部实现。 diff --git a/sidebars-tutorials.js b/sidebars-tutorials.js index e3881dc71..d88f08e15 100644 --- a/sidebars-tutorials.js +++ b/sidebars-tutorials.js @@ -61,6 +61,11 @@ module.exports = { id: "labs/topology-aware-scheduling", customProps: { level: "Intermediate", duration: "about 45 minutes" }, }, + { + type: "doc", + id: "labs/kai-scheduler-hami", + customProps: { level: "Advanced", duration: "about 50 minutes" }, + }, ], }, ], diff --git a/tutorials/labs/kai-scheduler-hami.md b/tutorials/labs/kai-scheduler-hami.md new file mode 100644 index 000000000..d194fffdc --- /dev/null +++ b/tutorials/labs/kai-scheduler-hami.md @@ -0,0 +1,355 @@ +--- +title: "Lab 11: KAI Scheduler + HAMi-core on a Fake GPU" +sidebar_label: "Lab 11: KAI + HAMi (nvml-mock)" +lab: + level: Advanced + duration: about 50 minutes + environment: Linux/macOS laptop with kind + nvml-mock · no real GPU required + cost: free + authors: + - maishivamhoo123 + verified: "2026-08-05" +tags: + - kai-scheduler + - hami-core +toc_max_heading_level: 2 +--- + +This lab installs **NVIDIA KAI Scheduler** with the `hamicore` plugin and the [kai-resource-isolator](https://github.com/Project-HAMi/KAI-resource-isolator) on top of the **nvml-mock** environment from [Lab 5](./nvml-mock.md) — 8 fake A100 GPUs in a local **kind** cluster — then walks the full control plane: whole-GPU scheduling, fractional GPU-memory accounting, and HAMi-core (`libvgpu.so`) injection. KAI owns scheduling and GPU sharing; HAMi-core is brought in only for GPU-memory isolation, exactly as the two projects divide the work in production. Because a fake GPU has no CUDA/NVML runtime inside pods, the lab verifies the scheduling and injection control plane and stops, on purpose, at KAI's reservation-Pod boundary (Step 8). + +:::note + +nvml-mock provides GPU discovery and node-level NVML, so KAI scheduling, memory accounting, and the isolator's injection all work. It does **not** inject `libnvidia-ml.so` into arbitrary pods, so a _running_ shared Pod, `nvidia-smi` memory slicing, and `cudaMalloc` enforcement still require a real GPU (or the full NVIDIA container toolkit / CDI). Step 8 shows exactly where that line is. + +::: + +## What You'll Learn + +- How the NVIDIA device-plugin advertises whole GPUs (`nvidia.com/gpu: 8`) against nvml-mock +- Why KAI needs the `nvidia.com/gpu.memory` label — and why it must exist before KAI reads the node +- How KAI queues, the `hamicore` plugin, and the isolator webhook fit together +- The exact point where GPU _sharing_ needs real (or toolkit-injected) NVML, and why + +## Lab Overview + +```mermaid +%% title: Lab Flow +flowchart LR + Step1["Step 1
kind cluster"] --> Step2["Step 2
nvml-mock"] + Step2 --> Step3["Step 3
device-plugin
+ gpu.memory"] + Step3 --> Step4["Step 4
KAI + queues"] + Step4 --> Step5["Step 5
isolator
+ RuntimeClass"] + Step5 --> Step6["Step 6
whole-GPU Pod"] + Step6 --> Step7["Step 7
shared Pod
+ injection"] + Step7 --> Step8["Step 8
reservation
boundary"] +``` + +## Prerequisites + +- A Linux or macOS laptop with **Docker** running and at least 4 CPU / 8 GB RAM free +- `kind` v0.20+, `kubectl` v1.31+, `helm` 3.x, `git`, `go` (install commands in [Lab 5](./nvml-mock.md)) +- **KAI Scheduler ≥ v0.17.0** (required by the `hamicore` plugin and the `1.1.0-chart` isolator) +- Access to GitHub, GHCR, Docker Hub, and NVCR/NGC (`nvcr.io` — the device-plugin image) + +## Step 1: Create the kind Cluster + +Bootstrap a single-node cluster and store its node name in a variable that later steps reuse. + +```bash +kind create cluster --name kai-hami-test +NODE_NAME=$(kubectl get nodes -o jsonpath='{.items[0].metadata.name}') +echo "NODE_NAME=${NODE_NAME}" +``` + +```plaintext +NODE_NAME=kai-hami-test-control-plane +``` + +## Step 2: Deploy nvml-mock + +nvml-mock provides a fake `libnvidia-ml.so`, virtual devices, and PCI topology so the node reports 8 A100 GPUs. This is the same simulator as Lab 5. + +```bash +git clone https://github.com/NVIDIA/k8s-test-infra.git +cd k8s-test-infra +docker build -t nvml-mock:local -f deployments/nvml-mock/Dockerfile . +kind load docker-image nvml-mock:local --name kai-hami-test + +helm install nvml-mock oci://ghcr.io/nvidia/k8s-test-infra/chart/nvml-mock \ + --set image.repository=nvml-mock --set image.tag=local \ + --wait --timeout 120s + +kubectl get node ${NODE_NAME} \ + -o custom-columns=NAME:.metadata.name,GPU_PRESENT:.metadata.labels.nvidia\\.com/gpu\\.present +``` + +```plaintext +NAME GPU_PRESENT +kai-hami-test-control-plane true +``` + +> `GPU_PRESENT=true` means the node is now a GPU node. If it is empty, the nvml-mock pods have not finished starting — wait and re-run the last command. + +## Step 3: Install the Device-Plugin and Publish GPU Memory + +KAI schedules against `nvidia.com/gpu`, so install the NVIDIA device-plugin. Use nvml-mock's own manifest — it hostPath-mounts `/var/lib/nvml-mock` and passes the right driver-root and NVML discovery flags, which the plain Helm chart does not (that fails with `ERROR_LIBRARY_NOT_FOUND`). Then publish per-GPU memory, which KAI needs to turn a `gpu-memory` request into a fraction. + +```bash +kubectl apply -f https://raw.githubusercontent.com/NVIDIA/k8s-test-infra/main/tests/e2e/device-plugin-mock.yaml +kubectl -n kube-system wait --for=condition=ready \ + pod -l name=nvidia-device-plugin-mock --timeout=120s + +kubectl get node ${NODE_NAME} -o jsonpath='{.status.allocatable.nvidia\.com/gpu}{"\n"}' + +kubectl label node ${NODE_NAME} \ + nvidia.com/gpu.memory=40960 \ + nvidia.com/gpu.product=NVIDIA-A100-SXM4-40GB \ + nvidia.com/gpu.count=8 --overwrite +``` + +```plaintext +8 +node/kai-hami-test-control-plane labeled +``` + +> `8` (not `80`) confirms KAI gets 8 whole GPUs to share itself. Apply the `nvidia.com/gpu.memory` label **before** installing KAI: KAI caches per-GPU memory when it first registers the node, so a label added later leaves memory at 0 and every shared Pod stays `Pending` with `didn't have enough resources: GPU memory` until the scheduler is restarted. + +## Step 4: Install KAI Scheduler and Create Queues + +`global.gpuSharing=true` enables GPU sharing; `binder.plugins.hamicore.enabled=true` makes KAI inject `CUDA_DEVICE_MEMORY_LIMIT` into shared containers. KAI will not schedule a Pod until the queue named in its `kai.scheduler/queue` label exists. + +```bash +helm install kai-scheduler oci://ghcr.io/kai-scheduler/kai-scheduler/kai-scheduler \ + --namespace kai-scheduler --create-namespace \ + --set global.gpuSharing=true \ + --set binder.plugins.hamicore.enabled=true \ + --version v0.17.0 + +# Wait for KAI (esp. the admission webhook) to be ready before creating Queues, +# or the Queue apply can be rejected while certs are still settling. +kubectl -n kai-scheduler wait --for=condition=available --timeout=180s deploy --all + +kubectl apply -f - <<'EOF' +apiVersion: scheduling.run.ai/v2 +kind: Queue +metadata: + name: default +spec: + resources: + cpu: { quota: -1, limit: -1, overQuotaWeight: 1 } + memory: { quota: -1, limit: -1, overQuotaWeight: 1 } + gpu: { quota: -1, limit: -1, overQuotaWeight: 1 } +--- +apiVersion: scheduling.run.ai/v2 +kind: Queue +metadata: + name: default-queue +spec: + parentQueue: default + resources: + cpu: { quota: -1, limit: -1, overQuotaWeight: 1 } + memory: { quota: -1, limit: -1, overQuotaWeight: 1 } + gpu: { quota: -1, limit: -1, overQuotaWeight: 1 } +EOF + +kubectl get pods -n kai-scheduler +kubectl get queues +``` + +```plaintext +NAME READY STATUS RESTARTS AGE +admission-759b9bb99c-4wx9q 1/1 Running 0 2m +binder-69bf5f648-k572n 1/1 Running 0 2m +kai-operator-997c6886c-dthws 1/1 Running 0 2m +kai-scheduler-default-5dfbc85f96-6kp9v 1/1 Running 0 2m +pod-grouper-68f4fb47-5q99f 1/1 Running 0 2m +podgroup-controller-5947b5b4dd-f66pj 1/1 Running 0 2m +queue-controller-6cc8c844c8-sdl67 1/1 Running 0 2m + +NAME PRIORITY PARENT CHILDREN DISPLAYNAME +default ["default-queue"] +default-parent-queue +default-queue default +``` + +> All KAI pods reach `Running` in about two minutes (the admission/webhook certs settle last). `default-parent-queue` is auto-created by the operator; your Pods target `default-queue`. + +## Step 5: Deploy the Isolator and a RuntimeClass Shim + +The isolator ships HAMi-core to the node and runs a webhook that injects `libvgpu.so` and `ld.so.preload` into shared Pods. That webhook also adds `runtimeClassName: nvidia`; a kind cluster has no such runtime, so map `nvidia` to `runc` or Pod creation is rejected with `RuntimeClass "nvidia" not found`. + +```bash +helm install kai-resource-isolator oci://docker.io/projecthami/kai-resource-isolator \ + --namespace kai-resource-isolator --create-namespace \ + --set monitor.enabled=true --version 1.1.0-chart + +kubectl apply -f - <<'EOF' +apiVersion: node.k8s.io/v1 +kind: RuntimeClass +metadata: + name: nvidia +handler: runc +EOF + +kubectl get pods -n kai-resource-isolator +kubectl logs -n kai-resource-isolator deploy/kai-resource-isolator-webhook | tail -1 +``` + +```plaintext +NAME READY STATUS RESTARTS AGE +kai-resource-isolator-libsync-w7vms 1/1 Running 0 40s +kai-resource-isolator-monitor-wmh4f 1/1 Running 0 40s +kai-resource-isolator-webhook-776dd4c45c-nk6sn 1/1 Running 0 40s + +2026/08/05 02:27:08 webhook starting listen=:8443 containerVgpuMount=/usr/local/vgpu annotationKeys=gpu-fraction|gpu-memory +``` + +> `annotationKeys=gpu-fraction|gpu-memory` confirms the webhook will mutate any Pod carrying either annotation. The `monitor` pod may briefly show `0/1` while it probes NVML on the mock; it does not affect the rest of the lab. + +## Step 6: Verify Base Scheduling with a Whole-GPU Pod + +A whole-GPU request (`nvidia.com/gpu: 1`, no sharing annotation) does not trigger a reservation Pod, so it runs cleanly and proves KAI scheduling, binding, and the RuntimeClass all work on the mock. + +```bash +kubectl apply -f - <<'EOF' +apiVersion: v1 +kind: Pod +metadata: + name: kai-whole-gpu + labels: + kai.scheduler/queue: default-queue +spec: + schedulerName: kai-scheduler + runtimeClassName: nvidia + containers: + - name: app + image: busybox + command: ["sleep","3600"] + resources: + limits: + nvidia.com/gpu: 1 +EOF + +kubectl get pod kai-whole-gpu -o wide -w +``` + +```plaintext +NAME READY STATUS RESTARTS AGE IP NODE +kai-whole-gpu 1/1 Running 0 10s 10.244.0.22 kai-hami-test-control-plane +``` + +> `Running` confirms the base GPU path end to end. Leave this Pod running — Step 7 uses it to show KAI's queue accounting. + +## Step 7: Schedule a Shared-GPU Pod and Inspect the Injection + +A `gpu-memory` annotation makes this a shared Pod, with no `nvidia.com/gpu` request — KAI reserves the fraction. `20480` MiB is half an A100. `CUDA_DISABLE_CONTROL=true` keeps HAMi-core from aborting on the missing CUDA driver. + +```bash +kubectl apply -f - <<'EOF' +apiVersion: v1 +kind: Pod +metadata: + name: gpu-sharing-with-isolation + labels: + kai.scheduler/queue: default-queue + annotations: + gpu-memory: "20480" +spec: + schedulerName: kai-scheduler + containers: + - name: gpu-workload + image: busybox + command: ["sleep", "3600"] + env: + - name: CUDA_DISABLE_CONTROL + value: "true" +EOF + +kubectl logs -n kai-scheduler deploy/kai-scheduler-default --tail=40 \ + | grep -iE 'resource division result for queue ' + +kubectl get pod gpu-sharing-with-isolation -o yaml \ + | grep -iE 'runtimeClassName|ld.so.preload|vgpu' +``` + +```plaintext +Resource division result for queue : ... GPU: requested: <1.51>, allocated: <1.51>, fairShare: <1.51> ... + + - name: CONTAINER_VGPU_MOUNT + value: /usr/local/vgpu + - mountPath: /usr/local/vgpu + name: kai-resource-isolator-vgpu + - mountPath: /etc/ld.so.preload + name: kai-resource-isolator-vgpu + subPath: ld.so.preload + - mountPath: /usr/local/vgpu/containers + - mountPath: /tmp/vgpulock + name: kai-resource-isolator-vgpulock + runtimeClassName: nvidia + path: /usr/local/vgpu + name: kai-resource-isolator-vgpu + path: /usr/local/vgpu/containers + path: /tmp/vgpulock + name: kai-resource-isolator-vgpulock +``` + +> KAI read the GPU memory and allocated the fraction: `1.51` is the queue total — the whole-GPU Pod from Step 6 (`1.0`) plus this shared Pod's slice. The slice is `20480` MiB of the card's `40960` MiB — about half; KAI converts it to a GPU fraction at two-decimal precision (rounding up, so `0.51` rather than a bare `0.50`). The second block is the isolator's injection: `/etc/ld.so.preload` and `/usr/local/vgpu` (HAMi-core) plus the `runtimeClassName: nvidia` that Step 5 accounts for. Both are fully verified on the mock. + +## Step 8: Observe the Reservation-Pod Boundary + +The shared Pod stays `Pending` — the honest limit of a fake GPU. To share a card, KAI launches a reservation Pod that calls NVML **inside its container** to claim the device, and that call fails. + +```bash +kubectl describe pod gpu-sharing-with-isolation | sed -n '/Events/,$p' | tail -3 + +# KAI recreates the reservation Pod on each bind attempt, so its name changes. +# Grab whichever one currently exists and read its log. +RPOD=$(kubectl get pods -n kai-resource-reservation \ + -o jsonpath='{.items[0].metadata.name}') +kubectl logs -n kai-resource-reservation "$RPOD" --all-containers +``` + +```plaintext + Warning BindingError ... binder Failed to bind pod default/gpu-sharing-with-isolation ...: + failed to reserve GPUs ...: failed waiting for GPU reservation pod to allocate: + kai-resource-reservation/gpu-reservation-fc9757493a36c25e + +INFO Looking for GPU device id for pod {"name": "gpu-reservation-5df74f36086ed6c5"} +Error while running the app: unable to initialize NVML: ERROR_LIBRARY_NOT_FOUND +``` + +> KAI recreates the reservation Pod on every bind attempt, so the name in the `BindingError` above and the one you read here will differ — that's expected, and each attempt fails the same way. The reservation Pod gets the GPU device nodes but not the driver _libraries_ — injecting nvml-mock's `libnvidia-ml.so` into a pod is the job of the NVIDIA container toolkit / CDI, which a plain kind cluster does not wire up. So NVML init fails and the bind never completes. Crossing this boundary needs a real GPU, or the full toolkit/CDI stack against the mock driver root — both out of scope for a laptop lab. + +## Cleanup + +```bash +kubectl delete pod gpu-sharing-with-isolation kai-whole-gpu --ignore-not-found +kubectl delete queue default-queue default --ignore-not-found +kubectl delete runtimeclass nvidia --ignore-not-found + +helm uninstall kai-resource-isolator -n kai-resource-isolator +helm uninstall kai-scheduler -n kai-scheduler +kubectl delete -f https://raw.githubusercontent.com/NVIDIA/k8s-test-infra/main/tests/e2e/device-plugin-mock.yaml --ignore-not-found +helm uninstall nvml-mock + +kind delete cluster --name kai-hami-test +``` + +## What This Lab Proved + +| Claim | Evidence | +| ------------------------------- | -------------------------------------------------------- | +| Mock advertises whole GPUs | `nvidia.com/gpu: 8` allocatable (Step 3) | +| KAI installs and runs | all `kai-scheduler` pods `Running` (Step 4) | +| Queues gate scheduling | `default` / `default-queue` created (Step 4) | +| Isolator webhook is active | `annotationKeys=gpu-fraction\|gpu-memory` (Step 5) | +| KAI base scheduling works | `kai-whole-gpu` reaches `Running` (Step 6) | +| KAI reads per-GPU memory | fractional `allocated: <1.51>` (Step 7) | +| Isolator injects HAMi-core | `ld.so.preload` + `/usr/local/vgpu` on the Pod (Step 7) | +| Boundary: shared Pod cannot run | reservation Pod `NVML: ERROR_LIBRARY_NOT_FOUND` (Step 8) | + +## Next Steps + +- Run the Step 7 manifest on a real GPU node to see the reservation Pod succeed and the shared Pod enforce its cap — see [How to use KAI Scheduler with HAMi](https://project-hami.io/docs/next/userguide/kai-scheduler/how-to-use-kai-scheduler). +- Compare with [Lab 5: nvml-mock](./nvml-mock.md), where HAMi's own scheduler and device-plugin slice each GPU into 10 — a path that needs no reservation Pod and runs fully on the mock. +- Read the isolator internals in the [kai-resource-isolator repository](https://github.com/Project-HAMi/KAI-resource-isolator). From a062462630beeddf76c7eb3f8ae54515b9c0f127 Mon Sep 17 00:00:00 2001 From: maishivamhoo123 Date: Wed, 5 Aug 2026 04:25:51 +0000 Subject: [PATCH 2/3] Added the description Signed-off-by: maishivamhoo123 --- tutorials/labs/kai-scheduler-hami.md | 1 + 1 file changed, 1 insertion(+) diff --git a/tutorials/labs/kai-scheduler-hami.md b/tutorials/labs/kai-scheduler-hami.md index d194fffdc..92476725b 100644 --- a/tutorials/labs/kai-scheduler-hami.md +++ b/tutorials/labs/kai-scheduler-hami.md @@ -1,5 +1,6 @@ --- title: "Lab 11: KAI Scheduler + HAMi-core on a Fake GPU" +description: "Install NVIDIA KAI Scheduler with HAMi-core isolation and verify the scheduling and injection control plane — no real GPU required." sidebar_label: "Lab 11: KAI + HAMi (nvml-mock)" lab: level: Advanced From c9a5f2bf019246d1509f891f6a7567adf918e57c Mon Sep 17 00:00:00 2001 From: maishivamhoo123 Date: Sat, 8 Aug 2026 08:37:44 +0000 Subject: [PATCH 3/3] Changed according to the suggestions Signed-off-by: maishivamhoo123 --- .../current/labs/kai-scheduler-hami.md | 7 ++----- tutorials/labs/kai-scheduler-hami.md | 7 ++----- 2 files changed, 4 insertions(+), 10 deletions(-) diff --git a/i18n/zh/docusaurus-plugin-content-docs-tutorials/current/labs/kai-scheduler-hami.md b/i18n/zh/docusaurus-plugin-content-docs-tutorials/current/labs/kai-scheduler-hami.md index 6cf803aeb..615d40428 100644 --- a/i18n/zh/docusaurus-plugin-content-docs-tutorials/current/labs/kai-scheduler-hami.md +++ b/i18n/zh/docusaurus-plugin-content-docs-tutorials/current/labs/kai-scheduler-hami.md @@ -189,7 +189,7 @@ apiVersion: node.k8s.io/v1 kind: RuntimeClass metadata: name: nvidia -handler: runc +handler: runc # 仅限 mock —— 切勿在真实 GPU 节点上应用 EOF kubectl get pods -n kai-resource-isolator @@ -243,7 +243,7 @@ kai-whole-gpu 1/1 Running 0 10s 10.244.0.22 kai-hami-test-c ## 步骤 7: 调度共享 GPU Pod 并检查注入 -`gpu-memory` 注解使其成为共享 Pod,且不带 `nvidia.com/gpu` 请求——由 KAI 预留分数。`20480` MiB 是半块 A100。`CUDA_DISABLE_CONTROL=true` 可避免 HAMi-core 因缺少 CUDA 驱动而中止。 +`gpu-memory` 注解使其成为共享 Pod,且不带 `nvidia.com/gpu` 请求——由 KAI 预留分数。`20480` MiB 是半块 A100。 ```bash kubectl apply -f - <<'EOF' @@ -261,9 +261,6 @@ spec: - name: gpu-workload image: busybox command: ["sleep", "3600"] - env: - - name: CUDA_DISABLE_CONTROL - value: "true" EOF kubectl logs -n kai-scheduler deploy/kai-scheduler-default --tail=40 \ diff --git a/tutorials/labs/kai-scheduler-hami.md b/tutorials/labs/kai-scheduler-hami.md index 92476725b..77fd328d1 100644 --- a/tutorials/labs/kai-scheduler-hami.md +++ b/tutorials/labs/kai-scheduler-hami.md @@ -189,7 +189,7 @@ apiVersion: node.k8s.io/v1 kind: RuntimeClass metadata: name: nvidia -handler: runc +handler: runc # mock-only — never apply this on a real GPU node EOF kubectl get pods -n kai-resource-isolator @@ -243,7 +243,7 @@ kai-whole-gpu 1/1 Running 0 10s 10.244.0.22 kai-hami-test-c ## Step 7: Schedule a Shared-GPU Pod and Inspect the Injection -A `gpu-memory` annotation makes this a shared Pod, with no `nvidia.com/gpu` request — KAI reserves the fraction. `20480` MiB is half an A100. `CUDA_DISABLE_CONTROL=true` keeps HAMi-core from aborting on the missing CUDA driver. +A `gpu-memory` annotation makes this a shared Pod, with no `nvidia.com/gpu` request — KAI reserves the fraction. `20480` MiB is half an A100. ```bash kubectl apply -f - <<'EOF' @@ -261,9 +261,6 @@ spec: - name: gpu-workload image: busybox command: ["sleep", "3600"] - env: - - name: CUDA_DISABLE_CONTROL - value: "true" EOF kubectl logs -n kai-scheduler deploy/kai-scheduler-default --tail=40 \