Skip to main content
Version: v2.9.0

Prerequisites

Before installing HAMi, make sure the following tools and dependencies are properly installed in your environment:

  • NVIDIA drivers >= 440
  • nvidia-docker version > 2.0
  • default runtime configured as nvidia for containerd/docker/cri-o container runtime
  • Kubernetes version >= 1.18
  • kernel version >= 3.10
  • helm > 3.0

Preparing your GPU Nodes

Execute the following steps on all your GPU nodes.

This guide assumes pre-installation of NVIDIA drivers and the nvidia-container-toolkit. Additionally, it assumes configuration of the nvidia-container-runtime as the default low-level runtime.

For details see Installing the NVIDIA Container Toolkit.

To inject NVIDIA GPUs through CDI, follow Enable NVIDIA CDI support for HAMi before installing HAMi. Verify the container runtime, driver root, and NVIDIA Container Toolkit path first.

Example for debian-based systems with Docker and containerd

Install the nvidia-container-toolkit

curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg \
&& curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list

sudo apt-get update && sudo apt-get install -y nvidia-container-toolkit

Configure Docker

When running Kubernetes with Docker, use the nvidia-ctk tool to automatically configure Docker:

sudo nvidia-ctk runtime configure --runtime=docker

Then restart Docker:

sudo systemctl daemon-reload && sudo systemctl restart docker

Configure containerd

When running Kubernetes with containerd, use the nvidia-ctk tool to automatically configure containerd:

sudo nvidia-ctk runtime configure --runtime=containerd

Then restart containerd:

sudo systemctl daemon-reload && sudo systemctl restart containerd

Label your nodes

Label your GPU nodes for scheduling with HAMi by adding the label "gpu=on". Without this label, the nodes cannot be managed by the HAMi scheduler.

kubectl label nodes <node-name> gpu=on
CNCFHAMi is a CNCF Incubating project