AWS AWSEC2: Add Amazon Linux 2023 and update NVIDIA/CUDA install steps
Summary
Adds Amazon Linux 2023 instructions and updates NVIDIA driver, CUDA, cuDNN, GDRCopy, and NIXL setup, including replacing apt-key with cuda-keyring.
Security assessment
The diff replaces deprecated apt-key adv commands with NVIDIA cuda-keyring packages and HTTPS repo URLs, improving package repository key management; no specific CVE or incident is referenced.
Evidence
+ $ wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb \
Diff
diff --git a/AWSEC2/latest/UserGuide/efa-start-nixl.md b/AWSEC2/latest/UserGuide/efa-start-nixl.md index 4cbe29313..33014fc6f 100644 --- a//AWSEC2/latest/UserGuide/efa-start-nixl.md +++ b//AWSEC2/latest/UserGuide/efa-start-nixl.md @@ -13 +13 @@ The NVIDIA Inference Xfer Library (NIXL) is a high-throughput, low-latency commu - * Only Ubuntu 24.04 and Ubuntu 22.04 base AMIs are supported. + * Supported base AMIs: Amazon Linux 2023, Ubuntu 24.04, and Ubuntu 22.04. @@ -127 +127,83 @@ Skip Step 3 if your AMI already includes Nvidia GPU drivers, the CUDA toolkit, a -###### To install the Nvidia GPU drivers, Nvidia CUDA toolkit, and cuDNN +Amazon Linux 2023 + + +###### To install the NVIDIA GPU drivers, NVIDIA CUDA Toolkit, and cuDNN + + 1. To ensure that all of your software packages are up to date, perform a quick software update on your instance. + + $ sudo dnf upgrade -y && sudo reboot + +After the instance has rebooted, reconnect to it. + + 2. Install the utilities that are needed to install the NVIDIA GPU drivers and the NVIDIA CUDA Toolkit. + + $ sudo dnf groupinstall 'Development Tools' -y && sudo dnf install -y dkms kernel-devel-$(uname -r) kernel-headers-$(uname -r) + + 3. Disable the `nouveau` open source drivers. + + 1. Install the required utilities and the kernel headers package for the version of the kernel that you are currently running. + + $ sudo yum install -y wget kernel-devel-$(uname -r) kernel-headers-$(uname -r) + + 2. Add `nouveau` to the `/etc/modprobe.d/blacklist.conf` deny list file. + + $ cat << EOF | sudo tee --append /etc/modprobe.d/blacklist.conf + blacklist vga16fb + blacklist nouveau + blacklist rivafb + blacklist nvidiafb + blacklist rivatv + EOF + + 3. Append `GRUB_CMDLINE_LINUX="rdblacklist=nouveau"` to the `grub` file and rebuild the GRUB configuration. + + $ echo 'GRUB_CMDLINE_LINUX="rdblacklist=nouveau"' | sudo tee -a /etc/default/grub \ + && sudo grub2-mkconfig -o /boot/grub2/grub.cfg + + 4. Reboot the instance and reconnect to it. + + 5. Add the CUDA network repository. + + $ sudo yum-config-manager --add-repo https://developer.download.nvidia.com/compute/cuda/repos/rhel8/x86_64/cuda-rhel8.repo + + 6. Download and install the NVIDIA GPU driver. + + $ wget https://us.download.nvidia.com/tesla/580.167.08/NVIDIA-Linux-x86_64-580.167.08.run \ + && sudo sh NVIDIA-Linux-x86_64-580.167.08.run -m kernel-open --no-drm --disable-nouveau --dkms --silent + + 7. Install the NVIDIA CUDA Toolkit and cuDNN. + + $ sudo dnf install -y cuda-toolkit-13-0 libcudnn9-cuda-13 libcudnn9-devel-cuda-13 + + 8. Reboot the instance and reconnect to it. + + 9. Install and start the NVIDIA Fabric Manager on instances with NVSwitch, such as P-series multi-GPU instances. G-series instances do not use NVSwitch and do not require Fabric Manager. + + $ sudo dnf install -y https://developer.download.nvidia.com/compute/cuda/repos/rhel8/x86_64/nvidia-fabricmanager-580.167.08-1.el8.x86_64.rpm \ + && sudo systemctl enable nvidia-fabricmanager && sudo systemctl start nvidia-fabricmanager + + 10. Ensure that the CUDA paths are set each time that the instance starts. + + * For _bash_ shells, add the following statements to `/home/`username`/.bashrc` and `/home/`username`/.bash_profile`. + + export PATH=/usr/local/cuda/bin:$PATH + export LD_LIBRARY_PATH=/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64:$LD_LIBRARY_PATH + + * For _tcsh_ shells, add the following statements to `/home/`username`/.cshrc`. + + setenv PATH /usr/local/cuda/bin:$PATH + setenv LD_LIBRARY_PATH /usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64:$LD_LIBRARY_PATH + + 11. To confirm that the NVIDIA GPU drivers are functional, run the following command. + + $ nvidia-smi -q | head + +The command should return information about the NVIDIA GPUs, NVIDIA GPU drivers, and NVIDIA CUDA Toolkit. + + + + +Ubuntu 24.04 and Ubuntu 22.04 + + +###### To install the NVIDIA GPU drivers, NVIDIA CUDA Toolkit, and cuDNN @@ -133 +215 @@ Skip Step 3 if your AMI already includes Nvidia GPU drivers, the CUDA toolkit, a - 2. Install the utilities that are needed to install the Nvidia GPU drivers and the Nvidia CUDA toolkit. + 2. Install the utilities that are needed to install the NVIDIA GPU drivers and the NVIDIA CUDA Toolkit. @@ -135 +217 @@ Skip Step 3 if your AMI already includes Nvidia GPU drivers, the CUDA toolkit, a - $ sudo apt-get install build-essential -y + $ sudo apt-get update && sudo apt-get install build-essential -y @@ -137 +219 @@ Skip Step 3 if your AMI already includes Nvidia GPU drivers, the CUDA toolkit, a - 3. To use the Nvidia GPU driver, you must first disable the `nouveau` open source drivers. + 3. To use the NVIDIA GPU driver, you must first disable the `nouveau` open source drivers. @@ -157 +239 @@ Skip Step 3 if your AMI already includes Nvidia GPU drivers, the CUDA toolkit, a - 4. Rebuild the Grub configuration. + 4. Rebuild the GRUB configuration. @@ -163,16 +245 @@ Skip Step 3 if your AMI already includes Nvidia GPU drivers, the CUDA toolkit, a - 5. Add the CUDA repository and install the Nvidia GPU drivers, NVIDIA CUDA toolkit, and cuDNN. - - $ sudo apt-key adv --fetch-keys http://developer.download.nvidia.com/compute/machine-learning/repos/ubuntu2004/x86_64/7fa2af80.pub \ - && wget -O /tmp/deeplearning.deb http://developer.download.nvidia.com/compute/machine-learning/repos/ubuntu2004/x86_64/nvidia-machine-learning-repo-ubuntu2004_1.0.0-1_amd64.deb \ - && sudo dpkg -i /tmp/deeplearning.deb \ - && wget -O /tmp/cuda.pin https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/cuda-ubuntu2004.pin \ - && sudo mv /tmp/cuda.pin /etc/apt/preferences.d/cuda-repository-pin-600 \ - && sudo apt-key adv --fetch-keys https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/3bf863cc.pub \ - && sudo add-apt-repository 'deb http://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/ /' \ - && sudo apt update \ - && sudo apt install nvidia-dkms-535 \ - && sudo apt install -o Dpkg::Options::='--force-overwrite' cuda-drivers-535 cuda-toolkit-12-3 libcudnn8 libcudnn8-dev -y - - 6. Reboot the instance and reconnect to it. - - 7. (`p4d.24xlarge` and `p5.48xlarge` only) Install the Nvidia Fabric Manager. + 5. Add the CUDA repository and install the NVIDIA GPU drivers, NVIDIA CUDA Toolkit, and cuDNN. @@ -180 +247 @@ Skip Step 3 if your AMI already includes Nvidia GPU drivers, the CUDA toolkit, a - 1. You must install the version of the Nvidia Fabric Manager that matches the version of the Nvidia kernel module that you installed in the previous step. + * Ubuntu 24.04 @@ -182 +249,4 @@ Skip Step 3 if your AMI already includes Nvidia GPU drivers, the CUDA toolkit, a -Run the following command to determine the version of the Nvidia kernel module. + $ wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb \ + && sudo dpkg -i cuda-keyring_1.1-1_all.deb \ + && sudo add-apt-repository -y 'deb https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64 /' \ + && sudo apt-get update @@ -184 +254 @@ Run the following command to determine the version of the Nvidia kernel module. - $ cat /proc/driver/nvidia/version | grep "Kernel Module" + $ sudo apt-get install -y nvidia-open-580 cuda-toolkit-13-0 libcudnn9-cuda-13 libcudnn9-dev-cuda-13 @@ -186 +256 @@ Run the following command to determine the version of the Nvidia kernel module. -The following is example output. + * Ubuntu 22.04 @@ -188 +258,4 @@ The following is example output. - NVRM version: NVIDIA UNIX x86_64 Kernel Module **450**.42.01 Tue Jun 15 21:26:37 UTC 2021 + $ wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.1-1_all.deb \ + && sudo dpkg -i cuda-keyring_1.1-1_all.deb \ + && sudo DEBIAN_FRONTEND=noninteractive add-apt-repository -y "deb https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/ /" \ + && sudo apt-get update @@ -190 +263 @@ The following is example output. -In the example above, major version `450` of the kernel module was installed. This means that you need to install Nvidia Fabric Manager version `450`. + $ sudo apt-get install -y nvidia-open-580 cuda-toolkit-13-0 libcudnn9-cuda-13 libcudnn9-dev-cuda-13 @@ -192,5 +265 @@ In the example above, major version `450` of the kernel module was installed. Th - 2. Install the Nvidia Fabric Manager. Run the following command and specify the major version identified in the previous step. - - $ sudo apt install -o Dpkg::Options::='--force-overwrite' nvidia-fabricmanager-major_version_number - -For example, if major version `450` of the kernel module was installed, use the following command to install the matching version of Nvidia Fabric Manager. + 6. Reboot the instance and reconnect to it. @@ -198 +267 @@ For example, if major version `450` of the kernel module was installed, use the - $ sudo apt install -o Dpkg::Options::='--force-overwrite' nvidia-fabricmanager-450 + 7. Install and start the NVIDIA Fabric Manager on instances with NVSwitch, such as P-series multi-GPU instances. G-series instances do not use NVSwitch and do not require Fabric Manager. @@ -200 +269 @@ For example, if major version `450` of the kernel module was installed, use the - 3. Start the service, and ensure that it starts automatically when the instance starts. Nvidia Fabric Manager is required for NV Switch Management. + * Ubuntu 24.04 and Ubuntu 22.04 @@ -202 +271,2 @@ For example, if major version `450` of the kernel module was installed, use the - $ sudo systemctl start nvidia-fabricmanager && sudo systemctl enable nvidia-fabricmanager + $ sudo apt-get install -y nvidia-fabricmanager-580 \ + && sudo systemctl enable nvidia-fabricmanager && sudo systemctl start nvidia-fabricmanager @@ -213,2 +283,2 @@ For example, if major version `450` of the kernel module was installed, use the - setenv PATH=/usr/local/cuda/bin:$PATH - setenv LD_LIBRARY_PATH=/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64:$LD_LIBRARY_PATH + setenv PATH /usr/local/cuda/bin:$PATH + setenv LD_LIBRARY_PATH /usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64:$LD_LIBRARY_PATH @@ -216 +286 @@ For example, if major version `450` of the kernel module was installed, use the - 9. To confirm that the Nvidia GPU drivers are functional, run the following command. + 9. To confirm that the NVIDIA GPU drivers are functional, run the following command. @@ -220 +290 @@ For example, if major version `450` of the kernel module was installed, use the -The command should return information about the Nvidia GPUs, Nvidia GPU drivers, and Nvidia CUDA toolkit. +The command should return information about the NVIDIA GPUs, NVIDIA GPU drivers, and NVIDIA CUDA Toolkit. @@ -230,0 +301,3 @@ Install GDRCopy to improve the performance of Libfabric on GPU-based platforms. +Amazon Linux 2023 + + @@ -235 +308 @@ Install GDRCopy to improve the performance of Libfabric on GPU-based platforms. - $ sudo apt -y install build-essential devscripts debhelper check libsubunit-dev fakeroot pkg-config dkms + $ sudo yum -y install dkms rpm-build make check check-devel @@ -239,3 +312,8 @@ Install GDRCopy to improve the performance of Libfabric on GPU-based platforms. - $ wget https://github.com/NVIDIA/gdrcopy/archive/refs/tags/v2.4.tar.gz \ - && tar xf v2.4.tar.gz \ - && cd gdrcopy-2.4/packages + $ wget https://github.com/NVIDIA/gdrcopy/archive/refs/tags/v2.5.2.tar.gz \ + && tar xf v2.5.2.tar.gz && cd gdrcopy-2.5.2/packages + + 3. Build the GDRCopy RPM packages. + + $ CUDA=/usr/local/cuda ./build-rpm-packages.sh + + 4. Install the GDRCopy RPM packages. @@ -243 +321,3 @@ Install GDRCopy to improve the performance of Libfabric on GPU-based platforms. - 3. Build the GDRCopy DEB packages. + $ sudo rpm -Uvh gdrcopy-kmod-2.5.2*dkms*.rpm \ + && sudo rpm -Uvh gdrcopy-2.5.2*.rpm \ + && sudo rpm -Uvh gdrcopy-devel-2.5.2*.rpm @@ -245 +324,0 @@ Install GDRCopy to improve the performance of Libfabric on GPU-based platforms. - $ CUDA=/usr/local/cuda ./build-deb-packages.sh