Skip to content

About

Build paddlepaddle-gpu from source for NVIDIA DGX Spark and GB10 (Grace Blackwell, aarch64, CUDA 13, sm_121). Prebuilt wheel plus the exact patches upstream is missing.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

PaddlePaddle GPU on NVIDIA DGX Spark / GB10 (aarch64, CUDA 13, sm_121)

Build paddlepaddle-gpu from source for the NVIDIA GB10 Grace Blackwell platform (DGX Spark and equivalents like the ASUS Ascent GX10): aarch64 / ARM64, CUDA 13, compute capability sm_121.

There is no prebuilt paddlepaddle-gpu wheel for aarch64 + CUDA 13 + Blackwell on PyPI, and upstream does not compile cleanly on this target. This repo has a working build, a prebuilt wheel, and the exact patches it needs.

Verification after install:

python -c "import paddle; paddle.utils.run_check()"
# PaddlePaddle works well on 1 GPU.

Target

Device NVIDIA GB10 (Grace Blackwell), 128 GB unified memory
Arch aarch64 / sm_121
OS Ubuntu 24.04 LTS (aarch64)
CUDA 13.0, driver 580.x
Python 3.12
PaddlePaddle develop branch (3.5.0.dev)

Install the prebuilt wheel

Grab the wheel from Releases:

pip install paddlepaddle_gpu-3.5.0.dev-cp312-cp312-linux_aarch64.whl
python -c "import paddle; paddle.utils.run_check()"

Build from source

sudo apt-get update && sudo apt-get install -y build-essential cmake ninja-build git
bash build.sh
pip install build/Paddle/build/python/dist/paddlepaddle_gpu-*-cp312-cp312-linux_aarch64.whl

build.sh creates a venv, installs cuDNN and patchelf, clones Paddle, applies the patches below, configures, and runs ninja. Expect 1 to 2 hours on the 20 core CPU.

Why upstream does not build on GB10, and the fixes

WITH_ARM=ON, the master switch

Set WITH_ARM=ON together with WITH_GPU=ON. On GB10 this keeps GPU and sm_121 while disabling the x86 assumptions that otherwise break the build (denormal.cc including pmmintrin.h, the xbyak JIT, the -m64 compiler flag). Verified: WITH_GPU=ON and WITH_ARM=ON coexist, and WITH_XBYAK auto disables. Set this first. It removes most of the individual fixes; the CUTLASS, warp, and SLEEF fixes below still apply.

Manual arch, or you get the garbage string 121(20)

Without -DCUDA_ARCH_NAME=Manual, Paddle ignores CUDA_ARCH_BIN and auto detects, producing the arch string 121(20) that nvcc rejects. Pass -DCUDA_ARCH_NAME=Manual -DCUDA_ARCH_BIN=121 for a clean -gencode arch=compute_121,code=sm_121.

CUTLASS attention generators crash on sm_121

generate_kernels.py and generate_variable_forward_kernels.py under paddle/phi/kernels/fusion/cutlass/memory_efficient_attention run int("121(20)") and DEFAULT_ARCH.index(121), which throw (their MAX_ARCH is 100). patch_arch_generators.py rewrites convert_to_arch_list to parse the leading integer and clamp to 80. The fMHA templates only exist up to sm_80 and run on sm_121 through CUDA forward compatibility.

warpctc and warprnnt emit compute_20

These speech libraries use legacy FindCUDA and pass compute_20, removed in CUDA 13. build.sh rewrites their CUDA_NVCC_FLAGS to -gencode arch=compute_90,code=sm_90.

-m64 breaks the aarch64 compiler

cmake/flags.cmake adds -m64 when NOT WITH_ARM. Setting WITH_ARM=ON handles it.

SLEEF tlfloat link error

WITH_ARM=ON pulls in SLEEF, which depends on tlfloat, but libtlfloat.a is never installed or linked, so the final link fails with undefined reference to tlfloat_fma. build.sh copies libtlfloat.a into SLEEF's install lib dir and attaches it to the sleef imported target via INTERFACE_LINK_LIBRARIES.

patchelf and cuDNN

patchelf must be on PATH. cuDNN dev headers come from pip nvidia-cudnn-cu13, passed as CUDNN_ROOT, with a libcudnn.so -> libcudnn.so.9 symlink.

Final cmake configuration

cmake -B build -GNinja \
  -DWITH_GPU=ON -DWITH_ARM=ON \
  -DCUDA_ARCH_NAME=Manual -DCUDA_ARCH_BIN=121 \
  -DCMAKE_CUDA_FLAGS="-U__ARM_NEON -DEIGEN_DONT_VECTORIZE=1" \
  -DPY_VERSION=3.12 -DPYTHON_EXECUTABLE="$VENV/bin/python" \
  -DWITH_TESTING=OFF -DWITH_DISTRIBUTE=OFF -DWITH_NCCL=OFF \
  -DCUDNN_ROOT="$CUDNN_DIR"
ninja -C build -j"$(nproc)"

Notes

sm_90 code in the warp libraries and the CUTLASS fMHA kernels runs on sm_121 through forward compatibility. The main graph compiles natively for sm_121. Tested with PaddleOCR-VL and other vision models on GPU.

License

Scripts and patches here are MIT. PaddlePaddle is Apache 2.0.

About

Build paddlepaddle-gpu from source for NVIDIA DGX Spark and GB10 (Grace Blackwell, aarch64, CUDA 13, sm_121). Prebuilt wheel plus the exact patches upstream is missing.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages