Build paddlepaddle-gpu from source for the NVIDIA GB10 Grace Blackwell platform
(DGX Spark and equivalents like the ASUS Ascent GX10): aarch64 / ARM64, CUDA 13,
compute capability sm_121.
There is no prebuilt paddlepaddle-gpu wheel for aarch64 + CUDA 13 + Blackwell on
PyPI, and upstream does not compile cleanly on this target. This repo has a working
build, a prebuilt wheel, and the exact patches it needs.
Verification after install:
python -c "import paddle; paddle.utils.run_check()"
# PaddlePaddle works well on 1 GPU.
| Device | NVIDIA GB10 (Grace Blackwell), 128 GB unified memory |
| Arch | aarch64 / sm_121 |
| OS | Ubuntu 24.04 LTS (aarch64) |
| CUDA | 13.0, driver 580.x |
| Python | 3.12 |
| PaddlePaddle | develop branch (3.5.0.dev) |
Grab the wheel from Releases:
pip install paddlepaddle_gpu-3.5.0.dev-cp312-cp312-linux_aarch64.whl
python -c "import paddle; paddle.utils.run_check()"
sudo apt-get update && sudo apt-get install -y build-essential cmake ninja-build git
bash build.sh
pip install build/Paddle/build/python/dist/paddlepaddle_gpu-*-cp312-cp312-linux_aarch64.whl
build.sh creates a venv, installs cuDNN and patchelf, clones Paddle, applies the
patches below, configures, and runs ninja. Expect 1 to 2 hours on the 20 core CPU.
Set WITH_ARM=ON together with WITH_GPU=ON. On GB10 this keeps GPU and sm_121 while
disabling the x86 assumptions that otherwise break the build (denormal.cc including
pmmintrin.h, the xbyak JIT, the -m64 compiler flag). Verified: WITH_GPU=ON and
WITH_ARM=ON coexist, and WITH_XBYAK auto disables. Set this first. It removes most
of the individual fixes; the CUTLASS, warp, and SLEEF fixes below still apply.
Without -DCUDA_ARCH_NAME=Manual, Paddle ignores CUDA_ARCH_BIN and auto detects,
producing the arch string 121(20) that nvcc rejects. Pass
-DCUDA_ARCH_NAME=Manual -DCUDA_ARCH_BIN=121 for a clean
-gencode arch=compute_121,code=sm_121.
generate_kernels.py and generate_variable_forward_kernels.py under
paddle/phi/kernels/fusion/cutlass/memory_efficient_attention run int("121(20)") and
DEFAULT_ARCH.index(121), which throw (their MAX_ARCH is 100). patch_arch_generators.py
rewrites convert_to_arch_list to parse the leading integer and clamp to 80. The fMHA
templates only exist up to sm_80 and run on sm_121 through CUDA forward compatibility.
These speech libraries use legacy FindCUDA and pass compute_20, removed in CUDA 13.
build.sh rewrites their CUDA_NVCC_FLAGS to -gencode arch=compute_90,code=sm_90.
cmake/flags.cmake adds -m64 when NOT WITH_ARM. Setting WITH_ARM=ON handles it.
WITH_ARM=ON pulls in SLEEF, which depends on tlfloat, but libtlfloat.a is never
installed or linked, so the final link fails with undefined reference to tlfloat_fma.
build.sh copies libtlfloat.a into SLEEF's install lib dir and attaches it to the
sleef imported target via INTERFACE_LINK_LIBRARIES.
patchelf must be on PATH. cuDNN dev headers come from pip nvidia-cudnn-cu13, passed
as CUDNN_ROOT, with a libcudnn.so -> libcudnn.so.9 symlink.
cmake -B build -GNinja \
-DWITH_GPU=ON -DWITH_ARM=ON \
-DCUDA_ARCH_NAME=Manual -DCUDA_ARCH_BIN=121 \
-DCMAKE_CUDA_FLAGS="-U__ARM_NEON -DEIGEN_DONT_VECTORIZE=1" \
-DPY_VERSION=3.12 -DPYTHON_EXECUTABLE="$VENV/bin/python" \
-DWITH_TESTING=OFF -DWITH_DISTRIBUTE=OFF -DWITH_NCCL=OFF \
-DCUDNN_ROOT="$CUDNN_DIR"
ninja -C build -j"$(nproc)"
sm_90 code in the warp libraries and the CUTLASS fMHA kernels runs on sm_121 through forward compatibility. The main graph compiles natively for sm_121. Tested with PaddleOCR-VL and other vision models on GPU.
Scripts and patches here are MIT. PaddlePaddle is Apache 2.0.