Skip to content

[Build] Build tilelang without host toolchain - #1833

Merged
LeiWang1999 merged 1 commit into
tile-ai:mainfrom
oraluben:copilot/add-build-option-without-nvcc
Feb 24, 2026
Merged

LeiWang1999 merged 1 commit into
tile-ai:mainfrom
oraluben:copilot/add-build-option-without-nvcc

Conversation

@oraluben

@oraluben oraluben commented Feb 11, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

This PR enables building tilelang with CUDA toolchain from pip (pip install "nvidia-cuda-nvcc>=13" "nvidia-cuda-cccl>=13" "nvidia-cuda-nvrtc>=13"), and would benefit some agentic workflow when e.g. user does not have permission to install cuda toolkit.

Summary by CodeRabbit

  • New Features

    • Added support for pip-provided CUDA toolchain alongside host CUDA installation
  • Documentation

    • Updated Python minimum version to >= 3.9
    • Enhanced CUDA setup documentation with new toolchain configuration workflows
  • Dependencies

    • Updated z3-solver, apache-tvm-ffi, and cython version constraints
    • Added platform-specific packages; refined development dependencies

@github-actions

Copy link
Copy Markdown

👋 Hi! Thank you for contributing to the TileLang project.

Please remember to run pre-commit run --all-files in the root directory of the project to ensure your changes are properly linted and formatted. This will help ensure your contribution passes the format check.

We appreciate you taking this step! Our team will review your contribution, and we look forward to your awesome work! 🚀

@coderabbitai

coderabbitai Bot commented Feb 11, 2026 •

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

This PR adds support for detecting and using pip-installed CUDA toolkits alongside traditional host installations. It introduces CMake utilities to dynamically locate CUDA, updates build requirements and documentation, removes obsolete libdevice lookup functions, and adjusts Python version constraints to >=3.9.

Changes

Cohort / File(s) Summary
CUDA Toolchain Detection
cmake/FindPipCUDAToolkit.cmake, cmake/find_pip_cuda.py, CMakeLists.txt, .gitignore
New CMake module and Python utility to detect CUDA via host installation, explicit pip-toolchain path, or Python environment; CMakeLists.txt integrates detection and extends link directories; .gitignore adds CMakeFiles/ pattern.
Installation Documentation
docs/get_started/Installation.md
Updated Python version prerequisite to >=3.9; broadened CUDA support from 12.0-13.0 to >=10.0 (host) or >=13.0 (pip); added workflows and WITH_PIP_CUDA_TOOLCHAIN environment variable documentation.
Dependency Management
requirements.txt, requirements-dev.txt, requirements-test.txt
Loosened apache-tvm-ffi lower bound (>=0.1.6 to >=0.1.2); tightened z3-solver upper bound (<4.15.5); raised cython to >=3.1.0; added torch-c-dlpack-ext with Python version marker; removed torch; added delocate for Darwin platform.
Code Cleanup
tilelang/contrib/nvcc.py
Removed find_libdevice_path() and callback_libdevice_path() helper functions and their TVM FFI registrations.

Sequence Diagram

sequenceDiagram
    participant CMake as CMake Build System
    participant FindPip as FindPipCUDAToolkit.cmake
    participant PythonScript as find_pip_cuda.py
    participant HostCUDA as Host CUDA Installation
    participant PipCUDA as Pip CUDA Package
    participant Compiler as Compiler Config

    CMake->>FindPip: Include module, detect CUDA
    FindPip->>HostCUDA: find_package(CUDAToolkit)
    alt Host CUDA Found
        HostCUDA-->>FindPip: Return toolkit paths
        FindPip->>Compiler: Set CMAKE_CUDA_COMPILER
    else Host CUDA Not Found
        FindPip->>PythonScript: Execute find_pip_cuda.py
        alt WITH_PIP_CUDA_TOOLCHAIN Set
            PythonScript->>PipCUDA: Use explicit path
        else
            PythonScript->>PipCUDA: Auto-detect from site-packages
        end
        PipCUDA-->>PythonScript: Return nvcc path & lib dirs
        PythonScript->>PythonScript: Ensure symlinks & stubs
        PythonScript-->>FindPip: Return JSON config
        FindPip->>Compiler: Set CMAKE_CUDA_COMPILER & link dirs
    end
    Compiler-->>CMake: CUDA toolchain configured
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • #1817: Synchronizes dependency updates (apache-tvm-ffi, z3-solver, torch-c-dlpack-ext) across multiple configuration files.
  • #1528: Implements complementary pip-installed CUDA detection via tilelang.env CUDA_HOME introspection.
  • #1373: Earlier modification of apache-tvm-ffi version constraints with overlapping dependency tuning.

Suggested labels

enhancement, dependencies

Suggested reviewers

  • XuehaiPan
  • LeiWang1999

Poem

🐰 A rabbit hops through CMake's maze,
Finding CUDA in pip's glaze,
Host or package, the choice is clear,
nvcc compiler appears right here!
No more lost libdevice ways,
Hopping forward through brighter days. ✨

🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 75.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately captures the main objective: enabling builds without requiring a host CUDA toolchain by supporting pip-provided CUDA packages.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Tip

Issue Planner is now in beta. Read the docs and try it out! Share your feedback on Discord.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@oraluben oraluben changed the title Build tilelang without host toolchain [Build] Build tilelang without host toolchain Feb 11, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Fix all issues with AI agents
In `@cmake/find_pip_cuda.py`:
- Around line 58-75: The _ensure_cuda_stub function currently swallows all
exceptions when trying to compile a libcuda.so stub which hides a missing gcc
and causes later cryptic link errors; change the error handling around
subprocess.check_call(["gcc", ...]) to specifically catch FileNotFoundError (or
OSError indicating missing executable) and emit a clear warning to stderr (or
via logging) that gcc is not available and the stub could not be built, while
allowing other exceptions (e.g., permission, disk full) to either be logged with
their details or re-raised so they aren't masked; keep cleanup of src in the
finally block and reference _ensure_cuda_stub, stubs_dir, stub, src, and
subprocess.check_call when making the change.

In `@docs/get_started/Installation.md`:
- Around line 8-9: Update the Installation.md CUDA requirement from ">= 10.0" to
">= 11.0" to reflect the compile-time assertion in src/target/stubs/cudart.cc
that requires CUDART_VERSION >= 11000; change the CUDA host installation line to
state "CUDA Version: >= 11.0 (host installation), or pip-provided CUDA toolchain
(>= 13.0)" so documentation matches the CUDART_VERSION check and prevents
build-time confusion.

In `@requirements-dev.txt`:
- Line 6: Confirm whether cython_wrapper.pyx actually requires Cython >=3.1.0
(check for any 3.1+ specific syntax) and then make the requirement consistent:
either lower the constraint in requirements-dev.txt and pyproject.toml to match
the bare "cython" used in requirements-test.txt, or keep >=3.1.0 and add the
same >=3.1.0 constraint to requirements-test.txt; reference cython_wrapper.pyx,
the language_level=3 directive, requirements-dev.txt, requirements-test.txt, and
pyproject.toml when making the update.
🧹 Nitpick comments (2)
cmake/FindPipCUDAToolkit.cmake (1)

23-26: find_program may locate a different Python than the one with pip CUDA packages.

find_program(... NAMES python3 python) searches the system PATH, which in isolated build environments (the default for pip install) may resolve to a different interpreter than the one that has the nvidia packages installed. This is mitigated by the doc guidance to use --no-build-isolation for the auto-detect path, but it might be worth adding a comment clarifying this assumption — or preferring ${Python_EXECUTABLE} if available (though find_package(Python) hasn't run yet at this point since it's before project()).

CMakeLists.txt (1)

337-337: Consider using target_link_directories instead of link_directories.

link_directories is directory-scoped and affects all subsequently defined targets, which can cause unintended side effects. Modern CMake (3.13+) prefers target_link_directories for more precise scoping. That said, since this is within the USE_CUDA block and the targets that need it are defined shortly after, the practical impact is minimal.

Suggested change
-  link_directories(${CUDAToolkit_LIBRARY_DIR} ${CUDAToolkit_LIBRARY_DIR}/stubs)
+  # Ensure pip-installed CUDA stubs are discoverable at link time.
+  set(TILELANG_CUDA_LINK_DIRS ${CUDAToolkit_LIBRARY_DIR} ${CUDAToolkit_LIBRARY_DIR}/stubs)

Then apply per-target:

target_link_directories(tilelang_objs PRIVATE ${TILELANG_CUDA_LINK_DIRS})

Comment thread cmake/find_pip_cuda.py
Comment thread docs/get_started/Installation.md
Comment thread requirements-dev.txt
@LeiWang1999
LeiWang1999 merged commit a61da71 into tile-ai:main Feb 24, 2026
17 of 18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants