Skip to content

CI: run the CUDA unit tests on a simulated GPU - #817

Open
saqibkh wants to merge 1 commit into
taskflow:masterfrom
saqibkh:ci-simulated-gpu
Open

saqibkh wants to merge 1 commit into
taskflow:masterfrom
saqibkh:ci-simulated-gpu

Conversation

@saqibkh

@saqibkh saqibkh commented Oct 2, 2026

Copy link
Copy Markdown

Every job builds with TF_BUILD_CUDA off, because the runners have no GPU, so the cudaGraph unit tests (unittests/cuda) never run in CI. This adds one job to ubuntu.yml that builds them and runs them on a simulated NVIDIA T4.

How: PantheonSim, via pantheongpu/setup-pantheonsim, simulates the GPU on the runner's CPU. The kernels run from their PTX and the CUDA Graph calls go through the simulator's runtime, so the results are real. There is no timing model, so this checks correctness, not performance.

Changes: one new job, cuda-test-simulated-gpu, in the same style as the others:

  • It configures with -DTF_BUILD_CUDA=ON -DCMAKE_CUDA_RUNTIME_LIBRARY=Shared, with examples and benchmarks off. The shared runtime is because the simulator stands in for it.
  • It builds only the seven CUDA test targets unittests/cuda/CMakeLists.txt enables.
  • It runs ctest -E NOT_BUILT, which skips the CPU tests this job doesn't build; the existing jobs already run those.

Nothing else changes.

What it can touch: the action is pinned by commit (v0.1.2), not by a tag, and it pins the simulator it builds by commit too, so nothing in the job changes until you move the pin. The new job's token is read-only (permissions: contents: read), and the action uses no token or secrets; it installs packages only from Ubuntu's and NVIDIA's repositories. The job isn't a required check, so a red run never blocks a merge unless you decide it should.

What it gives, from a run on my fork:

  • All 75 CUDA test cases pass (objects, basics, updates, matrix, kmeans, for_each, transform), in about 12 seconds.
  • The job takes about 3.5 minutes. The action caches the simulator it builds, per pinned version, so only the first run in the repository builds it (about 19 minutes that time).

Disclosure: I maintain PantheonSim. Running these tests found a gap in it that's now fixed: a graph's memset node on cudaMallocManaged memory. If the job is ever flaky or wrong, please open an issue at pantheongpu/pantheonsim and I'll fix it on our side.

🤖 Generated with Claude Code

TF_BUILD_CUDA is off in every job, since the runners have no GPU, so
the cudaGraph tests never run. PantheonSim simulates a GPU on the
runner's CPU; this job builds the CUDA unit tests and runs them there.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant