Skip to content

Support on-disk caching in the CUDA target - #8089

Merged
sklam merged 6 commits into
numba:mainfrom
gmarkall:cuda-caching-2-rebase
Jun 17, 2022
Merged

sklam merged 6 commits into
numba:mainfrom
gmarkall:cuda-caching-2-rebase

Conversation

@gmarkall

@gmarkall gmarkall commented May 22, 2022 •

Copy link
Copy Markdown
Member

This adds support for on-disk caching of CUDA kernels, reusing Numba's core caching infrastructure; only some small changes were needed to "hook up" the core caching implementation to the CUDA target.

Tests are added based on the CPU dispatcher caching tests, along with some new CUDA-specific tests that check behaviour when cached functions are both CUDA- and CPU-jitted, and when a function is CUDA_jitted for multiple compute capabilities.

See individual commit messages for further details.

@gmarkall gmarkall added the CUDA CUDA related issue/PR label May 24, 2022
@gmarkall

Copy link
Copy Markdown
Member Author

gpuci run tests

3 similar comments
@gmarkall

Copy link
Copy Markdown
Member Author

gpuci run tests

@gmarkall

Copy link
Copy Markdown
Member Author

gpuci run tests

@gmarkall

Copy link
Copy Markdown
Member Author

gpuci run tests

gmarkall added 4 commits May 30, 2022 15:44
This primarily uses Numba's existing infrastructure for on-disk,
caching, and only requires filling in / connecting up some small extra
parts in the CUDA target. These changes are:

- Fix `_reduce_states()` and `_rebuild()` ion `CUDACodeLibrary`. These
  were never used before, so some simple logic errors persisted.
- Add a `magic_tuple()` function to the `JITCUDACodegen`. This is used
  to uniquely describe the codegen behaviour, which for CUDA is runtime
  version- and compute capability-specific.
- Wire up the `cache` kwarg in the `@cuda.jit` decorator to enable
  caching.
- Add some properties to CUDA `_Kernel` objects that the caching
  implementation expects to find, and modify its `_reduce_states()` and
  `_rebuild()` methods to store the signature instead of the list of
  argument types (which aligns it more closely with the design of a
  `CompileResult`).
- Implement a `CUDACache` class for the core caching implementation to
  use for CUDA Dispatchers.
- Interpose use of the cache in the `CUDADispatcher.compile()` function,
  along with some small refactoring to create the `add_overload()`
  function.
These tests consist of existing cache tests for the CPU dispatcher,
modified for the CUDA target. These tests are based on the `TestCache`
and `TestMultiProcessCache` tests in `numba.tests.test_caching`. Any
tests that were orthogonal to the CUDA target (e.g. IPython use cases)
or irrelevant to its functionality (e.g. object mode) have been omitted.
There are two kinds of tests:

- Those where a CPU and CUDA dispatcher is used for the same function.
- Those where caching is used with multiple compute capabilities.
@gmarkall gmarkall changed the title [WIP] CUDA caching Support on-disk caching for CUDA May 30, 2022
@gmarkall gmarkall changed the title Support on-disk caching for CUDA Support on-disk caching in the CUDA target May 30, 2022
@gmarkall gmarkall added this to the Numba 0.56 RC milestone May 30, 2022
@gmarkall
gmarkall force-pushed the cuda-caching-2-rebase branch from 523bda8 to 1d29426 Compare May 30, 2022 15:14
@gmarkall

Copy link
Copy Markdown
Member Author

gpuci run tests

@gmarkall

Copy link
Copy Markdown
Member Author

@stuartarchibald This is now ready for review.

@stuartarchibald stuartarchibald added the Effort - medium Medium size effort needed label May 30, 2022

@stuartarchibald stuartarchibald left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the patch. Great to see this feature extended to another hardware target. I've reviewed the patch and don't have many comments to make on the basis that:

  • @gmarkall and I have spent time OOB discussing and working out various aspects of this PR's design prior to implementation.
  • The tests are a reasonable facsimile of the CPU cache tests and give confidence that the patch is working as expected.
  • The CUDA refactoring effort to reuse as much of the CPU target as possible has lead to the proposed changes being small (which is good!)

I'll give this a manual test once reviews are complete. Thanks again!

Comment thread numba/cuda/dispatcher.py
# The following are referred to by the cache implementation. Note:
# - There are no referenced environments in CUDA.
# - Kernels don't have lifted code.
# - reload_init is only for parfors.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Currently it's only used by parfors, but the use extends to anything which needs some sort of "state" reinitialising.

Comment thread numba/cuda/tests/cudapy/cache_usecases.py Outdated
Comment thread numba/cuda/tests/cudapy/cache_usecases.py Outdated
Comment thread numba/cuda/tests/cudapy/test_caching.py
Comment thread numba/cuda/tests/cudapy/cache_usecases.py
@stuartarchibald stuartarchibald added 4 - Waiting on author Waiting for author to respond to review and removed 3 - Ready for Review labels Jun 6, 2022
Co-authored-by: stuartarchibald <stuartarchibald@users.noreply.github.com>
@gmarkall

gmarkall commented Jun 7, 2022

Copy link
Copy Markdown
Member Author

gpuci run tests

@gmarkall

gmarkall commented Jun 7, 2022

Copy link
Copy Markdown
Member Author

@stuartarchibald Many thanks for the review - I've applied the two suggestions and responded to the other comments.

@stuartarchibald stuartarchibald added 4 - Waiting on reviewer Waiting for reviewer to respond to author and removed 4 - Waiting on author Waiting for author to respond to review labels Jun 7, 2022

@stuartarchibald stuartarchibald left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the fixes. Patch approved pending manual testing.

sklam
sklam previously approved these changes Jun 14, 2022

@sklam sklam left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did manual testing on a multi-GPUs (CC 6.0, CC 3.5) machine and can confirm that the cache is working as expected.

@sklam sklam added 5 - Ready to merge Review and testing done, is ready to merge and removed 4 - Waiting on reviewer Waiting for reviewer to respond to author labels Jun 14, 2022
@stuartarchibald

Copy link
Copy Markdown
Contributor

I did manual testing on a multi-GPUs (CC 6.0, CC 3.5) machine and can confirm that the cache is working as expected.

Thanks @sklam

@stuartarchibald stuartarchibald added Pending BuildFarm For PRs that have been reviewed but pending a push through our buildfarm 4 - Waiting on CI Review etc done, waiting for CI to finish and removed 5 - Ready to merge Review and testing done, is ready to merge labels Jun 15, 2022
@stuartarchibald

Copy link
Copy Markdown
Contributor

Buildfarm ID: numba_smoketest_cuda_yaml_135.

@sklam sklam added BuildFarm Passed For PRs that have been through the buildfarm and passed 5 - Ready to merge Review and testing done, is ready to merge and removed Pending BuildFarm For PRs that have been reviewed but pending a push through our buildfarm 4 - Waiting on CI Review etc done, waiting for CI to finish labels Jun 15, 2022
@sklam

sklam commented Jun 15, 2022

Copy link
Copy Markdown
Member

@gmarkall, there's now a small merge conflict

@gmarkall
gmarkall dismissed stale reviews from sklam and stuartarchibald via 944a531 June 16, 2022 11:46
@gmarkall

Copy link
Copy Markdown
Member Author

gpuci run tests

@stuartarchibald stuartarchibald left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Merge commit looks good, thanks @gmarkall.

@sklam
sklam merged commit 7147ce1 into numba:main Jun 17, 2022
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

5 - Ready to merge Review and testing done, is ready to merge BuildFarm Passed For PRs that have been through the buildfarm and passed CUDA CUDA related issue/PR Effort - medium Medium size effort needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants