Repository navigation
Support on-disk caching in the CUDA target - #8089
Conversation
|
gpuci run tests |
3 similar comments
|
gpuci run tests |
|
gpuci run tests |
|
gpuci run tests |
This primarily uses Numba's existing infrastructure for on-disk, caching, and only requires filling in / connecting up some small extra parts in the CUDA target. These changes are: - Fix `_reduce_states()` and `_rebuild()` ion `CUDACodeLibrary`. These were never used before, so some simple logic errors persisted. - Add a `magic_tuple()` function to the `JITCUDACodegen`. This is used to uniquely describe the codegen behaviour, which for CUDA is runtime version- and compute capability-specific. - Wire up the `cache` kwarg in the `@cuda.jit` decorator to enable caching. - Add some properties to CUDA `_Kernel` objects that the caching implementation expects to find, and modify its `_reduce_states()` and `_rebuild()` methods to store the signature instead of the list of argument types (which aligns it more closely with the design of a `CompileResult`). - Implement a `CUDACache` class for the core caching implementation to use for CUDA Dispatchers. - Interpose use of the cache in the `CUDADispatcher.compile()` function, along with some small refactoring to create the `add_overload()` function.
These tests consist of existing cache tests for the CPU dispatcher, modified for the CUDA target. These tests are based on the `TestCache` and `TestMultiProcessCache` tests in `numba.tests.test_caching`. Any tests that were orthogonal to the CUDA target (e.g. IPython use cases) or irrelevant to its functionality (e.g. object mode) have been omitted.
There are two kinds of tests: - Those where a CPU and CUDA dispatcher is used for the same function. - Those where caching is used with multiple compute capabilities.
523bda8 to
1d29426
Compare
|
gpuci run tests |
|
@stuartarchibald This is now ready for review. |
stuartarchibald
left a comment
There was a problem hiding this comment.
Thanks for the patch. Great to see this feature extended to another hardware target. I've reviewed the patch and don't have many comments to make on the basis that:
- @gmarkall and I have spent time OOB discussing and working out various aspects of this PR's design prior to implementation.
- The tests are a reasonable facsimile of the CPU cache tests and give confidence that the patch is working as expected.
- The CUDA refactoring effort to reuse as much of the CPU target as possible has lead to the proposed changes being small (which is good!)
I'll give this a manual test once reviews are complete. Thanks again!
| # The following are referred to by the cache implementation. Note: | ||
| # - There are no referenced environments in CUDA. | ||
| # - Kernels don't have lifted code. | ||
| # - reload_init is only for parfors. |
There was a problem hiding this comment.
Currently it's only used by parfors, but the use extends to anything which needs some sort of "state" reinitialising.
Co-authored-by: stuartarchibald <stuartarchibald@users.noreply.github.com>
|
gpuci run tests |
|
@stuartarchibald Many thanks for the review - I've applied the two suggestions and responded to the other comments. |
stuartarchibald
left a comment
There was a problem hiding this comment.
Thanks for the fixes. Patch approved pending manual testing.
sklam
left a comment
There was a problem hiding this comment.
I did manual testing on a multi-GPUs (CC 6.0, CC 3.5) machine and can confirm that the cache is working as expected.
Thanks @sklam |
|
Buildfarm ID: |
|
@gmarkall, there's now a small merge conflict |
|
gpuci run tests |
stuartarchibald
left a comment
There was a problem hiding this comment.
Merge commit looks good, thanks @gmarkall.
This adds support for on-disk caching of CUDA kernels, reusing Numba's core caching infrastructure; only some small changes were needed to "hook up" the core caching implementation to the CUDA target.
Tests are added based on the CPU dispatcher caching tests, along with some new CUDA-specific tests that check behaviour when cached functions are both CUDA- and CPU-jitted, and when a function is CUDA_jitted for multiple compute capabilities.
See individual commit messages for further details.