To get much faster NCCL comms at small-to-medium payloads, start using symmetric memory, which has been added recently to NCCL and PyTorch.
If you struggle to overlap comms and compute this could be your salvation in a few lines of code!
Details github.com/stas00/ml-engi…
Toolmaker. Software creator, optimizer and harmonizer.
Makes ML systems work and fly @ Snowflake.

