Speaker
Description
GPU programming in the legacy gr-cuda custom buffer OOT meant writing C++, plus synchronization code in every block's work(). An upgrade to gr-cuda now lets you write ordinary Python blocks that call cp.* instead of np.*, with buffers handed to them as CuPy arrays already resident in device memory, with no data copies, and nothing for the user to manage. This talk shows how the gr-cuda API got that small: synchronization moved out of blocks and into the buffer using CUDA events, and a double-mapped virtual memory ring buffer along the lines of GNU Radio's CPU version. The result is Python CuPy blocks that benchmark within 5% of hand-written C++ CUDA blocks, and Host <-> Device data transfer rates that saturate PCIe bandwidths using GNURadio 3.
https://github.com/cascade-space-co/gr-cuda