Back
DOCA GPUnetIO Unifies GPU Initiated Networking Across NVIDIA Software Stack
SiTech AI Team3 min read

DOCA GPUnetIO Unifies GPU Initiated Networking Across NVIDIA Software Stack

DOCA GPUnetIO gives CUDA kernels a shared GDA-KI base to drive Ethernet, RDMA, Verbs and DMA while keeping the CPU off the critical path. NCCL, NVSHMEM, UCX/NIXL and Holoscan now build on one implementation.

A shared base for GPU initiated networking

DOCA GPUnetIO is a GPU-centric networking SDK layer for real-time packet processing and data movement. It combines GPUDirect RDMA, GPUDirect Async Kernel-Initiated (GDA-KI) and GDRCopy so CUDA kernels can directly drive Ethernet, RDMA, Verbs and DMA operations while keeping the CPU out of the application critical path. NVIDIA describes the framework as a unified GDA-KI foundation used by NVSHMEM, NCCL, Aerial 5G SDK, UCX/NIXL, NVQLink with Holoscan Sensor Bridge, Holoscan Advanced Network Operator, DeepEP/HybridEP and others.

Open source and SDK forms

NVIDIA ships GPUnetIO as the full DOCA SDK superset and as a lighter open-source Verbs-focused implementation. The SDK version covers Verbs, Ethernet, DMA and Comm Channel integration, while the open version targets frameworks that want a fully open networking transport path. The open implementation can detect the DOCA SDK at runtime and call selected closed-source SDK functions through dlopen when the documented environment variable points to the correct SDK path; otherwise it runs standalone. Both forms expose a device-facing CUDA API, keeping the GPU programming model aligned.

Control path, data path and APIs

GPUnetIO applications start with a CPU control path that initializes GPU and network devices, allocates memory, creates transport objects and exports descriptors into GPU memory. High-level helpers such as doca_gpu_verbs_create_qp_hl condense RDMA queue pair creation and connection. On the GPU data path, CUDA threads post work queue entries, ring the network card doorbell and optionally poll completion queue entries. Doorbell modes include regular MMIO writes from CUDA memory, BlueFlame writes of the whole WQE for latency-sensitive cases with few queues, and a CPU-assisted mode for systems without a direct GPU-to-NIC connection such as DGX Spark. For Ethernet and RDMA Verbs, high-level calls such as doca_gpu_dev_verbs_put_signal and doca_gpu_dev_eth_txq_send handle synchronization and doorbell submission at thread, warp or block scope. Low-level primitives let applications prepare WQEs with different opcodes, submit queues and poll completions, but they are not thread safe and require application synchronization. Reference examples are available in the open-source repository, with broader SDK samples and an Ethernet application in the DOCA SDK.

NCCL and NVSHMEM integrations

NCCL GIN exposes device-side put, get, signal, wait and flush primitives. Since NCCL 2.27, GIN has used the open-source GPUnetIO Verbs path as a GDA-KI backend, mapping GIN operations to GPUnetIO calls for WQE preparation, doorbell ringing and completion polling. NVSHMEM 3.7 added a GPUnetIO transport that follows the IBGDA structure while replacing larger low-level code portions with GPUnetIO calls. It supports GDA-KI and host-initiated RDMA, selected with NVSHMEM_REMOTE_TRANSPORT=gpunetio and NVSHMEM_GPUNETIO_ENABLE_GDAKI=1. NVQLink uses the GPUnetIO-based GPU RoCE Transceiver operator for about 2.6 microseconds minimum round-trip latency on IGX Thor with Blackwell GPU and ConnectX-7.

Sources: Nvidia Dev

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.