Parallelize compilation of kernels used by a graph
The cold-start time of applications using Warp can be very high. Newton's MuJoCo ANYmal D example takes about a minute on a powerful workstation. Most of the time is spent compiling kernel modules. The issue is that they are compiled on demand, and serially. CUDA graph construction requires having a valid compiled kernel handle.
GH-1086 offers compiling modules in parallel through wp.load_module() and wp.force_load(), but this requires the user to know in advance which ones will be needed early on. They may specify too many, or too little, or the kernels may have specializations not easily known ahead of time.
APIC, Warp's own graph representation used for persistent capture and replay, as well as an equivalent of CUDA graphs for the CPU device, could offer a solution by deferring compilation of multiple modules until graph launch time.
GH-1659 tracks the creation of deferred CUDA graph creation itself, while this Issue is for implementing deferred compilation on top of it.
Source: NVIDIA/warp