warp · Issues· 331 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #1980
Loop unroll criteria off by one
bugUpdated Sep 22, 2026 - #1979
Adam and SGD partially update parameters before rejecting a short gradient list
Updated Sep 22, 2026 - #1978
FEM first-sample interpolation returns zero gradients with triplet construction
bugwarp.femUpdated Sep 22, 2026 - #1977
Parallelize compilation of kernels used by a graph
feature requestUpdated Sep 21, 2026 - #474
Add support for NCCL
Updated Sep 19, 2026 - #420
[BUG] Investigate Inconsistencies With Array Shape and Strides
buginteropUpdated Sep 19, 2026 - #1361
Optimize tiled CUDA launches by reading blockIdx.x directly
tileUpdated Sep 19, 2026 - #464
Add ability to update kernel parameters of a captured graph
feature requestruntimeUpdated Sep 19, 2026 - #1965
tile_matmul(): use the cuBLASDx register-accumulator API with suggested shared layouts for Tensor Core precisions
tileUpdated Sep 18, 2026 - #1837
APIC portability onto other devices
feature requestUpdated Sep 17, 2026 - #1972
Improve the error when NumPy attempts to convert a CUDA wp.array
docsUpdated Sep 17, 2026 - #937
[REQ] Option for TF32 calculation in `wp.tile_matmul()`
feature requestUpdated Sep 17, 2026 - #1971
Reduce redundant radix sorting in CUDA BSR transpose
warp.sparseUpdated Sep 17, 2026 - #1970
Support first-sample FEM interpolation with row compression
warp.femUpdated Sep 17, 2026 - #1968
@wp.kernel(launch_bounds=...) fails under wp.config.llvm_cuda
Updated Sep 17, 2026