Original release notes
cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) (#25545)
- cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel)
- chore : remove indentation of #pragma unroll
- cuda : remove unnecessary kernel template declarations
- cuda : add WARPS_PER_BLOCK and K_VECS_PER_BLOCK template parameters in lightning indexer kernels to avoid duplication of constants.
- cuda : relax MMA architecture requirements to Turing in lightning indexer implementation
- chore : renamed variables
- chore : rename ggml_cuda_op_lightning_indexer() to ggml_cuda_lightning_indexer()
- chore : TODO for AMD rocWMMA
- chore : whitespace formatting
- chore : another variable rename to fix problems caused by shadowing
- chore : yet another rename, this time uppercased all constants
- cuda : added alignment checks for Q and K tensors in lightning indexer implementation
---------
Co-authored-by: Stanisaw Szymczyk
UI:

