Back to Freedom.Tech Back
All llama.cpp releasesAll versions
Release Wed, Jul 15, 2026 1 min read

llama.cpp b10032

Original release notes

cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) (#25545)

  • cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel)
  • chore : remove indentation of #pragma unroll
  • cuda : remove unnecessary kernel template declarations
  • cuda : add WARPS_PER_BLOCK and K_VECS_PER_BLOCK template parameters in lightning indexer kernels to avoid duplication of constants.
  • cuda : relax MMA architecture requirements to Turing in lightning indexer implementation
  • chore : renamed variables
  • chore : rename ggml_cuda_op_lightning_indexer() to ggml_cuda_lightning_indexer()
  • chore : TODO for AMD rocWMMA
  • chore : whitespace formatting
  • chore : another variable rename to fix problems caused by shadowing
  • chore : yet another rename, this time uppercased all constants
  • cuda : added alignment checks for Q and K tensors in lightning indexer implementation

---------

Co-authored-by: Stanisaw Szymczyk

UI: