Back to Freedom.Tech Back
All llama.cpp releasesAll versions
Release Tue, Jun 23, 2026 1 min read

llama.cpp b9767

Original release notes

ggml-webgpu: improve MTP inference by using mat-vec path for small batches (#24811)

  • ggml-webgpu: improve small batches decoding
  • Add barrier to the NUM_COLS loop in mul-mat-vec

UI: