Back to Freedom.Tech Back
All llama.cpp releasesAll versions
Release Tue, Jun 9, 2026 1 min read

llama.cpp b9585

Original release notes

graph: Fix granite speech model inference by applying embedding scale when deepstack is not used (#24357)

  • llama-graph : apply embedding scale when deepstack is not used
  • nits: remove non-existant hunyuan-vl from the tests
  • apply suggestion from @gabe-l-hart

---------

Co-authored-by: Xuan Son Nguyen

UI: