Back to Freedom.Tech Back
All Ollama releasesAll versions
Release Mon, Jul 6, 2026 1 min read

Ollama 0.31.2

Original release notes

What's Changed

  • Enabled flash attention on older NVIDIA GPUs (compute capability 6.x)
  • iGPU can now offload vision models with padding to fit available memory
  • Fixed structured output for thinking models when thinking is disabled
  • Hardened GGUF model creation
  • ollama launch for Claude Code now disables telemetry by default
  • Fixed loading models on paths with non-UTF-8 characters
  • Updated the MLX and llama.cpp engines

New Contributors

  • @kevinpark1217 made their first contribution in https://github.com/ollama/ollama/pull/16949

Full Changelog: https://github.com/ollama/ollama/compare/v0.31.1...v0.31.2