Original release notes
What's Changed
- MLX: Qwen3.8 Flash Next support
- cmake: make external compat patches idempotent
- MLX and llama.cpp update
- mlxrunner: add structured output support
- mlxrunner: avoid Metal GPU timeouts when loading models from slow storage
New Contributors
- @pd95 made their first contribution in https://github.com/ollama/ollama/pull/17948
Full Changelog: https://github.com/ollama/ollama/compare/v0.33.0...v0.33.1

