Ollama
Local LLM runner releases. Ships new model support and inference improvements.
Release history
Ollama 0.33.1
Ollama 0.33.1 adds Qwen3.8 Flash support, structured output for MLX models, and fixes Metal GPU timeouts during model loading
Ollama 0.33.0
Claude Desktop integration: toggle individual Ollama models on and off from the menu bar, choose models directly within Claude
Ollama 0.32.15
Model metadata caching halves time-to-first-token, parser error handling fixed, Qwen system messages normalized
Ollama 0.32.11
Ollama added DeepSeek Harness and Meta Muse Code agent support, plus web search in OpenAI-compatible Responses API
Ollama 0.32.10
Ollama 0.32.10 changes default repeat penalty to 1.0 for faster speculative decoding, speeds NVFP4 prefill%, and fixes blob verification on shared digests
Ollama 0.32.9
Ollama adds NVIDIA's Nemotron 3.5 Lightning, a 30B mixture-of-experts model optimized for always-on agent deployment
Ollama 0.32.8
Muse Glimmer model now available across all platforms, optimized for coding agents and personal assistants with MLX performance on Apple Silicon
Ollama 0.32.7
Muse Glimmer, Meta's 30B multimodal agent model, now runs locally on Ollama with Apple Silicon optimization and image input support
Ollama 0.32.6
Ollama 0.32.6 speeds up Qwen3.5 on Apple GPUs via speculative decoding, fixes OpenAI API streaming format compliance, and improves TUI usability
Ollama 0.32.4
Ollama 0.32.4 adds Laguna support for Apple GPUs, fixes Qwen3 MoE expert quantization, and improves speculative decoding performance
Ollama 0.32.3
Ollama 0.32.3 fixes stalled downloads, expands GPU support to Windows ARM64 and B200, and adds chat/tool-calling for Laguna 2.1 models
Ollama 0.32.2
Ollama 0.32.2 adds agent skills system, fixes Anthropic thinking-block handling, updates llama.cpp and MLX backends, and enables unlimited tool rounds for cloud models
Ollama 0.32.1
Ollama 0.32.1 fixes MLX model cache leaks, improves Gemma 4 tool calling, and adds working directory context to agents
Ollama 0.32.0
Adds Qwen3.5 parser and renderer selection; Agent UI warns before loading old agent models
Ollama 0.31.2
Flash attention now works on older NVIDIA GPUs, vision model offloading improved, and non-UTF-8 model paths fixed
Ollama 0.31.1
Gemma 4 nearly 90 percent faster on Apple Silicon via multi-token prediction
Ollama 0.30.12
Fixed tool call parsing to handle JSON strings with braces correctly, and updated llama.cpp and MLX dependencies
Ollama 0.30.11
Ollama 0.30.11 fixes GPU detection on hybrid Windows systems, improves speculative decoding, and corrects memory reporting for partially offloaded models
Ollama 0.30.10
Command and Northfamily models now run on Apple Silicon via MLX engine, llama.cpp updated to build 9672
Ollama 0.30.9
Ollama 0.30.9 adds Cohere2Moe support, fixes token output limits in coding agents, and validates message sizes against context windows
Ollama 0.30.8
Ollama 0.30.8 improves prompt caching efficiency, fixes provider selection bugs, and hardens MLX inference stability with snapshot creation
Ollama 0.30.7
Ollama adds Hermes Desktop, a native UI for managing Hermes agent conversations and integrations via ollama launch
Ollama 0.30.6
Ollama adds Gemma 4 QAT models for lower memory use, integrates Oh My Pi coding agent, improves Apple Silicon quantization
Ollama 0.30.3
Ollama added Gemma 4 12B model support, expanding locally-runnable open-weight options for builders
Ollama 0.30.2
Auto-install Cline CLI; Qwen code integration added; Template logging improved for troubleshooting
Ollama 0.23.4
Ollama launch now handles vision models with image inputs. Local image paths work in Claude tool results
Ollama 0.30.0
Ollama now builds directly on llama.cpp with GGUF format support and MLX acceleration for Apple Silicon inference