LocalAI
Drop-in OpenAI-compatible self-hosted AI API. Local models, no GPU required.
Release history
LocalAI 4.9.0
Authentication now deny-by-default, chat adds context compression, MiniMax-H3 video generation, and improved operator visibility across 146 pull requests
LocalAI 4.8.2
LocalAI 4.8.2 adds NVIDIA NeMo-Speech.cpp backend for speech processing and improves model gallery resilience with mirror fallbacks
LocalAI 4.8.1
LocalAI 4.8.1 fixes VRAM handling for malformed GGUF metadata, restores vllm-cpp backend ABI compatibility, and adds Qwen3.5 variants to model gallery
LocalAI 4.8.0
LocalAI 4.8.0 ships vllm.cpp (C++20 inference engine), 3D generation, audio.cpp multi-family backend, VRAM budgets, and distributed mode hardening across 386 merged PRs
LocalAI 4.7.1
Scopes llama.cpp serving options to the llama.cpp code path only
LocalAI 4.6.2
Use slices.Contains to simplify code; Shard single-arch backend ma
LocalAI 4.6.0
AMD ROCm backends now offload to GPU correctly, distributed worker failures no longer lock model loads, and realtime sessions warm pipelines eagerly to eliminate cold-start stalls
LocalAI 4.5.6
LocalAI adds native voice and face detection backends, improves distributed state management across replicas, and fixes GGUF context sizing and gallery matching
LocalAI 4.5.5
Fixed CI build failures across multiple backends, whisper macOS library loading, and added new gallery models
LocalAI 4.5.4
LocalAI 4.5.4 fixes Darwin backend binary derivation from exec lines, resolving macOS model execution issues
LocalAI 4.5.2
LocalAI 4.5.2 fixes the Opus backend to build and package correctly on macOS/Darwin
LocalAI 4.5.0
LocalAI 4.5.0 adds depth perception, sound-event tagging, on-device multilingual TTS, PII filtering via NER, speaker-aware realtime conversations, and prefix caching for concurrent multi-user serving without configuration
LocalAI 4.4.1
LocalAI 4.4.1 restores vLLM 0.22+ compatibility and adds streaming pipeline stages for realtime LLM, TTS, and transcription workflows
LocalAI 4.4.0
LocalAI 4.4.0 adds multimodal capabilities: parakeet.cpp and CrispASR audio backends, video understanding, object detection, LTX-2 video generation, distributed prefix-cache routing, and security hardening
LocalAI 4.3.6
LocalAI 4.3.6 adds NVIDIA NeMo Parakeet ASR backend, hardens HTTP clients against redirects, and updates core dependencies
LocalAI 4.3.5
LocalAI 4.3.5 fixes tool-calling bugs with streaming and tokenizer templates, adds per-request reasoning effort control, and updates dependencies
LocalAI 4.3.3
LocalAI 4.3.3 updates core inference backends, fixes OpenAI API response compatibility, and optimizes UI bundle performance
LocalAI 4.3.1
Fixes build break in kokoros backend. Upgrade if that backend is in use
LocalAI 4.2.5
LocalAI 4.2.5 fixes Ollama model listing, TTS output handling, parameter parsing, and improves OpenAI API compliance with upstream dependency updates
LocalAI 4.2.4
LocalAI 4.2.4 fixes distributed node cleanup, proxy prefix handling, agent job persistence races, and OpenAI tool_choice parsing, plus adds Vulkan VRAM detection and Liquid Audio speech models
LocalAI 4.2.0
LocalAI 4.2.0 adds voice and face recognition, diarization, Ollama-compatible API, video generation, and eleven new backends for multimodal local AI workloads