LocalAI
PackageRun LLM, image and audio models on your own hardware behind OpenAI-style APIs
- Price
- Free, open source
- Access
- None; optional API key
About
MIT-licensed AI engine you run on your own machines, CPU-only or with NVIDIA, AMD, Intel or Apple GPUs. Pluggable backends such as llama.cpp, vLLM and whisper.cpp serve text, image and audio models behind OpenAI-, Anthropic- and ElevenLabs-compatible APIs, with a built-in web UI.
What you can do with it
- Serve open models behind an OpenAI-compatible API on your own servers
- Run chat, image generation and transcription on a CPU-only machine
- Pull GGUF models from Hugging Face, Ollama or the built-in model gallery
Get started
- Start the container with docker run, or install the macOS or Linux app
- Install a model from the web UI's model gallery or with local-ai run
- Call the OpenAI-compatible API on localhost:8080
Example
docker run -p 8080:8080 --name local-ai -ti localai/localai:latest
# install qwen3-4b from Models > Explore at http://localhost:8080, then:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-4b",
"messages": [{"role": "user", "content": "Hello!"}]
}'Details
- Hosting
- Self-hosted, Runs locally
- Available in
- Worldwide
- MCP server
- Local