LocalAI

Package

Run LLM, image and audio models on your own hardware behind OpenAI-style APIs

Price
Free, open source
Access
None; optional API key

About

MIT-licensed AI engine you run on your own machines, CPU-only or with NVIDIA, AMD, Intel or Apple GPUs. Pluggable backends such as llama.cpp, vLLM and whisper.cpp serve text, image and audio models behind OpenAI-, Anthropic- and ElevenLabs-compatible APIs, with a built-in web UI.

What you can do with it

  • Serve open models behind an OpenAI-compatible API on your own servers
  • Run chat, image generation and transcription on a CPU-only machine
  • Pull GGUF models from Hugging Face, Ollama or the built-in model gallery

Get started

  1. Start the container with docker run, or install the macOS or Linux app
  2. Install a model from the web UI's model gallery or with local-ai run
  3. Call the OpenAI-compatible API on localhost:8080

Example

docker run -p 8080:8080 --name local-ai -ti localai/localai:latest
# install qwen3-4b from Models > Explore at http://localhost:8080, then:
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-4b",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Details

Hosting
Self-hosted, Runs locally
Available in
Worldwide
MCP server
Local

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .