Ollama

Package

Run open models on your own machine behind a local REST API

Price
Free, open source; cloud plans paid
Access
None locally; API key for cloud

About

Open-source runtime that downloads and runs open-weight models on macOS, Windows and Linux, with an optional paid cloud for larger models. Suits local development and private data. Speed depends on your GPU or unified memory.

What you can do with it

  • Download and chat with open models like Gemma and Qwen offline
  • Serve a local OpenAI- and Anthropic-compatible API for apps and coding agents
  • Generate embeddings locally for RAG without sending data to a cloud

Get started

  1. Download Ollama for macOS, Windows or Linux
  2. Pull a model with ollama pull

Example

ollama pull gemma4:e2b
curl http://localhost:11434/api/chat \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma4:e2b",
    "messages": [{"role": "user", "content": "Say hello in one sentence."}],
    "stream": false
  }'

Details

Hosting
Hosted service, Self-hosted, Runs locally
Available in
Worldwide
Official SDKs
Python, JavaScript/TypeScript
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .