llama.cpp vs Ollama
Choose how to run language models locally, from a low-level inference runtime to a packaged model runner.
At a glance
| At a glance | llama.cpp | Ollama |
|---|---|---|
| License | MIT | MIT |
| Languages | C++, C | Go |
| Deployment | Runs locally / Self-hosted | Runs locally / Self-hosted |
| Maturity | Established | Established |
| Stars | 126k | 180k |
| Star growth over the last 7 days | +1.1k ★ | +545 ★ |
| Forks | 22.4k | 17.6k |
| Open issues | 2.3k | 3.8k |
| Last commit | 28 Aug 2026 | 28 Aug 2026 |
| Activity | Active | Active |
What each one does
llama.cpp
The quantisation work here is what put usable models on hardware without a datacentre GPU, and much of the local-model ecosystem is built on top of it. Working with it directly means dealing with build flags and hardware specifics that higher-level wrappers hide from you.
Full entry →Ollama
Pulls a quantised model — one compressed to run in less memory — and starts a local HTTP API in front of it, so running a model is a single command. The obvious starting point for building against a model on your own hardware. For serving many people at once, a server built for throughput will do considerably better.
Full entry →What you can do
llama.cpp
- Run a model in one command — llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF pulls the GGUF straight from Hugging Face and drops you into an interactive session, with no Python environment involved.
- Serve an OpenAI-compatible endpoint — llama serve -hf <repo> starts a REST server that existing OpenAI clients can be pointed at, and ships a built-in web UI for trying the model.
- Fit models larger than VRAM — Integer quantisation from 1.5-bit through 8-bit, combined with CPU+GPU hybrid inference, lets you offload only part of a model to the GPU and keep the rest in system RAM.
- Target the hardware you have — One source tree covers Metal and ARM NEON on Apple silicon, CUDA on NVIDIA, HIP on AMD, MUSA on Moore Threads, plus Vulkan, SYCL, OpenCL, WebGPU and CANN backends.
- Constrain output with GBNF grammars — A GBNF grammar file forces generation to match a syntax you define, which is how you get reliably structured output out of a small local model.
Ollama
- Run a model with one command —
ollama run gemma4fetches what it needs and drops you straight into a conversation. The models you can run are listed at ollama.com/library. - Choose between your own hardware and their cloud — The same commands either run the model on your machine or, for models too large for it, run them on Ollama's cloud instead.
- Call it from your own app — Send JSON to
http://localhost:11434/api/chatand a reply comes back. Official Python and JavaScript clients install withpip install ollamaornpm i ollama. - Point a coding assistant at your own model —
ollama launch claudestarts Claude Code against a model running on your machine, and the same command covers Codex, Copilot CLI, OpenCode, Droid and DeepSeek Harness. - Reach it from the chat apps you use —
ollama launch openclawputs your local model behind WhatsApp, Telegram, Slack and Discord as a personal assistant.