llama.cpp vs Ollama

Choose how to run language models locally, from a low-level inference runtime to a packaged model runner.

At a glance

At a glancellama.cppOllama
LicenseMITMIT
LanguagesC++, CGo
DeploymentRuns locally / Self-hostedRuns locally / Self-hosted
MaturityEstablishedEstablished
Stars126k180k
Star growth over the last 7 days+1.1k ★+545 ★
Forks22.4k17.6k
Open issues2.3k3.8k
Last commit28 Aug 202628 Aug 2026
ActivityActiveActive

What each one does

llama.cpp

The quantisation work here is what put usable models on hardware without a datacentre GPU, and much of the local-model ecosystem is built on top of it. Working with it directly means dealing with build flags and hardware specifics that higher-level wrappers hide from you.

Full entry →

Ollama

Pulls a quantised model — one compressed to run in less memory — and starts a local HTTP API in front of it, so running a model is a single command. The obvious starting point for building against a model on your own hardware. For serving many people at once, a server built for throughput will do considerably better.

Full entry →

What you can do

llama.cpp

  • Run a model in one command — llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF pulls the GGUF straight from Hugging Face and drops you into an interactive session, with no Python environment involved.
  • Serve an OpenAI-compatible endpoint — llama serve -hf <repo> starts a REST server that existing OpenAI clients can be pointed at, and ships a built-in web UI for trying the model.
  • Fit models larger than VRAM — Integer quantisation from 1.5-bit through 8-bit, combined with CPU+GPU hybrid inference, lets you offload only part of a model to the GPU and keep the rest in system RAM.
  • Target the hardware you have — One source tree covers Metal and ARM NEON on Apple silicon, CUDA on NVIDIA, HIP on AMD, MUSA on Moore Threads, plus Vulkan, SYCL, OpenCL, WebGPU and CANN backends.
  • Constrain output with GBNF grammars — A GBNF grammar file forces generation to match a syntax you define, which is how you get reliably structured output out of a small local model.

Ollama

  • Run a model with one command — ollama run gemma4 fetches what it needs and drops you straight into a conversation. The models you can run are listed at ollama.com/library.
  • Choose between your own hardware and their cloud — The same commands either run the model on your machine or, for models too large for it, run them on Ollama's cloud instead.
  • Call it from your own app — Send JSON to http://localhost:11434/api/chat and a reply comes back. Official Python and JavaScript clients install with pip install ollama or npm i ollama.
  • Point a coding assistant at your own model — ollama launch claude starts Claude Code against a model running on your machine, and the same command covers Codex, Copilot CLI, OpenCode, Droid and DeepSeek Harness.
  • Reach it from the chat apps you use — ollama launch openclaw puts your local model behind WhatsApp, Telegram, Slack and Discord as a personal assistant.