Model Serving
Running models on your own hardware behind an OpenAI-compatible endpoint. Agents are unusually chatty, so throughput and concurrency matter more here than single-request latency.
6 projects
Pulls a quantised model — one compressed to run in less memory — and starts a local HTTP API in front of it, so running a model is a single command. The obvious starting point for building against a model on your own hardware. For serving many people at once, a server built for throughput will do considerably better.
The quantisation work here is what put usable models on hardware without a datacentre GPU, and much of the local-model ecosystem is built on top of it. Working with it directly means dealing with build flags and hardware specifics that higher-level wrappers hide from you.
Paged attention and continuous batching let one GPU serve far more concurrent requests than a naive loop, which matters because agents are unusually chatty. The OpenAI-compatible endpoint means most agent frameworks point at it with a base-URL change. Expect real operational work around GPU memory sizing and model loading times.
The library half lets one function call reach any supported provider in the OpenAI request and response shape. The gateway half is the reason most teams arrive: a self-hosted server that issues virtual keys with their own budgets and rate limits, falls back when a provider fails, spreads load across deployments, and records cost per key. The price is an extra component in the request path that you now operate, and code under the enterprise directory is not covered by the MIT grant.
Point any OpenAI client at it and it answers: chat with tool calling, embeddings, image generation, transcription and speech, a reranker, even the realtime speech-to-speech API — plus imitations of the Anthropic, ElevenLabs and Ollama APIs. A small core detects your hardware and pulls inference backends on demand — llama.cpp, vLLM, MLX, whisper.cpp among many — a web gallery installs models by click, autonomous agents with MCP support are embedded, and API keys, OIDC and quotas cover multiple users. The honest comparison: Ollama stays simpler for chatting with a model on a laptop, a bare vLLM is leaner for one model at maximum throughput, and there is no native Windows build — containers or WSL only.
Its prefix cache pays off precisely in the agent case, where many calls share a long, identical system prompt. Also strong at constrained decoding when responses must match a schema. The main alternative in this space is vLLM; benchmark both on your own traffic shape.