← Back to all projects

LiteLLM

One OpenAI-shaped interface in front of a hundred model providers — as a library or a gateway

OfficialMIT
Stars
57.5k
Forks
11k
Open issues
4.9k
Last commit
28 Aug 2026

What is LiteLLM?

The library half lets one function call reach any supported provider in the OpenAI request and response shape. The gateway half is the reason most teams arrive: a self-hosted server that issues virtual keys with their own budgets and rate limits, falls back when a provider fails, spreads load across deployments, and records cost per key. The price is an extra component in the request path that you now operate, and code under the enterprise directory is not covered by the MIT grant.

What can you do with LiteLLM?

  • Call any provider with one function — A single completion call takes a model name and returns the OpenAI response shape, so switching between OpenAI, Anthropic, Vertex AI, Bedrock and the rest is a string change rather than a rewrite of the call site.
  • Put a gateway in front of everything — The self-hosted proxy exposes an OpenAI-compatible endpoint, which means existing clients, SDKs and tools that already speak OpenAI point at it unchanged and gain routing they never had to know about.
  • Hand out keys with limits attached — Virtual keys carry their own budgets and rate limits per key, per team and per user, so an experiment cannot spend the department's quota and the provider's real credential never leaves the gateway.
  • Keep serving when a provider does not — Fallbacks move a failing request to another deployment automatically, and load can be spread across several deployments of the same model rather than pinned to one.
  • See the spend and filter the traffic — Logging, cost tracking and an admin dashboard cover what went through and what it cost, and guardrails at the same layer can apply content filtering and mask personal data before a request leaves.

Before you choose LiteLLM

  • The MIT grant stops at the enterprise directory, which the licence file places under separate terms — so the open part and the commercial part live in the same repository and the boundary is a path.
  • Adopting the gateway puts a component you operate into the path of every model call, so its availability and its upgrades become your concern on top of the providers' own.

Star history

21 Aug to 28 Aug · +553

56.9k57.5k

Frequently asked questions

Is LiteLLM free for commercial use?

LiteLLM is released under the MIT licence — OSI-approved open source, which permits commercial use.

How can LiteLLM be deployed?

LiteLLM is available as Self-hosted / Runs locally.

Documentation

Reproduced from the BerriAI/litellm README, published under UNKNOWN — read the LICENSE file. Read the original ↗


What is LiteLLM

LiteLLM is an open source AI Gateway that gives you a single, unified interface to call 100+ LLM providers — OpenAI, Anthropic, Gemini, Bedrock, Azure, and more — using the OpenAI format.

Use it as a Python SDK for direct library integration, or deploy the AI Gateway (Proxy Server) as a centralized service for your team or organization.

Jump to LiteLLM Proxy (LLM Gateway) Docs Jump to Supported LLM Providers


Why LiteLLM

Managing LLM calls across providers gets complicated fast — different SDKs, auth patterns, request formats, and error types for every model. LiteLLM removes that friction:

  • Unified API — one interface for 100+ LLMs, no provider-specific SDK juggling
  • Drop-in OpenAI compatibility — swap providers without rewriting your code
  • Production-ready gateway — virtual keys, spend tracking, guardrails, load balancing, and an admin dashboard out of the box
  • 8ms P95 latency at 1k RPS (benchmarks)

OSS Adopters


Features

All Supported Endpoints - /chat/completions, /responses, /embeddings, /images, /audio, /batches, /rerank, /a2a, /messages and more.

Python SDK

uv add litellm
from litellm import completion
import os

os.environ["OPENAI_API_KEY"] = "your-openai-key"
os.environ["ANTHROPIC_API_KEY"] = "your-anthropic-key"

# OpenAI
response = completion(model="openai/gpt-4o", messages=[{"role": "user", "content": "Hello!"}])

# Anthropic  
response = completion(model="anthropic/claude-sonnet-4-20250514", messages=[{"role": "user", "content": "Hello!"}])

AI Gateway (Proxy Server)

Getting Started - E2E Tutorial - Setup virtual keys, make your first request

uv tool install 'litellm[proxy]'
litellm --model gpt-4o
import openai

client = openai.OpenAI(api_key="anything", base_url="http://0.0.0.0:4000")
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello!"}]
)

Docs: LLM Providers

Supported Providers - LangGraph, Vertex AI Agent Engine, Azure AI Foundry, Bedrock AgentCore, Pydantic AI

Python SDK - A2A Protocol

from litellm.a2a_protocol import A2AClient
from a2a.types import SendMessageRequest, MessageSendParams
from uuid import uuid4

client = A2AClient(base_url="http://localhost:10001")

request = SendMessageRequest(
    id=str(uuid4()),
    params=MessageSendParams(
        message={
            "role": "user",
            "parts": [{"kind": "text", "text": "Hello!"}],
            "messageId": uuid4().hex,
        }
    )
)
response = await client.send_message(request)

AI Gateway (Proxy Server)

Step 1. Add your Agent to the AI Gateway — set protocolVersion to 1.0 or 0.3 per agent

Step 2. Call Agent via A2A SDK (requires a2a-sdk>=1.1.0)

import httpx
from a2a.client import A2ACardResolver, ClientConfig, ClientFactory
from a2a.types import Message, Part, Role, SendMessageRequest
from a2a.utils.constants import TransportProtocol
from uuid import uuid4

base_url = "http://localhost:4000/a2a/my-agent"  # LiteLLM proxy + agent name
headers = {"Authorization": "Bearer sk-1234"}    # LiteLLM Virtual Key

async with httpx.AsyncClient(headers=headers, timeout=60.0) as http_client:
    resolver = A2ACardResolver(httpx_client=http_client, base_url=base_url)
    agent_card = await resolver.get_agent_card()
    config = ClientConfig(
        httpx_client=http_client,
        streaming=False,
        supported_protocol_bindings=[TransportProtocol.JSONRPC, TransportProtocol.HTTP_JSON],
    )
    client = ClientFactory(config).create(agent_card)

    request = SendMessageRequest(
        message=Message(
            message_id=uuid4().hex,
            role=Role.ROLE_USER,
            parts=[Part(text="Hello!")],
        )
    )
    async for event in client.send_message(request):
        populated = event.ListFields()
        if populated and populated[0][0].name in ("message", "msg"):
            print("".join(getattr(p, "text", "") or "" for p in populated[0][1].parts))

Docs: A2A Agent Gateway

Python SDK - MCP Bridge

from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
from litellm import experimental_mcp_client
import litellm

server_params = StdioServerParameters(command="python", args=["mcp_server.py"])

async with stdio_client(server_params) as (read, write):
    async with ClientSession(read, write) as session:
        await session.initialize()

        # Load MCP tools in OpenAI format
        tools = await experimental_mcp_client.load_mcp_tools(session=session, format="openai")

        # Use with any LiteLLM model
        response = await litellm.acompletion(
            model="gpt-4o",
            messages=[{"role": "user", "content": "What's 3 + 5?"}],
            tools=tools
        )

AI Gateway - MCP Gateway

Step 1. Add your MCP Server to the AI Gateway

Step 2. Call MCP tools via /chat/completions

curl -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
  -H 'Authorization: Bearer sk-1234' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Summarize the latest open PR"}],
    "tools": [{
      "type": "mcp",
      "server_url": "litellm_proxy/mcp/github",
      "server_label": "github_mcp",
      "require_approval": "never"
    }]
  }'

Use with Cursor IDE

{
  "mcpServers": {
    "LiteLLM": {
      "url": "http://localhost:4000/mcp/",
      "headers": {
        "x-litellm-api-key": "Bearer sk-1234"
      }
    }
  }
}

Docs: MCP Gateway

Supported Providers (Website Supported Models | Docs)

Provider/chat/completions/messages/responses/embeddings/image/generations/audio/transcriptions/audio/speech/moderations/batches/rerank
Abliteration (abliteration)✅
AI/ML API (aiml)✅✅✅✅✅
AI21 (ai21)✅✅✅
AI21 Chat (ai21_chat)✅✅✅
Aleph Alpha✅✅✅
Amazon Nova✅✅✅
Anthropic (anthropic)✅✅✅✅
Anthropic Text (anthropic_text)✅✅✅✅
Anyscale✅✅✅
AssemblyAI (assemblyai)✅✅✅✅
Auto Router (auto_router)✅✅✅
AWS - Bedrock (bedrock)✅✅✅✅✅
AWS - Sagemaker (sagemaker)✅✅✅✅
Azure (azure)✅✅✅✅✅✅✅✅✅
Azure AI (azure_ai)✅✅✅✅✅✅✅✅✅
Azure Text (azure_text)✅✅✅✅✅✅✅
Baseten (baseten)✅✅✅
Bytez (bytez)✅✅✅
Cerebras (cerebras)✅✅✅
Clarifai (clarifai)✅✅✅
Cloudflare AI Workers (cloudflare)✅✅✅
Codestral (codestral)✅✅✅
Cognition (cognition)✅✅✅
Cohere (cohere)✅✅✅✅✅
Cohere Chat (cohere_chat)✅✅✅
CometAPI (cometapi)✅✅✅✅
CompactifAI (compactifai)✅✅✅
Custom (custom)✅✅✅
Custom OpenAI (custom_openai)✅✅✅✅✅✅✅
Dashscope (dashscope)✅✅✅✅✅
Databricks (databricks)✅✅✅
DataRobot (datarobot)✅✅✅
Deepgram (deepgram)✅✅✅✅
DeepInfra (deepinfra)✅✅✅
Deepseek (deepseek)✅✅✅
ElevenLabs (elevenlabs)✅✅✅✅✅
Empower (empower)✅✅✅
Fal AI (fal_ai)✅✅✅✅
Featherless AI (featherless_ai)✅✅✅
Fireworks AI (fireworks_ai)✅✅✅
FriendliAI (friendliai)✅✅✅
Galadriel (galadriel)✅✅✅
GitHub Copilot (github_copilot)✅✅✅✅
GitHub Models (github)✅✅✅
Google - PaLM✅✅✅
Google - Vertex AI (vertex_ai)✅✅✅✅✅
Google AI Studio - Gemini (gemini)✅✅✅
GradientAI (gradient_ai)✅✅✅
Groq AI (groq)✅✅✅
Heroku (heroku)✅✅✅
Hosted VLLM (hosted_vllm)✅✅✅
Huggingface (huggingface)✅✅✅✅✅
Hyperbolic (hyperbolic)✅✅✅
IBM - Watsonx.ai (watsonx)✅✅✅✅
Infinity (infinity)✅
Jina AI (jina_ai)✅
Lambda AI (lambda_ai)✅✅✅
Lemonade (lemonade)✅✅✅
LiteLLM Proxy (litellm_proxy)✅✅✅✅✅
Llamafile (llamafile)✅✅✅
LM Studio (lm_studio)✅✅✅
Maritalk (maritalk)✅✅✅
Meta - Llama API (meta_llama)✅✅✅
Mistral AI API (mistral)✅✅✅✅
ModelScope (modelscope)✅✅✅✅
Moonshot (moonshot)✅✅✅
Morph (morph)✅✅✅
Nebius AI Studio (nebius)✅✅✅✅
NLP Cloud (nlp_cloud)✅✅✅
Novita AI (novita)✅✅✅
Nscale (nscale)✅✅✅
Nvidia NIM (nvidia_nim)✅✅✅
OCI (oci)✅✅✅
Ollama (ollama)✅✅✅✅
Ollama Chat (ollama_chat)✅✅✅
Oobabooga (oobabooga)✅✅✅✅✅✅✅
OpenAI (openai)✅✅✅✅✅✅✅✅✅
OpenAI-like (openai_like)✅
OpenRouter (openrouter)✅✅✅
OVHCloud AI Endpoints (ovhcloud)✅✅✅
Perplexity AI (perplexity)✅✅✅
Petals (petals)✅✅✅
Pinstripes (pinstripes)✅✅✅
Predibase (predibase)✅✅✅
Recraft (recraft)✅
Replicate (replicate)✅✅✅
Sagemaker Chat (sagemaker_chat)✅✅✅
Sambanova (sambanova)✅✅✅
Snowflake (snowflake)✅✅✅
Text Completion Codestral (text-completion-codestral)✅✅✅
Text Completion OpenAI (text-completion-openai)✅✅✅✅✅✅✅
Together AI (together_ai)✅✅✅
Topaz (topaz)✅✅✅
Triton (triton)✅✅✅
V0 (v0)✅✅✅
Vercel AI Gateway (vercel_ai_gateway)✅✅✅
VLLM (vllm)✅✅✅
Volcengine (volcengine)✅✅✅
Voyage AI (voyage)✅
WandB Inference (wandb)✅✅✅
Watsonx Text (watsonx_text)✅✅✅
xAI (xai)✅✅✅
Xinference (xinference)✅

Read the Docs


This README has been shortened. The full version is on GitHub. Read the original ↗