← プロジェクト一覧に戻る

LiteLLM

100社超のモデルプロバイダの前に立つ、OpenAI互換の単一インターフェース。ライブラリとしてもゲートウェイとしても使える

公式MIT
スター
57.5k
フォーク
11k
オープンIssue
4.9k
最終コミット
2026年8月28日

LiteLLMとは

ライブラリとしては、1つの関数呼び出しで対応プロバイダのどれにでも、OpenAI形式のリクエストとレスポンスのまま到達できます。多くのチームが導入する理由は後者のゲートウェイのほうです。自己ホストのサーバーが、個別の予算と流量制限を持つ仮想キーを発行し、プロバイダの障害時には代替へ切り替え、複数の配備先へ負荷を分散し、キーごとの費用を記録します。代償として、リクエスト経路に自分で運用する構成要素が1つ増えます。またenterpriseディレクトリ配下のコードはMITの対象外です。

LiteLLMで何ができますか?

  • どのプロバイダも1つの関数で呼ぶ — completionを1回呼ぶだけで、モデル名を指定してOpenAI形式の応答が返ります。OpenAI・Vertex AI・Bedrock・Anthropicなどの切り替えが、呼び出し箇所の書き換えではなく文字列の変更で済みます。
  • すべての手前にゲートウェイを置く — 自己ホストのプロキシがOpenAI互換のエンドポイントを提供します。既にOpenAI形式で話すクライアント、SDK、ツールは接続先を変えるだけで、意識せずに経路制御の恩恵を受けます。
  • 上限付きのキーを配る — 仮想キーはキー単位・チーム単位・利用者単位で予算と流量制限を持ちます。試験的な利用が部門の枠を使い切ることがなく、プロバイダの本物の認証情報はゲートウェイの外に出ません。
  • プロバイダが落ちても供給を続ける — 失敗したリクエストは自動的に別の配備先へ切り替わります。同じモデルの複数の配備先へ負荷を分散させることもでき、1か所に固定されません。
  • 支出を見て、通信を検査する — ログ、費用の集計、管理画面で、何が通り何にいくらかかったかを把握できます。同じ層のガードレールで、送信前に内容のフィルタリングや個人情報のマスキングを適用できます。

LiteLLMを選ぶ前に

  • MITの範囲はenterpriseディレクトリの手前までで、そこはライセンス文書が別条件と定めています。公開部分と商用部分が同じリポジトリに同居し、境界はパスで引かれます。
  • ゲートウェイを採用すると、すべてのモデル呼び出しの経路に自分で運用する構成要素が入ります。プロバイダ側の可用性に加えて、その稼働と更新も自分の担当になります。

スター推移

8月21日〜8月28日 · +553

56.9k57.5k

よくある質問

LiteLLMは商用利用できますか?

LiteLLMはMITライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。

LiteLLMはどの形で使えますか?

LiteLLMはセルフホスト・ローカル実行の形で利用できます。

ドキュメント

BerriAI/litellm のREADMEより転載(UNKNOWN — read the LICENSE file)。 原文を読む ↗


What is LiteLLM

LiteLLM is an open source AI Gateway that gives you a single, unified interface to call 100+ LLM providers — OpenAI, Anthropic, Gemini, Bedrock, Azure, and more — using the OpenAI format.

Use it as a Python SDK for direct library integration, or deploy the AI Gateway (Proxy Server) as a centralized service for your team or organization.

Jump to LiteLLM Proxy (LLM Gateway) Docs Jump to Supported LLM Providers


Why LiteLLM

Managing LLM calls across providers gets complicated fast — different SDKs, auth patterns, request formats, and error types for every model. LiteLLM removes that friction:

  • Unified API — one interface for 100+ LLMs, no provider-specific SDK juggling
  • Drop-in OpenAI compatibility — swap providers without rewriting your code
  • Production-ready gateway — virtual keys, spend tracking, guardrails, load balancing, and an admin dashboard out of the box
  • 8ms P95 latency at 1k RPS (benchmarks)

OSS Adopters


Features

All Supported Endpoints - /chat/completions, /responses, /embeddings, /images, /audio, /batches, /rerank, /a2a, /messages and more.

Python SDK

uv add litellm
from litellm import completion
import os

os.environ["OPENAI_API_KEY"] = "your-openai-key"
os.environ["ANTHROPIC_API_KEY"] = "your-anthropic-key"

# OpenAI
response = completion(model="openai/gpt-4o", messages=[{"role": "user", "content": "Hello!"}])

# Anthropic  
response = completion(model="anthropic/claude-sonnet-4-20250514", messages=[{"role": "user", "content": "Hello!"}])

AI Gateway (Proxy Server)

Getting Started - E2E Tutorial - Setup virtual keys, make your first request

uv tool install 'litellm[proxy]'
litellm --model gpt-4o
import openai

client = openai.OpenAI(api_key="anything", base_url="http://0.0.0.0:4000")
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello!"}]
)

Docs: LLM Providers

Supported Providers - LangGraph, Vertex AI Agent Engine, Azure AI Foundry, Bedrock AgentCore, Pydantic AI

Python SDK - A2A Protocol

from litellm.a2a_protocol import A2AClient
from a2a.types import SendMessageRequest, MessageSendParams
from uuid import uuid4

client = A2AClient(base_url="http://localhost:10001")

request = SendMessageRequest(
    id=str(uuid4()),
    params=MessageSendParams(
        message={
            "role": "user",
            "parts": [{"kind": "text", "text": "Hello!"}],
            "messageId": uuid4().hex,
        }
    )
)
response = await client.send_message(request)

AI Gateway (Proxy Server)

Step 1. Add your Agent to the AI Gateway — set protocolVersion to 1.0 or 0.3 per agent

Step 2. Call Agent via A2A SDK (requires a2a-sdk>=1.1.0)

import httpx
from a2a.client import A2ACardResolver, ClientConfig, ClientFactory
from a2a.types import Message, Part, Role, SendMessageRequest
from a2a.utils.constants import TransportProtocol
from uuid import uuid4

base_url = "http://localhost:4000/a2a/my-agent"  # LiteLLM proxy + agent name
headers = {"Authorization": "Bearer sk-1234"}    # LiteLLM Virtual Key

async with httpx.AsyncClient(headers=headers, timeout=60.0) as http_client:
    resolver = A2ACardResolver(httpx_client=http_client, base_url=base_url)
    agent_card = await resolver.get_agent_card()
    config = ClientConfig(
        httpx_client=http_client,
        streaming=False,
        supported_protocol_bindings=[TransportProtocol.JSONRPC, TransportProtocol.HTTP_JSON],
    )
    client = ClientFactory(config).create(agent_card)

    request = SendMessageRequest(
        message=Message(
            message_id=uuid4().hex,
            role=Role.ROLE_USER,
            parts=[Part(text="Hello!")],
        )
    )
    async for event in client.send_message(request):
        populated = event.ListFields()
        if populated and populated[0][0].name in ("message", "msg"):
            print("".join(getattr(p, "text", "") or "" for p in populated[0][1].parts))

Docs: A2A Agent Gateway

Python SDK - MCP Bridge

from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
from litellm import experimental_mcp_client
import litellm

server_params = StdioServerParameters(command="python", args=["mcp_server.py"])

async with stdio_client(server_params) as (read, write):
    async with ClientSession(read, write) as session:
        await session.initialize()

        # Load MCP tools in OpenAI format
        tools = await experimental_mcp_client.load_mcp_tools(session=session, format="openai")

        # Use with any LiteLLM model
        response = await litellm.acompletion(
            model="gpt-4o",
            messages=[{"role": "user", "content": "What's 3 + 5?"}],
            tools=tools
        )

AI Gateway - MCP Gateway

Step 1. Add your MCP Server to the AI Gateway

Step 2. Call MCP tools via /chat/completions

curl -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
  -H 'Authorization: Bearer sk-1234' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Summarize the latest open PR"}],
    "tools": [{
      "type": "mcp",
      "server_url": "litellm_proxy/mcp/github",
      "server_label": "github_mcp",
      "require_approval": "never"
    }]
  }'

Use with Cursor IDE

{
  "mcpServers": {
    "LiteLLM": {
      "url": "http://localhost:4000/mcp/",
      "headers": {
        "x-litellm-api-key": "Bearer sk-1234"
      }
    }
  }
}

Docs: MCP Gateway

Supported Providers (Website Supported Models | Docs)

Provider/chat/completions/messages/responses/embeddings/image/generations/audio/transcriptions/audio/speech/moderations/batches/rerank
Abliteration (abliteration)✅
AI/ML API (aiml)✅✅✅✅✅
AI21 (ai21)✅✅✅
AI21 Chat (ai21_chat)✅✅✅
Aleph Alpha✅✅✅
Amazon Nova✅✅✅
Anthropic (anthropic)✅✅✅✅
Anthropic Text (anthropic_text)✅✅✅✅
Anyscale✅✅✅
AssemblyAI (assemblyai)✅✅✅✅
Auto Router (auto_router)✅✅✅
AWS - Bedrock (bedrock)✅✅✅✅✅
AWS - Sagemaker (sagemaker)✅✅✅✅
Azure (azure)✅✅✅✅✅✅✅✅✅
Azure AI (azure_ai)✅✅✅✅✅✅✅✅✅
Azure Text (azure_text)✅✅✅✅✅✅✅
Baseten (baseten)✅✅✅
Bytez (bytez)✅✅✅
Cerebras (cerebras)✅✅✅
Clarifai (clarifai)✅✅✅
Cloudflare AI Workers (cloudflare)✅✅✅
Codestral (codestral)✅✅✅
Cognition (cognition)✅✅✅
Cohere (cohere)✅✅✅✅✅
Cohere Chat (cohere_chat)✅✅✅
CometAPI (cometapi)✅✅✅✅
CompactifAI (compactifai)✅✅✅
Custom (custom)✅✅✅
Custom OpenAI (custom_openai)✅✅✅✅✅✅✅
Dashscope (dashscope)✅✅✅✅✅
Databricks (databricks)✅✅✅
DataRobot (datarobot)✅✅✅
Deepgram (deepgram)✅✅✅✅
DeepInfra (deepinfra)✅✅✅
Deepseek (deepseek)✅✅✅
ElevenLabs (elevenlabs)✅✅✅✅✅
Empower (empower)✅✅✅
Fal AI (fal_ai)✅✅✅✅
Featherless AI (featherless_ai)✅✅✅
Fireworks AI (fireworks_ai)✅✅✅
FriendliAI (friendliai)✅✅✅
Galadriel (galadriel)✅✅✅
GitHub Copilot (github_copilot)✅✅✅✅
GitHub Models (github)✅✅✅
Google - PaLM✅✅✅
Google - Vertex AI (vertex_ai)✅✅✅✅✅
Google AI Studio - Gemini (gemini)✅✅✅
GradientAI (gradient_ai)✅✅✅
Groq AI (groq)✅✅✅
Heroku (heroku)✅✅✅
Hosted VLLM (hosted_vllm)✅✅✅
Huggingface (huggingface)✅✅✅✅✅
Hyperbolic (hyperbolic)✅✅✅
IBM - Watsonx.ai (watsonx)✅✅✅✅
Infinity (infinity)✅
Jina AI (jina_ai)✅
Lambda AI (lambda_ai)✅✅✅
Lemonade (lemonade)✅✅✅
LiteLLM Proxy (litellm_proxy)✅✅✅✅✅
Llamafile (llamafile)✅✅✅
LM Studio (lm_studio)✅✅✅
Maritalk (maritalk)✅✅✅
Meta - Llama API (meta_llama)✅✅✅
Mistral AI API (mistral)✅✅✅✅
ModelScope (modelscope)✅✅✅✅
Moonshot (moonshot)✅✅✅
Morph (morph)✅✅✅
Nebius AI Studio (nebius)✅✅✅✅
NLP Cloud (nlp_cloud)✅✅✅
Novita AI (novita)✅✅✅
Nscale (nscale)✅✅✅
Nvidia NIM (nvidia_nim)✅✅✅
OCI (oci)✅✅✅
Ollama (ollama)✅✅✅✅
Ollama Chat (ollama_chat)✅✅✅
Oobabooga (oobabooga)✅✅✅✅✅✅✅
OpenAI (openai)✅✅✅✅✅✅✅✅✅
OpenAI-like (openai_like)✅
OpenRouter (openrouter)✅✅✅
OVHCloud AI Endpoints (ovhcloud)✅✅✅
Perplexity AI (perplexity)✅✅✅
Petals (petals)✅✅✅
Pinstripes (pinstripes)✅✅✅
Predibase (predibase)✅✅✅
Recraft (recraft)✅
Replicate (replicate)✅✅✅
Sagemaker Chat (sagemaker_chat)✅✅✅
Sambanova (sambanova)✅✅✅
Snowflake (snowflake)✅✅✅
Text Completion Codestral (text-completion-codestral)✅✅✅
Text Completion OpenAI (text-completion-openai)✅✅✅✅✅✅✅
Together AI (together_ai)✅✅✅
Topaz (topaz)✅✅✅
Triton (triton)✅✅✅
V0 (v0)✅✅✅
Vercel AI Gateway (vercel_ai_gateway)✅✅✅
VLLM (vllm)✅✅✅
Volcengine (volcengine)✅✅✅
Voyage AI (voyage)✅
WandB Inference (wandb)✅✅✅
Watsonx Text (watsonx_text)✅✅✅
xAI (xai)✅✅✅
Xinference (xinference)✅

Read the Docs


このREADMEは一部を省略しています。全文はGitHubにあります。 原文を読む ↗