← プロジェクト一覧に戻る

vLLM

OpenAI互換APIを備えた高スループット推論サーバ

公式Apache-2.0
スター
89.1k
フォーク
20.7k
オープンIssue
6.7k
最終コミット
2026年8月15日

概要

PagedAttentionと継続バッチングにより、素朴な実装よりはるかに多くの同時リクエストを1枚のGPUで捌けます。エージェントは呼び出し回数が多くなりがちなため、この差が効いてきます。OpenAI互換エンドポイントを備えており、多くのエージェントフレームワークはベースURLの変更だけで接続できます。一方でGPUメモリの見積もりやモデルのロード時間など、運用面の作業は相応に発生します。

vLLMで何ができますか?

  • GPU1枚あたりの同時処理数PagedAttentionと継続バッチングにより、素朴な逐次処理よりはるかに多くのリクエストを1枚のGPUで同時に捌けます。
  • 既存クライアントをそのまま接続OpenAI互換のAPIサーバを備えるため、多くのエージェントフレームワークはベースURLの変更だけで接続先を切り替えられます。Anthropic Messages APIとgRPCにも対応します。
  • 手持ちのハードウェアで動かすuv pip install vllmで導入でき、NVIDIA・AMD・IntelのGPUに加えてx86/ARM/PowerPCのCPUでも動作します。Google TPU、Intel Gaudi、Huawei Ascend、Apple Siliconなどはハードウェアプラグインとして提供されます。
  • 量子化してカードに載せるFP8、MXFP8/MXFP4、NVFP4、INT8、INT4、GPTQ/AWQ、GGUF、compressed-tensorsといった量子化形式のまま配信できるため、扱えるモデルの上限がカードのメモリ容量だけで決まるわけではありません。
  • 出力形式の固定とツール呼び出し構造化出力の生成にはxgrammarとguidanceを利用でき、ツール呼び出しや推論過程を切り出すパーサも同梱されています。

ドキュメント

vllm-project/vllm のREADMEより転載(Apache-2.0)。 原文を読む ↗

🔥 We have built a vLLM website to help you get started with vLLM. Please visit vllm.ai to learn more. For events, please visit vllm.ai/events to join us.


About

vLLM is a fast and easy-to-use library for LLM inference and serving.

Originally developed in the Sky Computing Lab at UC Berkeley, vLLM has grown into one of the most active open-source AI projects built and maintained by a diverse community of many dozens of academic institutions and companies from over 2000 contributors.

vLLM is fast with:

  • State-of-the-art serving throughput
  • Efficient management of attention key and value memory with PagedAttention
  • Continuous batching of incoming requests, chunked prefill, prefix caching
  • Fast and flexible model execution with piecewise and full CUDA/HIP graphs
  • Quantization: FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF, compressed-tensors, ModelOpt, TorchAO, and more
  • Optimized attention kernels including FlashAttention, FlashInfer, TRTLLM-GEN, FlashMLA, and Triton
  • Optimized GEMM/MoE kernels for various precisions using CUTLASS, TRTLLM-GEN, CuTeDSL
  • Speculative decoding including n-gram, suffix, EAGLE, DFlash
  • Automatic kernel generation and graph-level transformations using torch.compile
  • Disaggregated prefill, decode, and encode

vLLM is flexible and easy to use with:

  • Seamless integration with popular Hugging Face models
  • High-throughput serving with various decoding algorithms, including parallel sampling, beam search, and more
  • Tensor, pipeline, data, expert, and context parallelism for distributed inference
  • Streaming outputs
  • Generation of structured outputs using xgrammar or guidance
  • Tool calling and reasoning parsers
  • OpenAI-compatible API server, plus Anthropic Messages API and gRPC support
  • Efficient multi-LoRA support for dense and MoE layers
  • Support for NVIDIA GPUs, AMD GPUs, Intel GPUs, and x86/ARM/PowerPC CPUs. Additionally, diverse hardware plugins such as Google TPUs, Intel Gaudi, IBM Spyre, Huawei Ascend, Rebellions NPU, Apple Silicon, MetaX GPU, and more.

vLLM seamlessly supports 200+ model architectures on Hugging Face, including:

  • Decoder-only LLMs (e.g., Llama, Qwen, Gemma)
  • Mixture-of-Expert LLMs (e.g., Mixtral, DeepSeek-V3, Qwen-MoE, GPT-OSS)
  • Hybrid attention and state-space models (e.g., Mamba, Qwen3.5)
  • Multi-modal models (e.g., LLaVA, Qwen-VL, Pixtral)
  • Embedding and retrieval models (e.g., E5-Mistral, GTE, ColBERT)
  • Reward and classification models (e.g., Qwen-Math)

Find the full list of supported models here.

Getting Started

Install vLLM with uv (recommended) or pip:

uv pip install vllm

Or build from source for development.

Visit our documentation to learn more.

Contact Us

  • For technical questions and feature requests, please use GitHub Issues
  • For discussing with fellow users, please use the vLLM Forum
  • For coordinating contributions and development, please use Slack
  • For security disclosures, please use GitHub’s Security Advisories feature
  • For collaborations and partnerships, please contact us at collaboration@vllm.ai

Media Kit