← プロジェクト一覧に戻る

LocalAI

OpenAI APIそっくりの入口をセルフホストで立て、裏の推論エンジンを差し替えられるサーバー。会話から音声・動画まで模します

MIT
スター
48.7k
フォーク
4.4k
オープンIssue
219
最終コミット
2026年8月28日

LocalAIとは

OpenAI向けのクライアントをそのまま向ければ応答します。ツール呼び出し付きのチャット、埋め込み、画像生成、文字起こしと読み上げ、リランカー、リアルタイム音声対話までで、AnthropicやElevenLabs、OllamaのAPIの模倣も持っています。小さな本体がハードウェアを検出し、llama.cpp・vLLM・MLX・whisper.cppといった推論バックエンドを必要になった時点で取得します。Webのギャラリーからモデルをクリックで導入でき、MCP対応の自律エージェントも同梱、APIキー・OIDC・クォータで複数ユーザーにも対応します。正直に比較すると、ノートPCでモデルと会話するだけならOllamaの方が簡単で、1モデルのスループットを追い込むなら素のvLLMの方が身軽です。そしてWindowsのネイティブ版はなく、コンテナかWSLに限られます。

LocalAIで何ができますか?

  • チャットだけではないOpenAI互換 — ツール呼び出しと文法制約付きのチャット、埋め込み、画像生成、文字起こし、読み上げ、リランキング、WebSocket越しのリアルタイム音声対話まで揃い、どのAPI向けに書かれたクライアントにも1つのデプロイで応答します。AnthropicやElevenLabs、OllamaのAPIも話せます。
  • バックエンドは必要になってから、機材に合わせて届く — 本体は小さなバイナリまたはコンテナで、各推論エンジンは別イメージとしてモデルが要求した時点で取得されます。GPUの検出は自動だとプロジェクトは述べています。CUDA・ROCm・Intel oneAPI・Vulkan・Apple Metal・Jetsonに個別のビルドがあります。
  • モデルはギャラリーから、あるいはどこからでも — Web UIのギャラリーにはおよそ1,700件の設定が索引され、クリックで導入できます。CLIでも同じ名前が使え、Hugging Face・Ollama・OCIレジストリからURI指定で直接取得することもできます。
  • エージェントと複数ユーザーの仕組みを同梱 — 組み込まれたLocalAGI層により、Web UIかJSON設定からツール・エージェント別の文書コレクション・MCPサーバーを持つエージェントを作れます。スキルは明示的に有効化するまで動きません。共有環境向けにAPIキー、OIDC、クォータ、ロール別権限も備わりました。
  • 1台に収まらなくなったときの道が2つ — P2Pのフェデレーションは複数インスタンスへ負荷を分散し、ワーカーモードは1つのllama.cppモデルの重みを複数台に分割します(対象は単一モデルのみ、推論開始後のワーカー追加は不可、と制限も明記)。別系統として、空きVRAMを見てルーティングする分散モードもあります。

LocalAIを選ぶ前に

  • Windowsのネイティブ版はなく、コンテナかWSLが前提です。この要望は2024年5月からトラッカーで最多リアクションの未解決Issueであり続けています。報告#2368
  • 単一モデルのスループットを追い込む用途では、これはエンジンではなく管理層です。裏でvLLMを動かすこともできますが、それだけが目的なら直接vLLMを立てる方が無駄がありません。
  • 開発は創始者にほぼ集中しており、コミット数は次点の20倍を超えます。現在は2人目のメンテナーが明記されるようになりました。

スター推移

8月19日〜8月28日 · +147

48.6k48.7k

よくある質問

LocalAIは商用利用できますか?

LocalAIはMITライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。

LocalAIはどの形で使えますか?

LocalAIはローカル実行・セルフホストの形で利用できます。

ドキュメント

mudler/LocalAI のREADMEより転載(MIT)。 原文を読む ↗

LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

A small core, not a bundle. Each backend wraps a best-in-class engine (llama.cpp, vLLM, whisper.cpp, stable-diffusion, MLX…) in its own image, pulled only when a model needs it. You install nothing you don’t use.

  • Composable by design: backends are separate and pulled on demand, so you install only what your model needs
  • Open and extensible: load any model, or build your own backend in any language against an open interface
  • Drop-in API compatibility: OpenAI, Anthropic, and ElevenLabs APIs across every backend
  • Any model, any modality: LLMs, vision, voice, image, and video behind one API
  • Any hardware: NVIDIA, AMD, Intel, Apple Silicon, Vulkan, or CPU-only
  • Multi-user ready: API key auth, user quotas, role-based access
  • Built-in AI agents: autonomous agents with tool use, RAG, MCP, and skills
  • Privacy-first: your data never leaves your infrastructure

A small LocalAI core with backends (llama.cpp, vLLM, MLX, whisper.cpp, stable-diffusion, kokoro, parakeet.cpp...) plugged in as separate on-demand images

Created by Ettore Di Giacinto and maintained by the LocalAI team.

:book: Documentation | :speech_balloon: Discord | 💻 Quickstart | 🖼️ Models | ❓FAQ

Guided tour

https://github.com/user-attachments/assets/08cbb692-57da-48f7-963d-2e7b43883c18

User and auth

https://github.com/user-attachments/assets/228fa9ad-81a3-4d43-bfb9-31557e14a36c

Agents

https://github.com/user-attachments/assets/6270b331-e21d-4087-a540-6290006b381a

Usage metrics per user

https://github.com/user-attachments/assets/cbb03379-23b4-4e3d-bd26-d152f057007f

Fine-tuning and Quantization

https://github.com/user-attachments/assets/5ba4ace9-d3df-4795-b7d4-b0b404ea71ee

WebRTC

https://github.com/user-attachments/assets/ed88e34c-fed3-4b83-8a67-4716a9feeb7b

Quickstart

macOS

Note: The DMG is not signed by Apple. After installing, run: sudo xattr -d com.apple.quarantine /Applications/LocalAI.app. See #6268 for details.

Containers (Docker, podman, …)

Already ran LocalAI before? Use docker start -i local-ai to restart an existing container.

CPU only:

docker run -ti --name local-ai -p 8080:8080 localai/localai:latest

NVIDIA GPU:

# CUDA 13
docker run -ti --name local-ai -p 8080:8080 --gpus all localai/localai:latest-gpu-nvidia-cuda-13

# CUDA 12
docker run -ti --name local-ai -p 8080:8080 --gpus all localai/localai:latest-gpu-nvidia-cuda-12

# NVIDIA Jetson ARM64 (CUDA 12, for AGX Orin and similar)
docker run -ti --name local-ai -p 8080:8080 --gpus all localai/localai:latest-nvidia-l4t-arm64

# NVIDIA Jetson ARM64 (CUDA 13, for DGX Spark)
docker run -ti --name local-ai -p 8080:8080 --gpus all localai/localai:latest-nvidia-l4t-arm64-cuda-13

AMD GPU (ROCm):

docker run -ti --name local-ai -p 8080:8080 --device=/dev/kfd --device=/dev/dri --group-add=video localai/localai:latest-gpu-hipblas

Intel GPU (oneAPI):

docker run -ti --name local-ai -p 8080:8080 --device=/dev/dri/card1 --device=/dev/dri/renderD128 localai/localai:latest-gpu-intel

Vulkan GPU:

docker run -ti --name local-ai -p 8080:8080 localai/localai:latest-gpu-vulkan

Loading models

# From the model gallery (see available models with `local-ai models list` or at https://models.localai.io)
local-ai run llama-3.2-1b-instruct:q4_k_m
# From Huggingface
local-ai run huggingface://TheBloke/phi-2-GGUF/phi-2.Q8_0.gguf
# From the Ollama OCI registry
local-ai run ollama://gemma:2b
# From a YAML config
local-ai run https://gist.githubusercontent.com/.../phi-2.yaml
# From a standard OCI registry (e.g., Docker Hub)
local-ai run oci://localai/phi-2:latest

To work with a running LocalAI server from the terminal, start the built-in agent from another shell. It answers questions, reads your files and runs commands on your machine, asking you to approve anything that changes state. Inside a session, /models lists installed models and /model <name> switches between them. See the Terminal agent docs.

# Terminal 1
local-ai run llama-3.2-1b-instruct:q4_k_m

# Terminal 2
local-ai chat --model llama-3.2-1b-instruct:q4_k_m

Automatic Backend Detection: LocalAI automatically detects your GPU capabilities and downloads the appropriate backend. For advanced options, see GPU Acceleration.

For more details, see the Getting Started guide.

Latest News

For older news and full release notes, see GitHub Releases and the blog.

Features

Supported Backends & Acceleration

LocalAI supports 60+ backends including llama.cpp, vLLM, SGLang, transformers, whisper.cpp, diffusers, MLX, MLX-VLM, and many more. Hardware acceleration is available for NVIDIA (CUDA 12/13), AMD (ROCm), Intel (oneAPI/SYCL), Apple Silicon (Metal), Vulkan, and NVIDIA Jetson (L4T). All backends can be installed on-the-fly from the Backend Gallery.

See the full Backend & Model Compatibility Table and GPU Acceleration guide.

Backends built by us

Most backends wrap a best-in-class upstream engine. A handful of them are native C/C++/GGML engines (no Python at inference) developed and maintained by the LocalAI project itself:

BackendWhat it does
vllm.cppFrom-scratch C++20 port of vLLM for text generation: paged KV cache, continuous batching, prefix caching, safetensors + GGUF loading, engine-enforced structured output, on CPU, CUDA, Metal and Vulkan. Also serves MiniMax-H3 joint video+audio generation
parakeet.cppC++/GGML port of NVIDIA NeMo Parakeet ASR (tdt/ctc/rnnt/hybrid), with cache-aware streaming transcription
moss-transcribe.cppC++/GGML port of OpenMOSS MOSS-Transcribe-Diarize: joint long-form transcription, speaker diarization and timestamping in a single pass
moss-tts.cppC++/GGML port of the OpenMOSS MOSS-TTS family: text-to-speech (MOSS-TTS-Local v1.5, 48 kHz stereo) with reference-audio voice cloning, through the MOSS-Audio-Tokenizer neural codec
magpie-tts.cppC++/GGML port of NVIDIA’s Magpie TTS Multilingual 357M: 22.05 kHz mono text-to-speech in 5 voices and 9+ languages, with the NanoCodec neural codec and tokenizer/G2P embedded in a single GGUF
ced.cppC++/GGML port of the CED audio-tagging models: sound-event classification (527-class AudioSet) over REST and the realtime API for live recognition
voice-detect.cppSpeaker recognition and voice analysis (ECAPA-TDNN, WeSpeaker, ERes2Net, CAM++, wav2vec2 age/gender/emotion), replacing the Python speaker-recognition backend
voxtral-tts.cMistral Voxtral-4B-TTS text-to-speech in pure C: 20 preset voices across 9 languages, 24 kHz WAV output, no dependencies beyond libc
vibevoice.cppNative port of Microsoft VibeVoice for TTS (voice cloning) and long-form ASR with speaker diarization
rf-detr.cppNative RF-DETR object detection and instance segmentation
locate-anything.cppOpen-vocabulary object detection and visual grounding (LocateAnything-3B)
depth-anything.cppDepth Anything 3 monocular metric depth + camera pose estimation
face-detect.cppFace detection, recognition, demographics and anti-spoofing (SCRFD/ArcFace, YuNet/SFace), replacing the Python insightface backend
free-splatter.cppPose-free 3D reconstruction (FreeSplatter): turns a handful of plain photos into 3D Gaussians, no camera poses or GPU required
trellis2.cppC++/GGML port of Microsoft TRELLIS.2: single-image to textured 3D mesh (GLB with PBR materials)
privacy-filter.cppStandalone GGML PII/NER token-classification engine powering LocalAI’s PII redaction tier
LocalVQEJoint acoustic echo cancellation, noise suppression, and dereverberation
local-storeLocal-first vector database for embeddings (shipped in-tree)

We also maintain apex-quant, a per-tensor, per-layer quantization recipe for Mixture-of-Experts models that exploits their structural sparsity to produce GGUFs matching or beating Q8_0 quality - and they run out of the box on stock llama.cpp.

Resources

Team

LocalAI is maintained by a small team of humans, together with the wider community of contributors.

A huge thank you to everyone who contributes code, reviews PRs, files issues, and helps users in Discord — LocalAI is a community-driven project and wouldn’t exist without you. See the full contributors list.

Individual sponsors

A special thanks to individual sponsors, a full list is on GitHub and buymeacoffee. Special shout out to drikster80 for being generous. Thank you everyone!