
MLflow
多くの組織がすでに動かしている実験管理基盤に、エージェントの実行記録・評価・プロンプトの版を同じ画面で載せる
MLflowとは
MLflowがエージェント向けに効いてくるのは、すでにそこにあるからです。モデルの学習管理として動かしている組織が多く、GenAI向けの機能はエージェントの実行記録、評価の実行、プロンプトの版を、その同じ実験画面に並べます。計測は多数のエージェント関連ライブラリに対して自動で行われ、1行足すだけでプロンプト、検索、ツール呼び出し、応答が記録されます。形式はOpenTelemetryのGenAI規約に沿うため、他の道具からも読めます。記録された実行はそのまま評価用のデータになり、既製の判定役や自作の判定役で採点できます。難点は、MLflowがライブラリではなく基盤である点です。実行記録だけが目的でも、データベースと成果物置き場を伴う管理サーバーを動かすことになります。
MLflowで何ができますか?
- 計測コードを書かずに実行を記録する — 多数のエージェントフレームワークとモデル提供元に自動計測が対応しており、設定を1行足すだけでプロンプトや検索、ツール呼び出しが記録されます。
- 他の道具でも読める形式で残す — 記録はOpenTelemetryの生成AI向け規約に沿うため、既存の監視基盤へ書き出せます。この製品の中だけに閉じません。
- 記録した実行を評価データに変える — 実際の利用がそのまま試験項目になります。利用者が本当に送ってくる入力に対して改善したかを確かめる方法は、これ以外にありません。
- 既製または自作の判定役で採点する — 事実と異なる出力や的外れな回答など、よくある失敗には既製の判定役があります。自作の採点基準を書けば、その業務にとっての良さを定義できます。
- プロンプトの版を実行と結び付けて保存する — プロンプトの雛形を保存・比較し、それを使った実行までたどれます。品質が落ちたとき、原因となった編集を特定できます。
MLflowを選ぶ前に
- MLflowはクライアントライブラリではなく基盤です。実行記録だけが目的でも、データベースと成果物置き場を伴う管理サーバーとその更新作業を抱えることになります。
- 資料はオープンな本体とDatabricksのマネージド版をまとめて説明しています。監視や統制の機能を前提に設計する前に、自社運用でも使えるかを確認してください。
よくある質問
MLflowは商用利用できますか?
MLflowはApache-2.0ライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。
MLflowはどの形で使えますか?
MLflowはセルフホスト・ローカル実行・マネージドクラウドの形で利用できます。
ドキュメント
mlflow/mlflow のREADMEより転載(Apache-2.0)。 原文を読む ↗
MLflow is the largest open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data. With over 60 million monthly downloads, thousands of organizations rely on MLflow each day to ship AI to production with confidence.
MLflow’s comprehensive feature set for agents and LLM applications includes production-grade observability, evaluation, prompt management, prompt optimization and an AI Gateway for managing costs and model access. Learn more at MLflow for LLMs and Agents.
Get Started in 3 Simple Steps
From zero to full-stack LLMOps in minutes. No complex setup or major code changes required. Get Started →
Fastest start — set up tracing with our CLI
uvx mlflow@latest agent setupOne command installs the MLflow skills and launches your coding agent of choice to add tracing to your app. Prefer to wire it up yourself? Follow the three steps below.
1. Start MLflow Server
uvx mlflow server
2. Enable Logging
import mlflow
mlflow.set_tracking_uri("http://localhost:5000")
mlflow.openai.autolog()
3. Run Your Code
from openai import OpenAI
client = OpenAI()
client.responses.create(
model="gpt-5.4-mini",
input="Hello!",
)
Explore traces and metrics in the MLflow UI at http://localhost:5000.
LLMs & Agents
MLflow provides everything you need to build, debug, evaluate, and deploy production-quality LLM applications and AI agents. Supports Python, TypeScript/JavaScript, Java and any other programming language. MLflow also natively integrates with OpenTelemetry and MCP.
Model Training
For machine learning and deep learning model development, MLflow provides a full suite of tools to manage the ML lifecycle:
- Experiment Tracking — Track models, parameters, metrics, and evaluation results across experiments
- Model Evaluation — Automated evaluation tools integrated with experiment tracking
- Model Registry — Collaboratively manage the full lifecycle of ML models
- Deployment — Deploy models to batch and real-time scoring on Docker, Kubernetes, Azure ML, AWS SageMaker, and more
Learn more at MLflow for Model Training.
Integrations
MLflow supports all agent frameworks, LLM providers, tools, and programming languages. We offer one-line automatic tracing for more than 60 frameworks. See the full integrations list.
OpenTelemetry
Agent Frameworks (Python)
Agent Frameworks (TypeScript)
Agent Frameworks (Java)
Model Providers
このREADMEは一部を省略しています。全文はGitHubにあります。 原文を読む ↗