← Back to all projects

MLflow

The experiment tracker teams already run, extended to hold agent traces, judges and prompt versions too

Apache-2.0
Stars
27.7k
Forks
6.2k
Open issues
2.1k
Last commit
27 Aug 2026

What is MLflow?

MLflow's usefulness for agents comes from where it already is: many organisations run it for model training, and the GenAI half puts agent traces, evaluation runs and prompt versions in that same experiment view. Instrumentation is automatic for a long list of agent libraries — one line and every prompt, retrieval, tool call and response is recorded, in OpenTelemetry's GenAI conventions so the traces are readable by other tools too. From there the recorded runs become evaluation sets, scored by built-in judges or ones you write. The cost is that MLflow is a platform, not a library: even for traces alone you are running a tracking server with a database and an artifact store behind it.

What can you do with MLflow?

  • Record a run without instrumenting it — Automatic instrumentation covers a long list of agent frameworks and model providers, so the prompts, retrievals and tool calls appear from one line of setup.
  • Traces other tools can read — The recorded spans follow OpenTelemetry's conventions for generative AI, so they can be exported into an existing observability stack instead of staying in a silo.
  • Turn recorded runs into an evaluation set — Real traffic becomes the test cases, which is the only reliable way to find out whether a change helped on the inputs users actually send.
  • Score with judges, built-in or your own — Ready-made judges cover the common failures such as hallucination and irrelevance; a custom scorer encodes what "good" means for your own task.
  • Version prompts with the runs that used them — Prompt templates are stored, compared and traced back to the runs they produced, so a regression can be tied to the edit that caused it.

Before you choose MLflow

  • MLflow is a platform, not a client library. Even if all you want is traces, you take on a tracking server with a backing database and an artifact store, and their upgrades.
  • The documentation covers the open project and Databricks' managed version together, so check which of the monitoring and governance features apply to a self-hosted install before planning around them.

Frequently asked questions

Is MLflow free for commercial use?

MLflow is released under the Apache-2.0 licence — OSI-approved open source, which permits commercial use.

How can MLflow be deployed?

MLflow is available as Self-hosted / Runs locally / Managed cloud.

Documentation

Reproduced from the mlflow/mlflow README, published under Apache-2.0. Read the original ↗

MLflow is the largest open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data. With over 60 million monthly downloads, thousands of organizations rely on MLflow each day to ship AI to production with confidence.

MLflow’s comprehensive feature set for agents and LLM applications includes production-grade observability, evaluation, prompt management, prompt optimization and an AI Gateway for managing costs and model access. Learn more at MLflow for LLMs and Agents.

Get Started in 3 Simple Steps

From zero to full-stack LLMOps in minutes. No complex setup or major code changes required. Get Started →

Fastest start — set up tracing with our CLI

uvx mlflow@latest agent setup

One command installs the MLflow skills and launches your coding agent of choice to add tracing to your app. Prefer to wire it up yourself? Follow the three steps below.

1. Start MLflow Server

uvx mlflow server

2. Enable Logging

import mlflow

mlflow.set_tracking_uri("http://localhost:5000")
mlflow.openai.autolog()

3. Run Your Code

from openai import OpenAI

client = OpenAI()
client.responses.create(
    model="gpt-5.4-mini",
    input="Hello!",
)

Explore traces and metrics in the MLflow UI at http://localhost:5000.

LLMs & Agents

MLflow provides everything you need to build, debug, evaluate, and deploy production-quality LLM applications and AI agents. Supports Python, TypeScript/JavaScript, Java and any other programming language. MLflow also natively integrates with OpenTelemetry and MCP.

Model Training

For machine learning and deep learning model development, MLflow provides a full suite of tools to manage the ML lifecycle:

  • Experiment Tracking — Track models, parameters, metrics, and evaluation results across experiments
  • Model Evaluation — Automated evaluation tools integrated with experiment tracking
  • Model Registry — Collaboratively manage the full lifecycle of ML models
  • Deployment — Deploy models to batch and real-time scoring on Docker, Kubernetes, Azure ML, AWS SageMaker, and more

Learn more at MLflow for Model Training.

Integrations

MLflow supports all agent frameworks, LLM providers, tools, and programming languages. We offer one-line automatic tracing for more than 60 frameworks. See the full integrations list.

OpenTelemetry

Agent Frameworks (Python)

Agent Frameworks (TypeScript)

Agent Frameworks (Java)

Model Providers

This README has been shortened. The full version is on GitHub. Read the original ↗