
Phoenix
Tracing and evaluation that runs inside a notebook
Overview
Built on OpenTelemetry, so traces are portable rather than locked to one vendor, and it will run locally next to the code you are debugging. Licensed under Elastic 2.0 — fine for internal use, restrictive if you intend to offer it as a managed service.
What can you do with Phoenix?
- Instrument an app automatically —
npx @arizeai/phoenix-cli setup(orpx setuponce Phoenix is installed) detects your framework and LLM provider, installs the matching OpenInference instrumentation and wires up trace export. - Run the platform locally —
uvx arize-phoenix servebrings up the full platform with nothing installed, and the same build ships as a Docker Hub image and a Helm chart for cluster deployment. - Replay a captured LLM call — The Playground reruns a traced call with a different prompt, model or parameters, and prompt management keeps those changes under version control with tagging.
- Measure changes as experiments —
arize-phoenix-evalsscores response and retrieval quality against versioned datasets, so a prompt or retrieval change is compared as a tracked experiment rather than by eye. - Query traces from a coding agent — The remote MCP server built into Phoenix exposes a
/mcpendpoint that Claude Code, Cursor and other MCP clients can use to read traces, datasets and experiments.
Documentation
Reproduced from the Arize-ai/phoenix README, published under Elastic-2.0. Read the original ↗
Phoenix is an open-source AI observability platform designed for experimentation, evaluation, and troubleshooting. It provides:
- Tracing - Trace your LLM application’s runtime using OpenTelemetry-based instrumentation.
- Evaluation - Leverage LLMs to benchmark your application’s performance using response and retrieval evals.
- Datasets - Create versioned datasets of examples for experimentation, evaluation, and fine-tuning.
- Experiments - Track and evaluate changes to prompts, LLMs, and retrieval.
- Playground- Optimize prompts, compare models, adjust parameters, and replay traced LLM calls.
- Prompt Management- Manage and test prompt changes systematically using version control, tagging, and experimentation.
- PXI (Phoenix Intelligence) - An AI engineering agent built into Phoenix for debugging traces, iterating on prompts, and navigating the product.
- Remote MCP Server - Connect Claude Code, Cursor, and other MCP clients directly to your Phoenix instance’s
/mcpendpoint to query traces, datasets, experiments, and more.
Phoenix is vendor and language agnostic with out-of-the-box support for popular frameworks (OpenAI Agents SDK, Claude Agent SDK, LangGraph, Vercel AI SDK, Mastra, CrewAI, LlamaIndex, DSPy) and LLM providers (OpenAI, Anthropic, Google GenAI, Google ADK, AWS Bedrock, OpenRouter, LiteLLM, and more). For details on auto-instrumentation, check out the OpenInference project.
Phoenix runs practically anywhere, including your local machine, a containerized deployment, or in the cloud. See Environments for a walkthrough of each option, or jump straight into the Tracing Quickstart.
Table of Contents
- Run Locally
- Trace Your Application
- Deploy
- Packages
- Tracing Integrations
- Sandboxes
- For Humans and Coding Agents
- Security & Privacy
- Community
Run Locally
Install Phoenix via pip or conda and have a fully functional Phoenix. For all installation and hosting options, see the install guide.
pip install arize-phoenix
phoenix serve
Or run it with no install using uvx:
uvx arize-phoenix serve
Trace Your Application
The fastest way to send traces is to let your coding agent (Claude Code, Codex, Cursor, and others) instrument your app. From your project directory, run:
npx @arizeai/phoenix-cli setup
# or, with Phoenix installed: px setup
Setup detects your framework and LLM provider, installs the right OpenInference instrumentation, and wires up trace export. Prefer to wire it up in code? See the tracing documentation.
Deploy
Phoenix container images are available via Docker Hub and can be deployed using Docker or Kubernetes via the Helm chart.
For Docker Compose, Kubernetes/Helm, and other deployment options, see the self-hosting documentation.
[!NOTE] The Google Cloud button builds Phoenix from source in Cloud Shell rather than deploying the prebuilt Docker Hub image. The Azure template serves plain HTTP (Azure Container Instances does not terminate TLS) — front it with a TLS proxy such as an Application Gateway before production use.
Packages
The arize-phoenix package includes the entire Phoenix platform. However, if you have deployed the Phoenix platform, there are lightweight Python sub-packages and TypeScript packages that can be used in conjunction with the platform.
Python Subpackages
| Package | Version & Docs | Description |
|---|---|---|
| arize-phoenix-otel | Provides a lightweight wrapper around OpenTelemetry primitives with Phoenix-aware defaults | |
| arize-phoenix-client | Lightweight client for interacting with the Phoenix server via its OpenAPI REST interface | |
| arize-phoenix-evals | Tooling to evaluate LLM applications including RAG relevance, answer relevance, and more |
TypeScript Subpackages
| Package | Version & Docs | Description |
|---|---|---|
| @arizeai/phoenix-otel | Provides a lightweight wrapper around OpenTelemetry primitives with Phoenix-aware defaults | |
| @arizeai/phoenix-client | Client for the Arize Phoenix API | |
| @arizeai/phoenix-evals | TypeScript evaluation library for LLM applications (alpha release) | |
| @arizeai/phoenix-mcp | Standalone stdio MCP server for older Phoenix versions (maintenance mode — superseded by the remote MCP server built into Phoenix) | |
| @arizeai/phoenix-cli | CLI for fetching traces, datasets, and experiments for use with Claude Code, Cursor, and other coding agents |