Observability & Evals

Seeing what an agent actually did — traces, costs, latencies — and turning those runs into evaluation sets. Without this, an agent in production is a black box you cannot improve.

19 projects

LiteLLMModel Serving

The library half lets one function call reach any supported provider in the OpenAI request and response shape. The gateway half is the reason most teams arrive: a self-hosted server that issues virtual keys with their own budgets and rate limits, falls back when a provider fails, spreads load across deployments, and records cost per key. The price is an extra component in the request path that you now operate, and code under the enterprise directory is not covered by the MIT grant.

OfficialMIT57.5k+553
ClickHouseObservability

A column-oriented engine for append-heavy, high-cardinality data, which is the shape agent telemetry takes: one row per model call, filtered later by model, cost or error. Langfuse moved its traces here from Postgres in December 2024 and ClickHouse acquired the project in January 2026, so a self-hosted observability stack increasingly means running this underneath. It is the wrong choice for anything transactional — updates and deletes rewrite whole column parts rather than edit rows — so treat it as where agent runs land, not as the database your application writes state to.

OfficialApache-2.049.5k+133
LangfuseObservability

Records what an agent actually did — every call, tool invocation and cost — and turns those traces into evaluation datasets. The managed service runs the same codebase you can self-host with Docker Compose, so migrating in either direction stays cheap. Some enterprise features sit outside the MIT core; check which tier you need before committing.

OfficialMIT33.9k+344
MLflowObservability

MLflow's usefulness for agents comes from where it already is: many organisations run it for model training, and the GenAI half puts agent traces, evaluation runs and prompt versions in that same experiment view. Instrumentation is automatic for a long list of agent libraries — one line and every prompt, retrieval, tool call and response is recorded, in OpenTelemetry's GenAI conventions so the traces are readable by other tools too. From there the recorded runs become evaluation sets, scored by built-in judges or ones you write. The cost is that MLflow is a platform, not a library: even for traces alone you are running a tracking server with a database and an artifact store behind it.

Apache-2.027.7k
MastraFrameworks

Agents are declared with instructions and a model, tools are validated on both ends by a schema, and workflows are chained from steps that can suspend mid-run and resume later because their execution state is written to storage. The breadth is the appeal — retrieval, memory, evaluation and tracing are in the same package rather than assembled from four — but breadth also means a large surface to learn, and the Apache-2.0 grant stops short of a few directories the project reserves under a separate enterprise licence.

OfficialApache-2.027.5k+195
promptfooObservability

Test cases are YAML: a prompt, some inputs, and assertions about the output. That is enough to produce a side-by-side matrix comparing several prompts or models on the same inputs, and to fail a build when a change makes things worse. The same tool also attacks the application, generating adversarial inputs to surface vulnerabilities and compliance risks. It works from the outside — prompt in, answer out — so when the question is which step inside a running agent went wrong, this is not the instrument for it.

OfficialMIT24.6k+208
OpikObservability

A decorator on a function or a wrapper around a client is enough to start recording traces, and from there the same data feeds experiments that score responses automatically instead of by hand — hallucination, answer relevance, context recall and the rest. More than thirty integrations cover the usual frameworks and providers, and the whole platform can run locally or on Kubernetes rather than only as a hosted service. Its difficulty is not capability but overlap: it competes directly with two platforms already in this catalogue, so the choice turns on integrations and hosting rather than on what is missing.

OfficialApache-2.021.7k+135
skillsTools & Integrations

Over a hundred skill directories of Google's own written guidance, installed selectively with `npx skills add google/skills`. The bulk sits under `skills/cloud` — GKE, BigQuery, Agent Platform, the six Well-Architected pillars — while `skills/ads` and `skills/analytics` cover the Google Ads API, the Mobile Ads and IMA SDKs and the two Google Analytics APIs. Depth is uneven: `skills/cloud/gke-*` accounts for twenty-nine directories on its own, Cloud Run and Firebase get one each, and the README labels the repository as under active development. A stack on neither Google Cloud nor Google's ads and analytics products gets nothing here, and app-side coverage stops at the Mobile Ads SDK — Android, Flutter, Dart, Genkit and ADK skills are separate repositories the README only links to.

Agent skillOfficialApache-2.018.9k+332
DeepEvalObservability

The shape is deliberately pytest's: assertions inside test functions, a runner that discovers them, parameterised cases — except the assertion is about whether an answer was faithful to its context or whether an agent finished the task. That framing is the point, because it puts quality regressions in front of the same gate that catches syntax errors. The thing to plan for is that most of these metrics are scored by a model, so the suite has a bill and a variance that ordinary tests do not.

OfficialApache-2.017.9k+174
RagasObservability

The failure most retrieval systems have is invisible from the outside — the answer reads well and is not supported by anything that was retrieved. Ragas measures that directly. It splits the pipeline in two and scores each half: whether the retrieved passages actually contained what was needed, and whether the generated answer stayed within them. It can also build a starting test set out of your own documents, so the first evaluation does not wait for someone to hand-write a hundred questions. Nearly every metric is itself a model call, which is the thing to plan around: a full run costs money, and two runs on the same data will not produce identical numbers.

Apache-2.015.5k
PhoenixObservability

Built on OpenTelemetry, so traces are portable rather than locked to one vendor, and it will run locally next to the code you are debugging. Licensed under Elastic 2.0 — fine for internal use, restrictive if you intend to offer it as a managed service.

OfficialSource-available11.2k+90
OpenLLMetryObservability

Every other option in this category is a product with its own backend, which means adopting one adds a system to run and a place where half your telemetry lives apart from the rest. This is the instrumentation layer instead: it emits standard OpenTelemetry spans for model and vector-database calls, and those go to Datadog, Grafana, New Relic, Splunk, Honeycomb or wherever your existing traces go. The trade is that a general-purpose backend shows spans and latencies, not the prompt-versioning and evaluation views a dedicated platform is built around.

OfficialApache-2.07.4k+18
NeMo GuardrailsGuardrails

Most guardrail tools check one thing in one place. This defines five interception points instead — before the model sees a request, over what retrieval pulled in, over the shape of the conversation, around tool calls, and before a reply reaches the user — and it sits between the application and the model, so it can be added to a system already in production without changing either side. The costs are latency and vocabulary: several of the stronger rails are model calls of their own, and dialogue rules are written in Colang, a purpose-built language.

OfficialApache-2.07k+36
HeliconeObservability

Most tracing tools ask you to wrap your calls, which is fine until the code you need to see is inside a library, a background job or a service somebody else owns. Helicone takes the other route: point the base URL at it and every request is logged on the way through, with no SDK in the path. Because it is already in that position it does gateway work too — one key for many providers, automatic fallback when one is down, caching for repeated prompts. The trade is the position itself: a proxy is now between your application and the model, which is a component to run and a hop to account for. Agent runs group into sessions so a multi-step trace reads as one thing rather than thirty unrelated calls.

Apache-2.06.1k
GiskardGuardrails

Evaluation tells you how well an agent does on the questions you thought of. That is the wrong shape for safety, where what matters is the question you did not think of. Giskard inverts it: the scan generates hostile inputs itself and reports the ones your agent answered when it should have refused, so the test set is produced by the tool rather than by your imagination. Alongside it sits a scan for the ways a retrieval system quietly goes wrong, and a way of writing behavioural expectations as ordinary tests that pass or fail under pytest, so a check that mattered once becomes a check that runs on every change. The maintainers are explicit that a clean scan is not a safety or compliance guarantee.

Apache-2.05.8k
agentgatewayGuardrails

A Rust data plane that sits between agents and everything they call — model providers, MCP servers, other agents — so authentication, RBAC, rate limits and OpenTelemetry export are configured once instead of reimplemented in every agent. Solo.io donated it to the Linux Foundation in August 2025 and it moved under the Agentic AI Foundation in June 2026, alongside MCP and goose. It is young for something that sits in the request path: per-user identity onto downstream MCP servers is still an open issue, so confirm the authentication model you need before putting it in front of production traffic.

OfficialApache-2.04.6k+162
HolmesGPTTools & Integrations

The slow part of an incident is not the fix, it is the twenty minutes of pulling up dashboards, matching a spike to a deploy and finding the pod that actually failed. HolmesGPT does that part: it connects to what you already run — Prometheus, Grafana, Datadog, Kubernetes, any REST API — and works through the question in a loop, fetching what it needs and following what it finds, then writes the conclusion back to the alert it came from. Notably it is built for the scale that breaks naive tools: results are filtered on the server, large outputs are streamed to disk and each tool has a memory limit, so querying a big observability dataset does not fill a context window or kill the process. It is a CNCF sandbox project.

Apache-2.03.2k
Inspect AIBenchmarks

Where the other entries in this category are each one measurement, this is the machinery for building measurements: a dataset supplies inputs and targets, a solver produces a response (a single model call or a full multi-turn agent), a scorer decides whether it was right, and a task binds the three together. More than two hundred pre-built evaluations come with it, model-written code runs inside a sandbox, and a log viewer shows what actually happened on each sample. What it cannot do is tell you what is worth measuring.

OfficialMIT2.7k+59
numbatGuardrails

numbat is a single binary that watches supported desktop, CLI, IDE and gateway agents through hooks, plugins and OTLP/HTTP log exporters, normalizes their activity into one event model, and evaluates it with a local CEL rule engine that writes versioned NDJSON events, findings and enforcement decisions. It also reconstructs past sessions from on-disk artifacts via `numbat scan`, so you can investigate an agent that was never instrumented. Blocking is deliberately conservative: enforce mode is off by default and every shipped rule is monitor-only, so denying an action means copying the rule YAML into your own `--rules-dir`, keeping its id, adding `enforce: true` and bumping the version. It is a poor fit if you expect coverage of arbitrary agents or a definitive account of what ran — the coverage matrix decides what is observable per host and surface, at-rest reconstruction cannot recover activity an agent never persisted, and findings are rule matches rather than proof that an action completed.

OfficialApache-2.0977+30