← Back to all projects

HolmesGPT

An incident-investigation agent that queries your existing monitoring instead of asking you to paste logs into a chat

Apache-2.0
Stars
3.2k
Forks
456
Open issues
401
Last commit
26 Aug 2026

What is HolmesGPT?

The slow part of an incident is not the fix, it is the twenty minutes of pulling up dashboards, matching a spike to a deploy and finding the pod that actually failed. HolmesGPT does that part: it connects to what you already run — Prometheus, Grafana, Datadog, Kubernetes, any REST API — and works through the question in a loop, fetching what it needs and following what it finds, then writes the conclusion back to the alert it came from. Notably it is built for the scale that breaks naive tools: results are filtered on the server, large outputs are streamed to disk and each tool has a memory limit, so querying a big observability dataset does not fill a context window or kill the process. It is a CNCF sandbox project.

What can you do with HolmesGPT?

  • Investigate against live data, not a pasted excerpt — It queries the monitoring you already run, so the answer reflects the current state rather than whatever someone happened to copy into a message.
  • Write the finding back where the alert came from — Alerts are read from AlertManager, PagerDuty, OpsGenie or Jira and the conclusion is written back, so the investigation lands in the ticket rather than in a chat window.
  • Survive a large query — Server-side filtering, per-tool memory limits and streaming of big results to disk keep an investigation from being defeated by the size of the dataset.
  • Watch without being asked — An operator mode runs continuously, checking service health on a schedule or after a deploy and raising what it finds instead of waiting for a human to notice.
  • Cover more than Kubernetes — Virtual machines, cloud services, databases and SaaS platforms are all reachable, so it fits an estate that was never fully containerised.
  • Add a source it does not know — Any REST API can be registered as a toolset, so an in-house monitoring system does not put the whole approach out of reach.

Before you choose HolmesGPT

  • It is only as good as what it can reach — an estate whose useful signal lives in dashboards nobody exported, or in an undocumented internal system, gives it little to work with until those are wired up.
  • The always-on operator mode itself runs in Kubernetes, so teams on other infrastructure get the on-demand investigation but not the unattended monitoring half without more work.

Frequently asked questions

Is HolmesGPT free for commercial use?

HolmesGPT is released under the Apache-2.0 licence — OSI-approved open source, which permits commercial use.

How can HolmesGPT be deployed?

HolmesGPT is available as Self-hosted / Runs locally.

Documentation

Reproduced from the HolmesGPT/holmesgpt README, published under Apache-2.0. Read the original ↗

Open-source AI agent for investigating production incidents and finding root causes. Works with any stack — Kubernetes, VMs, cloud providers, databases, and SaaS platforms. We are a Cloud Native Computing Foundation sandbox project. Originally created by Robusta.Dev, with major contributions from Microsoft.

New: Operator Mode — Find Problems 24/7 in the Background

Most AI agents are great at troubleshooting problems, but still need a human to notice something is wrong and trigger an investigation. Operator mode fixes that — HolmesGPT runs in the background 24/7, spots problems before your customers notice, and messages you in Slack with the fix. Connect the GitHub integration and it can even open PRs to fix what it finds.

While the operator itself runs in Kubernetes, health checks can query any data source Holmes is connected to — VMs, cloud services, databases, SaaS platforms, and more.

Features

  • Petabyte-scale data: Server-side filtering, JSON tree traversal, and tool output transformers keep large payloads out of context windows
  • Memory-safe execution: Per-tool memory limits, streaming large results to disk, and automatic output budgeting prevent OOM kills when querying large observability datasets
  • Deep integrations: Prometheus, Grafana, Datadog, Kubernetes, and many more—plus any REST API
  • Bidirectional alert integrations: Fetch alerts from AlertManager, PagerDuty, OpsGenie, or Jira—and write findings back
  • Any LLM provider: OpenAI, Anthropic, Azure, Bedrock, Gemini, and more
  • No Kubernetes required: Works with any infrastructure — VMs, bare metal, cloud services, or containers

How it Works

HolmesGPT uses an agentic loop to query live observability data from multiple sources and identify root causes.

HolmesGPT Investigation Demo

🔗 Data Sources

HolmesGPT integrates with popular observability and cloud platforms. The following data sources (“toolsets”) are built-in. Add your own.

Data SourceNotes
AKSAzure Kubernetes Service cluster and node health diagnostics
Atlassian RovoJira issues and Confluence pages via Atlassian’s hosted server (MCP)
ArgoCDGet status, history and manifests and more of apps, projects and clusters
AWSRDS events, instances, slow query logs, and more (MCP)
AzureAzure resources and diagnostics (MCP)
ConfluencePrivate runbooks and documentation
Confluence (MCP)Private runbooks and documentation (MCP)
CoralogixRetrieve logs for any resource
CrossplaneTroubleshoot Crossplane providers, compositions, claims, and managed resources
DatadogQuery logs, metrics, and traces
DockerGet images, logs, events, history and more
Elasticsearch / OpenSearchQuery logs, cluster health, shard and index diagnostics
GCPGoogle Cloud Platform resources (MCP)
GitHubRepositories, issues, and pull requests (MCP)
GitLabProjects, merge requests, issues, and CI/CD pipelines (MCP)
Jenkins (MCP)Build status, pipeline logs, and job history (MCP)
GrafanaQuery and analyze dashboard configurations and panels
HelmRelease status, chart metadata, and values
InternetPublic runbooks, community docs, etc.
KafkaFetch metadata, list consumers and topics or find lagging consumer groups
KubernetesPod logs, K8s events, and resource status (kubectl describe)
Kubernetes Remediation (MCP)Apply fixes like scaling, rollbacks, and resource edits (MCP)
LokiQuery logs for Kubernetes resources or any query
MariaDBMariaDB database queries and diagnostics
MongoDBQuery data, diagnose performance, inspect schemas, find slow operations
MongoDB AtlasCluster health, slow queries, and performance diagnostics
NewRelicInvestigate alerts, query tracing data
OpenShiftProjects, routes, builds, security context constraints, and deployment configs
Prefect (MCP)Workflow orchestration monitoring, flow runs, and worker health (MCP)
PrometheusInvestigate alerts, query metrics and generate PromQL queries
RabbitMQPartitions, memory/disk alerts, troubleshoot split-brain scenarios and more
RobustaMulti-cluster monitoring, historical change data, runbooks, PromQL graphs and more
ServiceNowQuery tables and incident records
SentryError tracking, issues, and performance monitoring (MCP)
SlabTeam knowledge base and runbooks on demand
SplunkLog search and analysis (MCP)
SQL DatabasesPostgreSQL, MySQL, ClickHouse, MariaDB, SQL Server, Azure SQL, SQLite
TempoFetch trace info, debug issues like high latency in application
VictoriaLogsQuery logs from VictoriaLogs using LogsQL
VictoriaMetricsQuery metrics from a Prometheus-compatible TSDB (vmsingle / vmcluster)
ZabbixMonitor hosts, problems, events, triggers, and historical metrics

See the full list of built-in toolsets for additional integrations including Cilium, KubeVela, Notion, and more.

🚀 End-to-End Automation

HolmesGPT can fetch alerts/tickets to investigate from external systems, then write the analysis back to the source or Slack.

IntegrationStatusNotes
Slack✅Demo. Available via Robusta
Microsoft Teams✅Available via Robusta
Prometheus/AlertManager✅Robusta or HolmesGPT CLI
PagerDuty✅HolmesGPT CLI only
OpsGenie✅HolmesGPT CLI only
Jira✅HolmesGPT CLI only
GitHub✅HolmesGPT CLI only

Installation

Read the installation documentation to learn how to install HolmesGPT.

Supported LLM Providers

Read the LLM Providers documentation to learn how to set up your LLM API key.

Using HolmesGPT

See the walkthrough documentation for usage guides, including:

🔐 Data Privacy

By design, HolmesGPT has read-only access and respects RBAC permissions. It is safe to run in production environments.

Community

Join our community to discuss the HolmesGPT roadmap and share feedback:

Support

If you have any questions, feel free to message us on HolmesGPT Slack Channel

How to Contribute

Please read our CONTRIBUTING.md for guidelines and instructions.

For help, contact us on Slack or ask DeepWiki AI your questions.

Please make sure to follow the CNCF code of conduct - details here.

OpenSSF Best Practices OpenSSF Scorecard