Tools & Integrations
Ready-made capabilities an agent can call: search, code execution, file handling and service integrations.
42 projects
OpenClaw installs a Gateway on your own machine and makes it the local control plane for sessions, tools, events and channel connections, so the Control UI, CLI, TUI and every connected messaging channel drive the same assistant. Models can be hosted or local, and the project expects new capabilities to arrive as plugins built on the plugin SDK rather than as changes to the core. The constraints are stated plainly in the README: it is designed for a single operator, tools run on the host for the main session unless you configure sandboxing, and DM-capable channels leave the assistant reachable by unknown senders until you approve a pairing with `openclaw pairing approve <channel> <code>`. The README tells you to read the security, exposure and sandboxing guides before connecting other users or exposing the Gateway remotely, which is a fair signal that shared or multi-user deployment is not the case it was built for.
Hermes Agent is Nous Research's own agent: a terminal UI plus a single gateway process that carries the same conversation into Telegram, Discord, Slack, WhatsApp, Signal and Email, with a cron scheduler, isolated subagents, and seven terminal backends from local and Docker to Modal and Vercel Sandbox. Its distinguishing piece is a closed learning loop — the agent curates its own memory on periodic nudges, creates skills after complex tasks, and searches past sessions through FTS5 with LLM summarization — which means the state that makes it useful accumulates under `~/.hermes`, outside your version control, and needs occasional pruning. It is an assistant you run and configure rather than a library you build on: the README documents `hermes` subcommands and slash commands, not an embedding API, so it is a poor fit if you wanted an agent loop to call from your own code. Putting a shell-capable agent behind chat platforms also makes the command-approval and DM-pairing settings load-bearing rather than optional.
Connects hundreds of services as workflows you draw on screen — each step in the flow is a node — and increasingly puts an AI agent node inside those workflows. The licence is the thing to check first: it is source-available under the Sustainable Use License, not OSI-approved open source. Internal use and self-hosting are fine; reselling it as a hosted service is not.
Its real value is coverage: whatever model, vector store or API you need, an adapter probably already exists, which shortens the distance from idea to prototype. The abstractions have a reputation for indirection, so many teams use it for the integrations and reach for something more explicit once the control flow gets complicated.
Extracts the interactive elements of a page and hands the model a structured list, instead of asking it to guess click coordinates from a screenshot. That makes it markedly more reliable for form flows and scraping. It degrades on canvas-heavy applications and on sites with aggressive bot protection, where there is no clean DOM to read.
It hooks into the agent's lifecycle, captures every tool call, and has a small model — Claude Haiku by default, on your existing subscription or key — compress the stream into titled observations that are injected back at the next session's start; search goes through MCP tools staged from cheap index lines to full detail. Storage is local SQLite with an optional vector index, installs are documented beyond Claude Code for OpenCode, Cursor and Codex CLI, and a private tag keeps marked text out of storage. Adopt it knowing what it is: a resident daemon plus Bun and Python tooling rather than a passive plugin, compression that spends real model quota every session, and a project that went through thirteen major versions in its first year.
Left alone, a coding agent starts writing code the moment you finish the sentence. This is a set of skills that puts the usual engineering order back: nine commands cover the stages from working out what to build to shipping it, and the skills behind them attach on their own when the work calls for one — designing an interface pulls in the interface skill, touching the front end pulls in that one. Notably the automatic mode does not remove the checks, only the person stepping between tasks: each task is still test-driven, committed on its own, and paused on failure. It is Markdown, so it installs into around seventy agents through the open skills command.
The clearest way to understand the Model Context Protocol is to read servers that implement it correctly, and this is that collection. Useful both as working connectors and as the template to copy when writing your own.
Ask for a library by name and Context7 resolves it to an identifier like /vercel/next.js, then answers queries with version-specific snippets from a continuously re-indexed documentation corpus — so the model writes against the API that exists today rather than the one in its training data. Setup is one command or a URL pasted into any MCP client, and anyone can submit a public repository to the index. Know what you are adopting: the open-source part is a thin client, while the crawler, parser and index are Upstash's private hosted service — the free tier stops at 1,000 calls a month, and when the hosted endpoint went down for a stretch in August 2026, the tool had nothing to answer from.
It hands an agent a live Chrome plus the instruments of DevTools: it records a performance trace and reads out Core Web Vitals with named insights, lists network requests and source-mapped console errors, runs Lighthouse audits, and still clicks and fills forms from a text snapshot of the page the way browser MCPs do. It can attach to the Chrome you already have open, profile and all — which is also where care is needed. The trade against Playwright MCP is scope: this is Chrome-only, and what that buys is analysis no cross-browser server offers.
The problem with asking a model to make its own writing sound less like a model is that it has no list of what gives it away. Humanizer supplies one: thirty-five patterns from the checklist Wikipedia's editors maintain for spotting AI-written articles. It drafts a rewrite without treating the original structure as fixed, then checks that draft against the same patterns and against the original claims before going again, and it shows the intermediate pass and its own critique rather than only the result. It does not invent: names, numbers, dates and citations must come from the source. Pointed at a file it changes prose only, leaving code, data and link targets alone.
An agent that only produces text wastes the application it lives in. CopilotKit connects the two: a hook publishes part of your application state so the agent can read it, another registers a frontend function the agent may call, and a third lets the agent answer with a real React component rather than a paragraph — a form to confirm, a chart, a diff to approve. It talks to backends through AG-UI, the protocol the same team publishes, so the agent can be LangGraph, CrewAI, the Claude Agent SDK or several others without changing the frontend. What it does not do is be the agent: you still bring one, and the smooth path is React.
An agent can already call any Python library, but it will use a specialist bioinformatics package the way someone reads the docs for the first time — plausibly, and wrong in the ways that matter. This supplies the missing half: over 160 skills that carry the curated usage and worked examples for specific scientific tools, plus access to a hundred or so public databases, across genomics, chemistry, drug discovery, proteomics, clinical research, imaging, physics and geospatial work. It installs as one plugin or skill by skill. Read the scope notes before relying on it, though — the clinical and healthcare skills are written for research and retrospective validation, and say plainly that they are not for patient-specific decisions.
Playwright MCP puts a real Chrome, Firefox, WebKit or Edge behind an MCP server and returns each page as an accessibility snapshot: a text outline of roles, names and reference ids that the agent clicks and types into, with no vision model and no coordinate guessing. Past plain navigation it can mock network requests, read and write cookies and web storage, record a trace or a video, and emit Playwright locators and assertions, which makes it a practical way to turn a session you drove by hand into a test you can check in. The cost is context: Microsoft's own README points coding agents at the Playwright CLI with skills instead, because the tool definitions and the snapshots compete for the same window as your codebase. It earns its keep in long-running loops, where holding one browser open across many turns is worth paying that.
It keeps no index of its own; a query is fanned out to as many as 269 search services and the results are merged, which is how an agent gets web search without a commercial search API behind it. A JSON endpoint makes it usable from code rather than only from the browser. Two consequences follow from having no index: result quality and availability are inherited from the upstream engines, and when one of them starts blocking your instance that is your problem to solve.
Point Claude, Copilot, Cursor or any other MCP client at it and the model can triage an issue, open and review a pull request, search code it has never downloaded, and read the logs of a failed Actions run — all through the GitHub API, so nothing is cloned. The remote server GitHub hosts needs no local setup or runtime at all; the same local server runs as a container or a single binary, and is the only route for GitHub Enterprise Server. The trade-off is breadth: more than eighty tools sit across twenty-odd toolsets, and only five toolsets are on by default for good reason — turning everything on gives most models more choices than they handle well.
Formerly Danswer, this is the whole surface a company usually assembles by hand — a chat interface, indexing from more than fifty sources, custom agents with their own instructions and actions, a multi-step research mode that returns a report, web search, code execution and artifacts — deployable by one command and pointed at any model provider, self-hosted or proprietary. Two things to weigh: the full deployment is a stack of index, workers, inference servers, cache and blob store, and directories named ee carry a separate enterprise licence rather than MIT.
Composio hands an agent a catalogue of ready-made actions spread across more than a thousand apps — create a GitHub issue, fetch Gmail threads — and keeps each of your users' account connections stored and refreshed, so you never write a sign-in flow yourself. A session is scoped to one user and by default gives the model only a handful of meta tools for finding, authenticating and running things at runtime, so hundreds of tool definitions never fill the context window; app events arrive as signed webhooks, and the agent gets a persistent Python sandbox to work in. The trade-off is that none of that machinery is yours: you gain a thousand integrations you did not have to build, and give up the ability to change any of them.
Asking an agent for a diagram usually returns rounded boxes and arrows that look nothing like the document they are going into, so the diagram gets dropped. This skill gives the agent a vocabulary instead: thirty-nine named layouts — architecture, sequence, swimlane, quadrant, Sankey, fishbone, Wardley map, story map and more — each shipped in light, dark and full-editorial variants as one HTML file with inline SVG. No build step, no JavaScript, no external images: open the file in a browser or drop it into a page. It can also read your site to pick up the brand, and redraw an existing draw.io or Mermaid source at a chosen size and level of detail.
A decorator over a typed Python function is a complete tool: the schema is derived from the type hints and docstring, and both the arguments coming in and the value going out are validated for you. Beyond that it covers the parts a second week of work runs into — a typed client, OAuth and remote identity providers, mounting several servers as one, proxying an existing server, and middleware. The confusion to be aware of is the name: an earlier version of this API was absorbed into the official MCP Python SDK, so FastMCP refers to two different things depending on which import you read.
Four functions carry most of the work — generateText and streamText for prose, generateObject and streamObject for data shaped by a schema — and they keep the same signature whether the model behind them is OpenAI, Anthropic, Google, Bedrock or a local one. A second half supplies framework hooks so a streaming chat UI is a few lines rather than a hand-rolled parser. It is a model-access and UI layer, not an agent runtime: persistent memory, retrieval and durable orchestration are somebody else's job.
Blender is powerful and its learning curve is the reason most people never get past a cube. This bridges it to any MCP-capable assistant: an add-on inside Blender opens a connection, and from the other side the model can create and move objects, apply materials, inspect the scene and run Python against Blender's own API. That last part is the important detail and the main risk — the assistant is not limited to a fixed set of operations, so it can do things you did not anticipate in a file you care about. Installation is genuinely three steps, but they span two programs: a package runner, an MCP client entry, and an add-on enabled inside Blender itself.
When the person who knew why the system is like that leaves, the documentation they wrote turns out to be the small half of what they knew. Distilly is an attempt at the other half: you give it their messages, documents, interviews or public writing, and it produces a profile of the observable pattern — how they decided things, what they insisted on, how they said it — packaged as a skill an agent can load. The maintainers are careful about what it claims, and so should you be: the output is grounded in the material you supplied and does not claim to be the person. The obvious caution is consent, since the material it works from is usually someone else's.
A decorator on a function or a wrapper around a client is enough to start recording traces, and from there the same data feeds experiments that score responses automatically instead of by hand — hallucination, answer relevance, context recall and the rest. More than thirty integrations cover the usual frameworks and providers, and the whole platform can run locally or on Kubernetes rather than only as a hosted service. Its difficulty is not capability but overlap: it competes directly with two platforms already in this catalogue, so the choice turns on integrations and hosting rather than on what is missing.
You write a character file — a name, a system prompt, a few bio lines and the list of plugins to load — and the runtime turns it into an agent that stores each message, composes context from the providers the turn actually routes to, lets the model pick an action, and files what it learned through evaluators that run after the reply. First-party plugins ship in the same repository for the chat platforms people already use, MCP servers, browser and desktop control, calendar and inbox assistants, non-custodial EVM and Solana wallets, and an on-device model path that answers with the network off. Two things to weigh before adopting it: the repository is a whole product stack — desktop and mobile apps, a hosted cloud, native device bridges — so you take on far more than an agent library, and the stable release on npm is still the older 1.x line while the stack described here ships only under beta and alpha tags.
Over a hundred skill directories of Google's own written guidance, installed selectively with `npx skills add google/skills`. The bulk sits under `skills/cloud` — GKE, BigQuery, Agent Platform, the six Well-Architected pillars — while `skills/ads` and `skills/analytics` cover the Google Ads API, the Mobile Ads and IMA SDKs and the two Google Analytics APIs. Depth is uneven: `skills/cloud/gke-*` accounts for twenty-nine directories on its own, Cloud Run and Firebase get one each, and the README labels the repository as under active development. A stack on neither Google Cloud nor Google's ads and analytics products gets nothing here, and app-side coverage stops at the Mobile Ads SDK — Android, Flutter, Dart, Genkit and ADK skills are separate repositories the README only links to.
The shape is deliberately pytest's: assertions inside test functions, a runner that discovers them, parameterised cases — except the assertion is about whether an answer was faithful to its context or whether an agent finished the task. That framing is the point, because it puts quality regressions in front of the same gate that catches syntax errors. The thing to plan for is that most of these metrics are scored by a model, so the suite has a bill and a variance that ordinary tests do not.
An agent working against an unfamiliar library reads the same documentation site over and over, a page at a time, paying for it on every run. Skill Seekers converts that source once: point it at a documentation site, a repository, a PDF, a wiki or a video and it produces a structured knowledge asset, which it can then export as an agent skill or into a retrieval pipeline. Eighteen kinds of source go in and it can package for a couple of dozen destinations, so the same extraction serves a coding assistant and a search index. It also checks for conflicts, which is the failure mode when two versions of the same documentation end up loaded at once.
The headline feature is AI Services: you declare a Java interface, annotate it, and the library supplies the implementation that builds the prompt, calls the model and maps the reply back onto your return type. Underneath, a plain ChatModel API is there whenever the declarative layer gets in the way, and RAG, tool calling and integrations for Spring Boot, Quarkus, Micronaut and Helidon are first-party. Despite the name it is not a port of the Python library, so material from that ecosystem does not transfer.
Agents call models far more often than a chat application does, which turns a provider's occasional rate limit into a daily outage. Portkey's gateway sits between your code and the providers and absorbs that: a failed call retries, a rate-limited key gives way to another, an unavailable model falls back to a second choice, and repeated prompts can be served from a cache instead of being paid for twice. Routing rules can send different requests to different models — cheap for classification, expensive for the hard step — without that decision spreading through your code. The project states it adds under a millisecond of latency in a footprint of about 122KB. One thing to check before planning around a feature: the documentation describes the open gateway and the company's hosted platform together.
Connects to an MCP server you are developing and lets you call its tools, read its resources and render its prompts one at a time, checking each response as it comes back. It starts from npx with nothing installed, as a web UI, a scriptable CLI for CI, or a terminal TUI — all from the same command. It is a development-time check, not production monitoring or evals.
Most MCP servers return text and leave the client to render it. mcp-use is built around the opposite idea: a tool can carry a React view, so what the user sees inside ChatGPT or Claude is an interface rather than a paragraph. Tool inputs and outputs are declared with Zod schemas and those types flow through to the view's props, so a change in one place shows up as a type error in the other. A scaffolded project comes with an inspector at a local URL for calling tools and looking at views while you work, and a tunnel for testing against a real client. Version 2 is a rewrite around this server-and-view model; the earlier Python client library is no longer where the project's attention is.
`npx @openai/codex-security scan .` walks a checkout and leaves JSON results on stdout, while `--mode deep` spreads discovery across workers and subagents until it stops turning up anything new. Across runs, `scans compare BEFORE_SCAN_ID AFTER_SCAN_ID` matches findings by root cause and labels them new, persisting, reopened, resolved or unknown, so a second scan reports movement instead of restating the whole report. It is a poor fit as a pre-commit gate: deep discovery runs until `--max-time-hours`, which defaults to 96. Access is also gated — the CLI needs access to Codex Security, and some cybersecurity requests and protected findings require Trusted Access for Cyber approval, so cloning the repo on its own scans nothing.
Point Claude, Cursor, Kiro or any other MCP client at one of these servers and the model can search current AWS documentation, write CloudFormation and CDK against AWS's own guidance, query DynamoDB or PostgreSQL, read CloudWatch logs and estimate costs — each capability is its own small server, started with one uvx command and your AWS credentials. Two need no setup at all: the hosted Knowledge server answers documentation questions without an AWS account, and the preview AWS MCP Server reaches the AWS APIs under IAM permissions with CloudTrail logging. The open question is longevity: AWS announced the Agent Toolkit for AWS in May 2026 as this suite's successor, recommends it for production agent work, and says the most useful servers here will move into it over time.
Point it at a chatbot or a model and it works through families of attacks — prompt injection, jailbreaks, encoding-based bypasses, data leaks and replay, toxicity, false reasoning — and reports which ones got through. Keeping the attacks, the judgement and the connection to the target as separate components is what lets a new probe or a new backend be added without rewriting the rest. It is the finding half of security, not the fixing half: the output tells you which probes succeeded, not what policy to write in response.
Every other option in this category is a product with its own backend, which means adopting one adds a system to run and a place where half your telemetry lives apart from the rest. This is the instrumentation layer instead: it emits standard OpenTelemetry spans for model and vector-database calls, and those go to Datadog, Grafana, New Relic, Splunk, Honeycomb or wherever your existing traces go. The trade is that a general-purpose backend shows spans and latencies, not the prompt-versioning and evaluation views a dedicated platform is built around.
A coding agent working on an iOS project can write Swift all day and never find out whether it compiles, because the commands that would tell it are long, particular and easy to get subtly wrong. This exposes them properly: build a scheme, boot a simulator, install and launch the app, read the logs, run the tests. Because the agent gets a result it can act on, the loop closes — it fixes the error it just caused instead of handing you a diff to try. It ships as both an MCP server and a plain command-line tool, with optional skills that tell the agent how to use whichever one you installed. It is macOS-only and tracks a current Xcode.
Most tracing tools ask you to wrap your calls, which is fine until the code you need to see is inside a library, a background job or a service somebody else owns. Helicone takes the other route: point the base URL at it and every request is logged on the way through, with no SDK in the path. Because it is already in that position it does gateway work too — one key for many providers, automatic fallback when one is down, caching for repeated prompts. The trade is the position itself: a proxy is now between your application and the model, which is a component to run and a hop to account for. Agent runs group into sessions so a multi-step trace reads as one thing rather than thirty unrelated calls.
The expensive habit of a coding agent is grep, then read the eight files it matched, then read three more because the answer was not there. Semble replaces that with one question in plain language and a handful of relevant snippets back — the maintainers put the saving at roughly 99% of the tokens against grep plus reading. It runs entirely on the machine, on CPU, with no API key and no service to call, and indexing a full repository takes under a second, which is what makes it practical to point at whatever the agent happens to be working on. Installation detects the agents you have and offers three ways in: an MCP tool, a line in your agent instructions, or a dedicated search sub-agent.
The MCP problem an organisation hits is not writing a server, it is that after a year there are forty of them, nobody has a list, and each client is configured by hand. ContextForge is the registry and proxy for that situation: servers are registered once and discovered from one place, existing REST and gRPC services can be exposed as MCP without being rewritten, A2A agents are routed the same way, and authentication, rate limits and tracing are applied centrally instead of per server. It is itself a compliant MCP server, so clients see one endpoint. The weight is the obvious cost — this is infrastructure with a database, a cache and a Kubernetes story, which is a lot to run if you have three servers rather than forty.
The slow part of an incident is not the fix, it is the twenty minutes of pulling up dashboards, matching a spike to a deploy and finding the pod that actually failed. HolmesGPT does that part: it connects to what you already run — Prometheus, Grafana, Datadog, Kubernetes, any REST API — and works through the question in a loop, fetching what it needs and following what it finds, then writes the conclusion back to the alert it came from. Notably it is built for the scale that breaks naive tools: results are filtered on the server, large outputs are streamed to disk and each tool has a memory limit, so querying a big observability dataset does not fill a context window or kill the process. It is a CNCF sandbox project.
numbat is a single binary that watches supported desktop, CLI, IDE and gateway agents through hooks, plugins and OTLP/HTTP log exporters, normalizes their activity into one event model, and evaluates it with a local CEL rule engine that writes versioned NDJSON events, findings and enforcement decisions. It also reconstructs past sessions from on-disk artifacts via `numbat scan`, so you can investigate an agent that was never instrumented. Blocking is deliberately conservative: enforce mode is off by default and every shipped rule is monitor-only, so denying an action means copying the rule YAML into your own `--rules-dir`, keeping its id, adding `enforce: true` and bumping the version. It is a poor fit if you expect coverage of arbitrary agents or a definitive account of what ran — the coverage matrix decides what is observable per host and surface, at-rest reconstruction cannot recover activity an agent never persisted, and findings are rule matches rather than proof that an action completed.