Guardrails & Security
Constraining what an agent may say and do: input and output filtering, prompt-injection defence, and policy enforcement around tool calls.
13 projects
The library half lets one function call reach any supported provider in the OpenAI request and response shape. The gateway half is the reason most teams arrive: a self-hosted server that issues virtual keys with their own budgets and rate limits, falls back when a provider fails, spreads load across deployments, and records cost per key. The price is an extra component in the request path that you now operate, and code under the enterprise directory is not covered by the MIT grant.
Test cases are YAML: a prompt, some inputs, and assertions about the output. That is enough to produce a side-by-side matrix comparing several prompts or models on the same inputs, and to fail a build when a change makes things worse. The same tool also attacks the application, generating adversarial inputs to surface vulnerabilities and compliance risks. It works from the outside — prompt in, answer out — so when the question is which step inside a running agent went wrong, this is not the instrument for it.
Installing a skill or an MCP server is closer to installing a browser extension than to adding a library: it arrives as instructions and scripts that an agent will follow with your credentials and your files in reach, and almost nobody reads them first. SkillSpector reads them. It scans a repository, archive or directory against a catalogue of patterns covering prompt injection, data being sent somewhere it should not go, privilege escalation, poisoned tool descriptions and supply-chain risk, then returns findings with a score and a recommendation. NVIDIA cites research finding vulnerabilities in 26.1% of skills examined and likely malicious intent in 5.2%. A pattern scanner reports suspicion rather than proof, so expect to review findings and record the ones you have accepted.
Agents call models far more often than a chat application does, which turns a provider's occasional rate limit into a daily outage. Portkey's gateway sits between your code and the providers and absorbs that: a failed call retries, a rate-limited key gives way to another, an unavailable model falls back to a second choice, and repeated prompts can be served from a cache instead of being paid for twice. Routing rules can send different requests to different models — cheap for classification, expensive for the hard step — without that decision spreading through your code. The project states it adds under a millisecond of latency in a footprint of about 122KB. One thing to check before planning around a feature: the documentation describes the open gateway and the company's hosted platform together.
`npx @openai/codex-security scan .` walks a checkout and leaves JSON results on stdout, while `--mode deep` spreads discovery across workers and subagents until it stops turning up anything new. Across runs, `scans compare BEFORE_SCAN_ID AFTER_SCAN_ID` matches findings by root cause and labels them new, persisting, reopened, resolved or unknown, so a second scan reports movement instead of restating the whole report. It is a poor fit as a pre-commit gate: deep discovery runs until `--max-time-hours`, which defaults to 96. Access is also gated — the CLI needs access to Codex Security, and some cybersecurity requests and protected findings require Trusted Access for Cyber approval, so cloning the repo on its own scans nothing.
Cloudflare OS pairs an agent chat with "Gadgets" — small apps the built-in coding agent writes for you, each running as a Dynamic Worker with its internet access disabled, its client confined to a sandboxed iframe that reaches the server only over Cap'n Web via `postMessage()`. External services are reached through Gatekeepers, per-service Workers that wrap an API in Cap'n Web, log every call, and simulate side-effecting actions so the agent keeps queueing work while approvals wait for you to review them in bulk. The commitment to Workers runs deep: every workspace is a Durable Object, every Gadget a Dynamic Worker Facet, and Facets and Dynamic Workers were added to the runtime specifically for this project — so hosting means a Cloudflare account, and the deploy-to-your-own-servers-on-`workerd` path is still marked COMING SOON in the README. The maintainers call the August 2026 v2 rewrite an early access release with many rough edges, and many Gatekeepers need configuration of their own — including OAuth client credentials per service — before GitHub or Google will connect.
Point it at a chatbot or a model and it works through families of attacks — prompt injection, jailbreaks, encoding-based bypasses, data leaks and replay, toxicity, false reasoning — and reports which ones got through. Keeping the attacks, the judgement and the connection to the target as separate components is what lets a new probe or a new backend be added without rewriting the rest. It is the finding half of security, not the fixing half: the output tells you which probes succeeded, not what policy to write in response.
Most guardrail tools check one thing in one place. This defines five interception points instead — before the model sees a request, over what retrieval pulled in, over the shape of the conversation, around tool calls, and before a reply reaches the user — and it sits between the application and the model, so it can be added to a system already in production without changing either side. The costs are latency and vocabulary: several of the stronger rails are model calls of their own, and dialogue rules are written in Colang, a purpose-built language.
The permissions an agent inherits are the ones its credentials carry, which is almost never the set of things it should be allowed to do: a key that can read a table can usually drop it. This toolkit closes that gap at the point of the call. You write the rules as a YAML file — deny anything destructive, require a named group's approval before an email goes out — wrap the tool function, and every call is evaluated, written to an audit trail, and refused with an explanation when a rule blocks it. Alongside that sit an identity scheme for attributing an action to a specific agent and a sandboxing layer for tools you would rather not run in your own process. It is a public preview, and the package layout has already moved once.
Evaluation tells you how well an agent does on the questions you thought of. That is the wrong shape for safety, where what matters is the question you did not think of. Giskard inverts it: the scan generates hostile inputs itself and reports the ones your agent answered when it should have refused, so the test set is produced by the tool rather than by your imagination. Alongside it sits a scan for the ways a retrieval system quietly goes wrong, and a way of writing behavioural expectations as ordinary tests that pass or fail under pytest, so a check that mattered once becomes a check that runs on every change. The maintainers are explicit that a clean scan is not a safety or compliance guarantee.
A Rust data plane that sits between agents and everything they call — model providers, MCP servers, other agents — so authentication, RBAC, rate limits and OpenTelemetry export are configured once instead of reimplemented in every agent. Solo.io donated it to the Linux Foundation in August 2025 and it moved under the Agentic AI Foundation in June 2026, alongside MCP and goose. It is young for something that sits in the request path: per-user identity onto downstream MCP servers is still an open issue, so confirm the authentication model you need before putting it in front of production traffic.
Where the other entries in this category are each one measurement, this is the machinery for building measurements: a dataset supplies inputs and targets, a solver produces a response (a single model call or a full multi-turn agent), a scorer decides whether it was right, and a task binds the three together. More than two hundred pre-built evaluations come with it, model-written code runs inside a sandbox, and a log viewer shows what actually happened on each sample. What it cannot do is tell you what is worth measuring.
numbat is a single binary that watches supported desktop, CLI, IDE and gateway agents through hooks, plugins and OTLP/HTTP log exporters, normalizes their activity into one event model, and evaluates it with a local CEL rule engine that writes versioned NDJSON events, findings and enforcement decisions. It also reconstructs past sessions from on-disk artifacts via `numbat scan`, so you can investigate an agent that was never instrumented. Blocking is deliberately conservative: enforce mode is off by default and every shipped rule is monitor-only, so denying an action means copying the rule YAML into your own `--rules-dir`, keeping its id, adding `enforce: true` and bumping the version. It is a poor fit if you expect coverage of arbitrary agents or a definitive account of what ran — the coverage matrix decides what is observable per host and surface, at-rest reconstruction cannot recover activity an agent never persisted, and findings are rule matches rather than proof that an action completed.