
Giskard
Attacks your own agent with hostile inputs and reports the ones it answered when it should have refused
What is Giskard?
Evaluation tells you how well an agent does on the questions you thought of. That is the wrong shape for safety, where what matters is the question you did not think of. Giskard inverts it: the scan generates hostile inputs itself and reports the ones your agent answered when it should have refused, so the test set is produced by the tool rather than by your imagination. Alongside it sits a scan for the ways a retrieval system quietly goes wrong, and a way of writing behavioural expectations as ordinary tests that pass or fail under pytest, so a check that mattered once becomes a check that runs on every change. The maintainers are explicit that a clean scan is not a safety or compliance guarantee.
What can you do with Giskard?
- Let the tool write the hostile inputs — The scan generates the attacks rather than asking you to think of them, which is the only way to cover the case you would not have imagined.
- Find where retrieval quietly fails — A separate scan looks for the ways a document-grounded system goes wrong without looking wrong, which ordinary accuracy numbers do not surface.
- Keep a finding as a test — Expectations are written as scenarios and checks that run under pytest, so something you fixed once is checked again on every change.
- Run it where your tests already run — Because it is a Python library rather than a platform, a scan fits into an existing pipeline without a service to stand up first.
Before you choose Giskard
- The maintainers state plainly that scan results are not a safety or compliance guarantee — it finds some of what is wrong, and a clean report is not evidence that nothing is.
- The open library covers the scans and the tests; continuous red teaming, dataset management and scheduled evaluation are part of the company's separate hosted product.
Frequently asked questions
Is Giskard free for commercial use?
Giskard is released under the Apache-2.0 licence — OSI-approved open source, which permits commercial use.
How can Giskard be deployed?
Giskard is available as Runs locally / Self-hosted.
Documentation
Reproduced from the Giskard-AI/giskard-oss README, published under Apache-2.0. Read the original ↗
[!IMPORTANT] Giskard v3 is a fresh rewrite designed for dynamic, multi-turn testing of AI agents. This release drops heavy dependencies for better efficiency while introducing a more powerful AI vulnerability scanner and enhanced RAG evaluation — both now shipping natively in
giskard-scan, with no dependency on v2. Only the legacy scan for tabular/ML models remains v2-only. Giskard v2 remains available but is no longer actively maintained. Follow progress → Read the v3 Announcement · Roadmap
Install
pip install giskard # checks (+ agents, llm, core)
pip install "giskard[scan]" # + vulnerability / quality scan
pip install "giskard[openai]" # provider SDK for LLM judges / generators
Requires Python 3.12+.
| Extra | Adds |
|---|---|
| (none) | giskard-checks and dependencies |
scan | giskard-scan |
openai / anthropic / … | provider SDKs (see pyproject.toml optional deps) |
Telemetry: optional aggregated analytics via giskard-core. No prompts or outputs are sent.
Opt out before importing Giskard: export DO_NOT_TRACK=1 or export GISKARD_TELEMETRY_DISABLED=1.
Details: giskard-core README.
Giskard is an open-source Python library for testing and evaluating agentic systems. The v3 architecture is a modular set of focused packages — each carrying only the dependencies it needs — built from scratch to wrap anything: an LLM, a black-box agent, or a multi-step pipeline.
| Status | Package | Description |
|---|---|---|
| ✅ Stable | giskard-checks | Testing & evaluation — scenario API, built-in checks, LLM-as-judge |
| ✅ Stable | giskard-scan | Agent vulnerability scanner + RAG/quality evaluation — red teaming, prompt injection, jailbreaks & harmful content (vulnerability_scan, successor of v2 Scan), plus knowledge-base quality eval (quality_scan, successor of v2 RAGET) |
These build on three foundational libraries — giskard-core (shared utilities & telemetry), giskard-llm (provider-agnostic LLM routing), and giskard-agents (agent & workflow orchestration) — which are pulled in automatically and rarely used directly.
Giskard Checks — create and apply evals for testing agents
pip install giskard-checks
Giskard Checks is a lightweight library for creating evaluations (evals) that test LLM-based systems — from simple assertions to LLM-as-judge assessments. Unlike traditional unit tests, evals are designed for non-deterministic outputs where the same input can produce different valid responses.
Use Giskard Checks to:
- Catch regressions — verify your system still behaves correctly after changes
- Validate RAG quality — check if answers are grounded in retrieved context
- Enforce safety rules — ensure outputs conform to your content policies
- Evaluate multi-turn agents — test full conversations, not just single exchanges
Built-in evals include string matching, comparisons, regex, semantic similarity, and LLM-as-judge checks (Groundedness, Conformity, LLMJudge).
Concepts
- Target — your system under test: any sync/async callable
(inputs) -> outputs(optionally withtrace) - Scenario — one eval: interactions + checks
- Check — assertion or LLM judge over the trace
- Suite — many scenarios run together
giskard.agents.Generator is an LLM client for workflows/judges — not the same as
giskard.checks input generators (LLMGenerator) that synthesize user messages.
Quickstart
import asyncio
from giskard.checks import Scenario, Groundedness
def get_answer(inputs: str) -> str:
return "Paris" # replace with your model / agent
async def main() -> None:
scenario = (
Scenario("test_france_capital")
.interact(inputs="What is the capital of France?", outputs=get_answer)
.check(
Groundedness(
name="answer is grounded",
context="France is in Western Europe. Its capital is Paris.",
)
)
)
result = await scenario.run()
result.print_report()
asyncio.run(main())
Groundedness is an LLM judge — install a provider extra (e.g. pip install "giskard[openai]") and set the matching API key. Default model: openai/gpt-4o-mini.
See the full docs for Suites, LLMJudge, multi-turn scenarios, and more.
Giskard Scan — vulnerability scanner for AI agents
pip install "giskard[scan]" # or: pip install giskard-scan
Giskard Scan is the red-teaming and vulnerability scanning layer for agentic systems. It generates adversarial test suites automatically from a plain-language description of your agent, covering prompt injection, harmful content, stereotypes, misinformation, and more.
Use Giskard Scan to:
- Red-team your agent — automatically generate adversarial inputs across OWASP LLM Top-10 threat categories
- Run prompt-injection probes — built-in dataset of injection payloads ready to use
- Extend with custom generators — pass your own
ScenarioGeneratorinstances togenerate_suite, or register them onvulnerability_suite_generator_registry
Quickstart
import asyncio
from giskard.scan import vulnerability_scan
async def my_agent(inputs: str) -> str:
# Replace with your agent / model call
return f"Echo: {inputs}"
async def main() -> None:
await vulnerability_scan(
target=my_agent,
description="A customer support chatbot for an e-commerce platform.",
languages=["en"],
)
asyncio.run(main())
Scan generators also need an LLM provider extra and API key (same as Checks judges above).
Looking for Giskard v2?
Giskard v2 included Scan (automatic vulnerability detection) and RAGET (RAG evaluation test set generation).
For LLM agents, both are superseded in v3 by giskard-scan: use vulnerability_scan in place of the v2 LLM scan, and quality_scan (with KnowledgeBase) in place of RAGET.
v3 works with ML models too — wrap one as a target and evaluate it with giskard-checks or giskard-scan. What the examples below cover is the v2-only automatic tabular scan — the detector suite that introspects a giskard.Model + giskard.Dataset to auto-detect performance, bias, and robustness issues — along with the giskard.testing ML test suite and the Giskard Hub. These are not planned for v3.
pip install "giskard[llm]>2,<3"
Scan — automatically detect performance, bias & security issues
Wrap your model and run the scan:
import giskard
import pandas as pd
# Replace my_llm_chain with your actual LLM chain or model inference logic
def model_predict(df: pd.DataFrame):
"""The function takes a DataFrame and must return a list of outputs (one per row)."""
return [my_llm_chain.run({"query": question}) for question in df["question"]]
giskard_model = giskard.Model(
model=model_predict,
model_type="text_generation",
name="My LLM Application",
description="A question answering assistant",
feature_names=["question"],
)
scan_results = giskard.scan(giskard_model)
display(scan_results)
RAGET — generate evaluation datasets for RAG applications
Automatically generate questions, reference answers, and context from your knowledge base:
import pandas as pd
from giskard.rag import generate_testset, KnowledgeBase
# Load your knowledge base documents
df = pd.read_csv("path/to/your/knowledge_base.csv")
knowledge_base = KnowledgeBase.from_pandas(df, columns=["column_1", "column_2"])
testset = generate_testset(
knowledge_base,
num_questions=60,
language="en",
agent_description="A customer support chatbot for company X",
)
We welcome contributions from the AI community! Read this guide to get started, and join our thriving community on Discord.
Follow the progress and share feedback: v3 Announcement · Roadmap
🌟 Leave us a star, it helps the project to get discovered by others and keeps us motivated to build awesome open-source tools! 🌟
❤️ If you find our work useful, please consider sponsoring us on GitHub. With a monthly sponsoring, you can get a sponsor badge, display your company in this readme, and get your bug reports prioritized. We also offer one-time sponsoring if you want us to get involved in a consulting project, run a workshop, or give a talk at your company.