
haystack
文書に基づく回答もエージェントも、コンポーネントをパイプラインにつないで自分で組み立てるPythonのライブラリです
haystackとは
Haystackは、文書の取り込みから検索、モデルの呼び出しまでを1つずつコンポーネントとして置き、パイプラインの上で出力と入力を名前で指定してつないでいくフレームワークです。型の合わない組み合わせはつないだ瞬間にエラーになり、条件による分岐も、回数の上限つきで前のコンポーネントへ戻すループも、同じパイプラインの中に書けます。Agentもコンポーネントの一つなので、ツール呼び出しの繰り返しをパイプラインへ組み込んだり、別のAgentのツールとして渡したりできます。裏を返せば配線はすべて自分で書くことになり、PDFを置くだけで対話システムができあがる近道はありません。ライブラリなので自分のプロセスの中で動き、パイプラインをHTTPの窓口やMCPサーバーとして公開するには別プロジェクトのHayhooksを使うか、受け口を自分で用意することになります。
haystackで何ができますか?
- つなぎ間違いは実行前に分かる — パイプラインにコンポーネントを登録し、出力と入力を名前で指定してつなぎます。型の合わない組み合わせはつないだ瞬間にエラーになります。ファイルの種類で経路を分ける分岐や、条件を満たすまで前のコンポーネントへ戻すループも同じパイプラインに書け、1つのコンポーネントが回れる回数には既定で100回の上限が設けられています。
- Agentがツール呼び出しの繰り返しを受け持つ — Agentはチャットモデルの呼び出しとツールの実行を、終了条件を満たすか手数の上限(既定で100手)に達するまで繰り返します。ツール同士が共有する状態は型を宣言して持ち回り、応答は逐次流れ、複数のツール呼び出しは同時に走ります。Agent自体もコンポーネントなので、ComponentToolで包めば別のAgentのツールとして渡せます。
- フックで6か所に割り込む — 実行前、モデル呼び出しの前、ツール実行の前と後、終了時、実行後の6か所にフックを差し込めます。ツールを実行する前に人へ確認を取るConfirmationHookと、大きくなったツールの戻り値をファイルなどの保管先へ退避し、会話には冒頭の抜粋と参照先だけを残すToolResultOffloadHookが同梱されており、内部に手を入れずに済みます。
- 関数もコンポーネントもパイプラインもツールになる — Pythonの関数はデコレーターを付けるだけでツールになり、既存のコンポーネントやパイプライン全体もそれぞれ専用のクラスで包めばツールとして渡せます。数が増えて説明文がプロンプトを圧迫する場合は、SearchableToolsetがモデルにまず検索用のツールだけを見せ、必要になったものを読み込みます。手順書をSKILL.mdとして置いておくSkillToolsetも同じ考え方です。
- Document Storeを入れ替えても残りは動く — 取り込み側は、ファイルを読み取るConverter、長い文書を切り分けるSplitter、ベクトルに変換するEmbedderが順に並び、結果をDocument Storeへ書き込みます。問い合わせ側はRetrieverが候補を取り出し、キーワードの一致でもベクトルの近さでも、両方を混ぜても検索できます。手元のInMemoryDocumentStoreからQdrantやElasticsearch、pgvectorへ移すとき、差し替えるのはDocument Storeと対応するRetrieverだけです。
- 検索と回答の出来を数値で確かめる — 標準の評価コンポーネントで、必要な文書が返ってきたか、どれだけ上位に来たかを測れます。回答の側は、渡した文書から外れていないかを見るほか、判定基準を自分で書いてモデルに採点させることもできます。すでにRagasやDeepEvalを使っているなら、それらも同じくコンポーネントとして組み込めます。
haystackを選ぶ前に
- 3.0で旧来のGeneratorが削除され、30のコンポーネントが別途インストールする統合パッケージへ移り、2.31への修正は2026年10月末までの安全性対応と重大バグ修正に限るとdeepsetは述べています。
- 既製のパイプラインテンプレート、Helmチャートと導入手順、開発チームの直接サポート、プロンプトインジェクション対策などの先行提供は有償のHaystack Enterpriseの範囲で、公開されている本体には含まれません。
スター推移
8月17日〜8月28日 · +116
よくある質問
haystackは商用利用できますか?
haystackはApache-2.0ライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。
haystackはどの形で使えますか?
haystackはセルフホスト・ローカル実行の形で利用できます。
ドキュメント
deepset-ai/haystack のREADMEより転載(Apache-2.0)。 原文を読む ↗
| CI/CD | |
| Docs | |
| Package | |
| Meta |
🎉🎊✨ Haystack 3.0 is out! ✨🎊🎉
Read the announcement here!
🥳 🎈 🎆 🪅 🎇 🍾 🥂 🎁 🌈
Haystack is an open-source AI orchestration framework for building production-ready LLM applications in Python.
Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Build scalable RAG systems, multimodal applications, semantic search, question answering, and autonomous agents, all in a transparent architecture that lets you experiment, customize deeply, and deploy with confidence.
Table of Contents
- Installation
- Documentation
- Features
- Haystack Enterprise: Support & Platform
- Telemetry
- 🖖 Community
- Contributing to Haystack
- Organizations using Haystack
Installation
The simplest way to get Haystack is via pip:
pip install haystack-ai
Install nightly pre-releases to try the newest features:
pip install --pre haystack-ai
Haystack supports multiple installation methods, including Docker images. For a comprehensive guide, please refer to the documentation.
Documentation
If you’re new to the project, check out “What is Haystack?” then go through the “Get Started Guide” and build your first LLM application in a matter of minutes. Keep learning with the tutorials. For more advanced use cases, or just to get some inspiration, you can browse our Haystack recipes in the Cookbook.
At any given point, hit the documentation to learn more about Haystack, what it can do for you, and the technology behind.
Features
Agents built for production
Extend agent behavior with lifecycle hooks (before_llm, before_tool, on_exit, …) for guardrails and custom logic, and track step_count, token_usage, and tool calls out of the box for monitoring and cost control. Get started fast with ready-made agents from Agent Pack (e.g., a deep research agent, or an advanced RAG agent) or give your own agents progressive skill discovery via SkillToolset, so skill descriptions only enter context when needed.
Built for context engineering
Design flexible systems with explicit control over how information is retrieved, ranked, filtered, combined, structured, and routed before it reaches the model. Define pipelines and agent workflows where retrieval, memory, tools, and generation are transparent and traceable.
Native Async Support
One Pipeline runs synchronously or asynchronously and streams token by token. Agent can run concurrent tool calls.
Modular and customizable
Use built-in components for retrieval, indexing, tool calling, memory, and evaluation, or create your own. Add loops, branches, and conditional logic to precisely control how context moves through your pipelines and agent workflows.
Model- and vendor-agnostic
Integrate with OpenAI, Mistral, Anthropic, Cohere, Hugging Face, Google, Azure OpenAI, AWS Bedrock, local models, and many others. Swap models or infrastructure components without rewriting your system.
Extensible ecosystem
Build and share custom components through a consistent interface that makes it easy for the community and third parties to extend Haystack and contribute to an open ecosystem.
[!TIP]
Would you like to deploy and serve Haystack pipelines as REST APIs or MCP servers? Hayhooks provides a simple way for you to wrap pipelines and agents with custom logic and expose them through HTTP endpoints or MCP. It also supports OpenAI-compatible chat completion endpoints and works with chat UIs like open-webui.
Haystack Enterprise: Support & Platform
Get expert support from the Haystack team, build faster with enterprise-grade templates, and scale securely with deployment guides for cloud and on-prem environments with Haystack Enterprise Starter. Read more about it in the announcement post.
👉 Get Haystack Enterprise Starter
Need a managed production setup for Haystack? The Haystack Enterprise Platform helps you build, test, deploy and operate Haystack pipelines with built-in observability, collaboration, governance, and access controls. It’s available as a managed cloud service or as a self-hosted solution.
👉 Learn more about Haystack Enterprise Platform or try it free
Telemetry
Haystack collects anonymous usage statistics of pipeline components. We receive an event every time these components are initialized. This way, we know which components are most relevant to our community.
Read more about telemetry in Haystack or how you can opt out in Haystack docs.
🖖 Community
If you have a feature request or a bug report, feel free to open an issue in GitHub. We regularly check these, so you can expect a quick response. If you’d like to discuss a topic or get more general advice on how to make Haystack work for your project, you can start a thread in Github Discussions or our Discord channel. We also check 𝕏 (Twitter) and Stack Overflow.
Organizations using Haystack
Haystack is used by thousands of teams building production AI systems across industries, including:
- Technology & AI Infrastructure: Apple, Meta, Databricks, NVIDIA, Intel
- Public Sector AI Initiatives: European Commission, German Federal Ministry of Research, Technology, and Space (BMFTR), PD, Baden-Württemberg State
- Enterprise & Industrial AI Applications: Airbus, Lufthansa Industry Solutions, Infineon, LEGO, Comcast, Accenture, TELUS Agriculture & Consumer Goods
- Knowledge & Content Platforms: Netflix, ZEIT Online, Rakuten, Oxford University Press, Manz, YPulse
Are you also using Haystack? Open a PR or tell us your story