Agent Frameworks
Libraries that give an agent its control loop: how it plans, calls tools, keeps state and recovers from failure. The choice here shapes everything downstream, so weigh debuggability and lock-in alongside how fast the first demo comes together.
39 projects
`dsh` is the agent harness DeepSeek AI develops itself. `npx @deepseek-ai/dsh web` brings up a Web UI on `http://127.0.0.1:3080`, and the architecture behind it is one where everything is a plugin, powered by the Cordis runtime, so it suits people who intend to extend a harness rather than only operate one. The caveat comes from the project itself: it is labelled a developer preview and warns in capitals that there will be compatibility-breaking changes, so plugins written today should budget for rework. It is also a poor fit if you want a documented scripted or embedded entry point right now, since the README covers only the `web` command and a source checkout, and links out to a development guide and architecture documentation without saying what either contains.
An agent here is a graph of blocks — one block calls a service, transforms data, prompts a model or branches — assembled on a canvas rather than written in code, then run on demand, on a schedule or from a trigger. One repository holds both the free self-hosted platform and the code behind the paid hosted one, and that is the catch: everything under the platform directory is PolyForm Shield, which permits use but forbids offering it as a competing product.
A complete platform — prompt orchestration, retrieval, agent tools and an admin UI — that non-engineers can operate once a team has set it up. The licence is a modified Apache 2.0 with additional conditions on multi-tenant hosting and branding, so it is source-available rather than open source. Read it before building a product on top.
Its real value is coverage: whatever model, vector store or API you need, an adapter probably already exists, which shortens the distance from idea to prototype. The abstractions have a reputation for indirection, so many teams use it for the integrations and reach for something more explicit once the control flow gets complicated.
Hand MetaGPT a one-line requirement and a team of LLM roles works through it in turn, leaving a requirements document, a system design and source files in a folder under `workspace/`. Which roles show up depends on where you installed from, and the gap is not cosmetic: the published 0.8.2 package hires a product manager, an architect, a project manager and an engineer, while current main — the README's clone-and-install route — hires a team leader, a product manager, an architect, an engineer and a data analyst, with the project manager commented out. Under the demo sits a general framework: you subclass Role and Action, declare which upstream actions each role watches, and a shared Environment routes the messages that put the team in order. A separate agent, the Data Interpreter, does data work alone, planning the steps, writing Python and retrying what fails. The catch is scope — MetaGPT's own FAQ says functions start being left unimplemented past roughly 500 lines of generated code.
Microsoft Research's take on multi-agent systems, where progress emerges from a structured conversation between specialised agents. Strong for exploration and research-flavoured problems. Two things to check before adopting: the API was substantially reworked between major versions, and GitHub reports the repository licence as CC-BY-4.0 — a content licence, not an OSI-approved software one. Confirm the terms with the maintainers before commercial use.
Models a task as a crew of agents with named roles and goals that hand work to each other. The metaphor makes a multi-agent design readable at a glance, which is why it demos so well. Whether several role-playing agents beat one well-prompted agent on your task is worth measuring before committing — the answer is often no.
Where most frameworks start from the agent loop, this one starts from the data: ingestion, indexing strategies, and the retrieval patterns that decide whether answers are actually grounded. Reach for it when the hard part of your problem is the corpus rather than the reasoning.
Agno is three layers you can adopt one at a time: a Python SDK for agents, teams and workflows; AgentOS, a FastAPI service that turns them into a REST API with sessions, JWT-scoped access, scheduling and OpenTelemetry traces written to a database you own; and a web control plane for watching runs. The first two are Apache-2.0 and run wherever containers run; apart from a usage ping you can switch off, nothing has to reach Agno. The control plane is the part that does not follow that rule — free against a runtime on your own machine, and a paid plan to point it at a deployed one. The project also moves quickly, with breaking changes in most minor releases, so pin a version before you build on it.
LangGraph has you define the agent's flow as a graph. Nodes are the units of work — an LLM call, a tool run — edges decide which node runs next, and every node reads and writes one shared state. Edges can branch on that state to form loops, and the state is checkpointed at each node, so a run can stop for a person's approval and resume later, or restart from the last checkpoint after a crash. The cost is code: a simple tool-calling agent takes noticeably more setup here than elsewhere.
You declare what each step should do and supply a metric; the optimiser searches for the prompts and few-shot examples that maximise it. The payoff is that improving a pipeline becomes measurable rather than superstitious. It requires an evaluation set — without one there is nothing to optimise against.
An agent that only produces text wastes the application it lives in. CopilotKit connects the two: a hook publishes part of your application state so the agent can read it, another registers a frontend function the agent may call, and a third lets the agent answer with a real React component rather than a paragraph — a form to confirm, a chart, a diff to approve. It talks to backends through AG-UI, the protocol the same team publishes, so the agent can be LangGraph, CrewAI, the Claude Agent SDK or several others without changing the frontend. What it does not do is be the agent: you still bring one, and the smooth path is React.
AgentScope starts where most frameworks start — an agent that reasons, calls a tool, looks at the result and goes round again — and then supplies the parts that usually have to be built afterwards. Tools run inside an isolated environment rather than in your process; context is compressed and offloaded as it grows instead of simply overflowing; a person can be asked to approve a step; and the finished agent can be served as a multi-tenant backend with its own interface. The price of that reach is that version 2.0 was a deliberate break from 1.0, so material written for the older releases does not carry over.
Few primitives — agents, handoffs, guardrails, tracing — and little else. The small surface area is the appeal: there is not much to learn and not much to fight. It is built around OpenAI's own models first, so weigh that if multi-provider portability matters to you.
The central idea is the CodeAgent: rather than returning a structured tool call, the model writes a short Python snippet that calls your tools directly, so nesting a call inside a loop or a conditional comes from the language itself instead of from framework machinery. The whole agent loop is about a thousand lines, which is both the appeal and the ceiling — durable state, workflow definitions and persistence are things you build yourself. A conventional ToolCallingAgent is included for the JSON style where that fits better.
Semantic Kernel puts a kernel at the centre of the application: a container you register model connections and plugins into, which the agents you build then draw on. Mark ordinary methods as kernel functions and the model can call them, import a whole API from an OpenAPI description, wrap filters around a call to log, cache, redact or stop it, and hand a conversation between several specialised agents. The trade-off is direction rather than quality. Microsoft has named Agent Framework the successor and says the majority of new features will be built there, so this codebase keeps getting fixes and will see some existing features reach general availability, but few new ideas. It suits a team extending something already written against it far better than a project starting from nothing today.
Agents are declared with instructions and a model, tools are validated on both ends by a schema, and workflows are chained from steps that can suspend mid-run and resume later because their execution state is written to storage. The breadth is the appeal — retrieval, memory, evaluation and tracing are in the same package rather than assembled from four — but breadth also means a large surface to learn, and the Apache-2.0 grant stops short of a few directories the project reserves under a separate enterprise licence.
A decorator over a typed Python function is a complete tool: the schema is derived from the type hints and docstring, and both the arguments coming in and the value going out are validated for you. Beyond that it covers the parts a second week of work runs into — a typed client, OAuth and remote identity providers, mounting several servers as one, proxying an existing server, and middleware. The confusion to be aware of is the name: an earlier version of this API was absorbed into the official MCP Python SDK, so FastMCP refers to two different things depending on which import you read.
Four functions carry most of the work — generateText and streamText for prose, generateObject and streamObject for data shaped by a schema — and they keep the same signature whether the model behind them is OpenAI, Anthropic, Google, Bedrock or a local one. A second half supplies framework hooks so a streaming chat UI is a few lines rather than a hand-rolled parser. It is a model-access and UI layer, not an agent runtime: persistent memory, retrieval and durable orchestration are somebody else's job.
Haystack builds document-grounded answering and agents out of components — file converters, splitters, embedders, retrievers, chat generators, tools — that you register in a Pipeline and connect output-name to input-name, with a mismatch raising an error at connection time rather than partway through a run. Branches and capped loops live in the same graph, and an Agent is itself a component, so a whole tool-calling loop drops into a pipeline or becomes a tool for another agent. The cost is that the wiring is yours: there is no one-call path from a folder of PDFs to a working assistant, and the library runs inside your own process — serving a pipeline over HTTP or as an MCP server means adding Hayhooks, a separate deepset project, or writing that wrapper yourself.
Treats the context window like RAM and everything else like disk, with the agent itself deciding what to page in and out. That framing makes indefinitely long-running agents tractable. It is a more opinionated commitment than bolting a memory library onto an existing stack.
You write an agent as an LlmAgent — a model, an instruction and a list of tools — then compose several into a hierarchy, or wire them into a workflow whose sequential, parallel and loop steps you control rather than leave to the model. A Runner drives each turn and keeps the session and its state, so a conversation survives across requests. It ships the parts most frameworks leave out: an evaluation harness that scores agents against saved cases, and a local web UI for stepping through a run. It is model-agnostic through LiteLLM and runs anywhere in a container, but the documentation, the tool catalogue and the managed runtime are Google's, so the further you get from Gemini and Google Cloud the more you assemble yourself.
Brings the validation model that Python developers already trust to agent output, so a malformed response fails where you can see it rather than three layers downstream. Deliberately smaller in scope than the all-in-one frameworks; that is the point, and it is a good fit if your codebase is already typed throughout.
You write a character file — a name, a system prompt, a few bio lines and the list of plugins to load — and the runtime turns it into an agent that stores each message, composes context from the providers the turn actually routes to, lets the model pick an action, and files what it learned through evaluators that run after the reply. First-party plugins ship in the same repository for the chat platforms people already use, MCP servers, browser and desktop control, calendar and inbox assistants, non-custodial EVM and Solana wallets, and an on-device model path that answers with the network off. Two things to weigh before adopting it: the repository is a whole product stack — desktop and mobile apps, a hosted cloud, native device bridges — so you take on far more than an agent library, and the stable release on npm is still the older 1.x line while the stack described here ships only under beta and alpha tags.
CAMEL's core idea is agents talking to each other: a RolePlaying session pairs an AI user that issues instructions with an AI assistant that carries them out, and a Workforce hands split-up subtasks to workers under a coordinator. Around that sit eighty-odd toolkits, around fifty model backends, benchmark harnesses such as GAIA and BrowseComp, and pipelines that turn the resulting conversations into fine-tuning data. That breadth is the point if you are running experiments; if you only want one dependable agent in production you will carry a great deal you never use, and the library has been public since 2023 yet is still on 0.2.x.
A request-and-response API cannot describe an agent that runs for minutes, streams partial work, changes its mind and occasionally needs a person. AG-UI defines that exchange as a stream of events over HTTP and WebSockets: tokens, tool calls and their results, shared state, attachments, and interrupts where the run stops to ask. Its usefulness scales with what already implements it — a first-party list that includes several frameworks in this catalogue — and drops sharply if neither your backend nor your frontend does.
Assembles the whole speech-to-speech loop — transcription, model, synthesis, interruption handling — as a pipeline of swappable services, which is the tedious part to get right. Latency is the entire game in voice, and turn-taking and barge-in are first-class here rather than bolted on. Running it well means caring about transport and network topology, not only which model you picked.
Builds server-side agents that join LiveKit rooms as programmable participants, with integrated WebRTC transport, dispatch, telephony, turn detection and provider plugins. It is a strong fit when media transport and production session orchestration need to work as one system. Teams that only need a text agent or want a transport-neutral pipeline may find the LiveKit architecture more infrastructure than necessary.
Written by the teams behind both predecessors, it keeps AutoGen's light agent abstractions and adds Semantic Kernel's session state, middleware and telemetry, then puts graph-based workflows on top for cases where the execution path should be explicit rather than left to a model. For a .NET team this is the first-party answer that previously did not exist. The cost is timing: it is young, and code already running on either predecessor faces a migration rather than an upgrade.
The headline feature is AI Services: you declare a Java interface, annotate it, and the library supplies the implementation that builds the prompt, calls the model and maps the reply back onto your return type. Underneath, a plain ChatModel API is there whenever the declarative layer gets in the way, and RAG, tool calling and integrations for Spring Boot, Quarkus, Micronaut and Helidon are first-party. Despite the name it is not a port of the Python library, so material from that ecosystem does not transfer.
PocketFlow is a reaction to frameworks you cannot see the bottom of. The whole thing is roughly a hundred lines with no dependencies: a node does one piece of work in three phases — gather what it needs, do the work, decide what happens next — and the string a node returns picks which edge to follow. Nodes share one plain dictionary. That is the framework. There is no model client, no retry policy for a provider's particular errors, no tracing; you write those, and the cookbook of worked examples stands in for a reference manual. It suits people who would rather own a small amount of code than configure a large amount of someone else's.
It covers the parts of a spoken agent that sit outside the model and usually get assembled by hand: detecting when someone is speaking, deciding when a turn has ended, separating speakers, connecting to the phone network, driving a lip-synced avatar, even reaching an embedded device. Read the licence before building on it. Apache 2.0 is qualified by additional conditions from Agora that prohibit hosting the framework on end-user devices, mobile terminals included, and prohibit deploying it in a way that competes with Agora's own offerings.
mcp-agent makes a narrow bet: if every tool an agent needs arrives over MCP, the framework can be small. It handles the tedious half of that — opening, holding and closing connections to MCP servers — and then supplies the well-known compositions from Anthropic's "Building Effective Agents" as pieces you can nest: a router that picks a specialist, an orchestrator that plans and delegates, an evaluator that sends work back for another pass. The same agent can be published as an MCP server itself. For runs that must survive a crash it can execute on Temporal, which brings pause and resume at the cost of a service to operate.
Rig is the answer for teams whose service is already Rust and who do not want a Python sidecar just to call a model. Tools are Rust functions, and their argument definitions are generated from the types, so a mismatch is a compile error rather than a malformed call discovered in production. The same applies to output: an extractor asks the model for a value and hands back an ordinary Rust struct. Providers and vector stores sit behind traits, so swapping either is a configuration change. What you give up is the ecosystem — the evaluation, tracing and agent tooling that exists ten times over in Python often has no Rust equivalent, and you will write more of the surrounding machinery yourself.
Swarms pairs one Agent class with a large stock of ready-made ways to run several agents together: a chain, a parallel fan-out, a directed graph, a director handing work to specialists, a panel that votes or debates. One router reaches fifteen of those shapes by name, so trying a different form of collaboration is usually an edit rather than a rewrite — though several also want a setting of their own, and one wants its three agents in a fixed order. Breadth is the cost as well as the appeal: in the 14.0.0 release the Agent constructor alone takes 89 named options, and the reference documentation has already drifted from the code on whether an agent's memory survives a restart. Read the defaults out of the source before the first production run, not out of the page that describes them.
Most guardrail tools check one thing in one place. This defines five interception points instead — before the model sees a request, over what retrieval pulled in, over the shape of the conversation, around tool calls, and before a reply reaches the user — and it sits between the application and the model, so it can be added to a system already in production without changing either side. The costs are latency and vocabulary: several of the stronger rails are model calls of their own, and dialogue rules are written in Colang, a purpose-built language.
Strands Agents takes the opposite bet from the graph frameworks. You supply a system prompt and a handful of ordinary functions marked as tools, and the model decides which to call, in what order, and when to stop. A first agent is a few lines rather than a wiring diagram — and when a run goes somewhere odd there is no graph to point at, so you read the trace instead. Around that loop sit the pieces a long-running agent needs: hooks that fire at every step so you can log or refuse a tool call, conversation managers that trim or summarise history before it reaches the model's limit, and ways to hand one agent to another as a tool.
Reinforcement learning compresses an entire execution into one number, which throws away the part that says what went wrong. GEPA keeps it: error messages, profiling output and reasoning logs go back to a model that diagnoses the failure and proposes a specific edit, and it maintains a frontier of candidate prompts each strong on different examples rather than collapsing to one winner. The project reports reaching results in 100 to 500 evaluations where reinforcement learning needs upwards of 10,000. Both the running and the reflecting are model calls, so the efficiency is relative to RL, not to zero.
When AutoGen went into maintenance mode it left two ways forward, and this is the community one: the same volunteers, the same conversation-driven idea of several agents talking a task through, carried on outside Microsoft. Anyone arriving from AutoGen needs to know what version 1.0 did, though. The classic classes people actually wrote code against — ConversableAgent, GroupChat, the `autogen` import name — have moved to a separate repository maintained alongside this one, and `pip install ag2` now gives you a newer protocol-driven framework instead. Existing code keeps working; it just no longer lives here.