Multi-Agent Orchestration
Coordinating several agents or steps: routing, delegation, parallel execution and merging results. Worth reaching for only once a single agent has demonstrably hit its limit.
18 projects
An agent here is a graph of blocks — one block calls a service, transforms data, prompts a model or branches — assembled on a canvas rather than written in code, then run on demand, on a schedule or from a trigger. One repository holds both the free self-hosted platform and the code behind the paid hosted one, and that is the catch: everything under the platform directory is PolyForm Shield, which permits use but forbids offering it as a competing product.
Version 2.0 is a complete application behind one port — web UI, gateway backend and all — where a lead agent spawns sub-agents, keeps persistent memory, executes work in sandboxes from local Docker to Kubernetes and E2B, and produces reports, slide decks, web pages and video through Markdown-defined skills. Models wire up through LangChain provider classes, search works keyless via DuckDuckGo and swaps for six alternatives, and the agent answers from Telegram, Slack, Feishu and other chat channels. Two facts should precede adoption: the celebrated v1 deep-research pipeline is frozen on a separate branch and is not what installs today, and the README scopes the design to loopback-only trusted use — gateway admin is equivalent to code execution on the host.
Hand MetaGPT a one-line requirement and a team of LLM roles works through it in turn, leaving a requirements document, a system design and source files in a folder under `workspace/`. Which roles show up depends on where you installed from, and the gap is not cosmetic: the published 0.8.2 package hires a product manager, an architect, a project manager and an engineer, while current main — the README's clone-and-install route — hires a team leader, a product manager, an architect, an engineer and a data analyst, with the project manager commented out. Under the demo sits a general framework: you subclass Role and Action, declare which upstream actions each role watches, and a shared Environment routes the messages that put the team in order. A separate agent, the Data Interpreter, does data work alone, planning the steps, writing Python and retrying what fails. The catch is scope — MetaGPT's own FAQ says functions start being left unimplemented past roughly 500 lines of generated code.
Microsoft Research's take on multi-agent systems, where progress emerges from a structured conversation between specialised agents. Strong for exploration and research-flavoured problems. Two things to check before adopting: the API was substantially reworked between major versions, and GitHub reports the repository licence as CC-BY-4.0 — a content licence, not an OSI-approved software one. Confirm the terms with the maintainers before commercial use.
Models a task as a crew of agents with named roles and goals that hand work to each other. The metaphor makes a multi-agent design readable at a glance, which is why it demos so well. Whether several role-playing agents beat one well-prompted agent on your task is worth measuring before committing — the answer is often no.
Agno is three layers you can adopt one at a time: a Python SDK for agents, teams and workflows; AgentOS, a FastAPI service that turns them into a REST API with sessions, JWT-scoped access, scheduling and OpenTelemetry traces written to a database you own; and a web control plane for watching runs. The first two are Apache-2.0 and run wherever containers run; apart from a usage ping you can switch off, nothing has to reach Agno. The control plane is the part that does not follow that rule — free against a runtime on your own machine, and a paid plan to point it at a deployed one. The project also moves quickly, with breaking changes in most minor releases, so pin a version before you build on it.
LangGraph has you define the agent's flow as a graph. Nodes are the units of work — an LLM call, a tool run — edges decide which node runs next, and every node reads and writes one shared state. Edges can branch on that state to form loops, and the state is checkpointed at each node, so a run can stop for a person's approval and resume later, or restart from the last checkpoint after a crash. The cost is code: a simple tool-calling agent takes noticeably more setup here than elsewhere.
ChatDev 2.0 builds a multi-agent system as a graph you draw in the browser or write as YAML: agent, python, human and subgraph nodes joined by edges that can branch, loop and fan out. You start a run from the Launch tab, watch each node move from pending to running to success or failed, answer a human node mid-run, and the same file runs from Python through the chatdev package on PyPI. The virtual software company from the 2023 paper is still here, rebuilt as one of the shipped workflows with nine agents from Chief Executive Officer to Software Test Engineer, but the 1.x code behind that paper now sits on a separate branch. The platform itself is young, and its own guides already disagree with the code on field names, so expect to learn it from the sample YAML files rather than the prose.
AgentScope starts where most frameworks start — an agent that reasons, calls a tool, looks at the result and goes round again — and then supplies the parts that usually have to be built afterwards. Tools run inside an isolated environment rather than in your process; context is compressed and offloaded as it grows instead of simply overflowing; a person can be asked to approve a step; and the finished agent can be served as a multi-tenant backend with its own interface. The price of that reach is that version 2.0 was a deliberate break from 1.0, so material written for the older releases does not carry over.
Semantic Kernel puts a kernel at the centre of the application: a container you register model connections and plugins into, which the agents you build then draw on. Mark ordinary methods as kernel functions and the model can call them, import a whole API from an OpenAPI description, wrap filters around a call to log, cache, redact or stop it, and hand a conversation between several specialised agents. The trade-off is direction rather than quality. Microsoft has named Agent Framework the successor and says the majority of new features will be built there, so this codebase keeps getting fixes and will see some existing features reach general availability, but few new ideas. It suits a team extending something already written against it far better than a project starting from nothing today.
Haystack builds document-grounded answering and agents out of components — file converters, splitters, embedders, retrievers, chat generators, tools — that you register in a Pipeline and connect output-name to input-name, with a mismatch raising an error at connection time rather than partway through a run. Branches and capped loops live in the same graph, and an Agent is itself a component, so a whole tool-calling loop drops into a pipeline or becomes a tool for another agent. The cost is that the wiring is yours: there is no one-call path from a folder of PDFs to a working assistant, and the library runs inside your own process — serving a pipeline over HTTP or as an MCP server means adding Hayhooks, a separate deepset project, or writing that wrapper yourself.
You write an agent as an LlmAgent — a model, an instruction and a list of tools — then compose several into a hierarchy, or wire them into a workflow whose sequential, parallel and loop steps you control rather than leave to the model. A Runner drives each turn and keeps the session and its state, so a conversation survives across requests. It ships the parts most frameworks leave out: an evaluation harness that scores agents against saved cases, and a local web UI for stepping through a run. It is model-agnostic through LiteLLM and runs anywhere in a container, but the documentation, the tool catalogue and the managed runtime are Google's, so the further you get from Gemini and Google Cloud the more you assemble yourself.
CAMEL's core idea is agents talking to each other: a RolePlaying session pairs an AI user that issues instructions with an AI assistant that carries them out, and a Workforce hands split-up subtasks to workers under a coordinator. Around that sit eighty-odd toolkits, around fifty model backends, benchmark harnesses such as GAIA and BrowseComp, and pipelines that turn the resulting conversations into fine-tuning data. That breadth is the point if you are running experiments; if you only want one dependable agent in production you will carry a great deal you never use, and the library has been public since 2023 yet is still on 0.2.x.
Written by the teams behind both predecessors, it keeps AutoGen's light agent abstractions and adds Semantic Kernel's session state, middleware and telemetry, then puts graph-based workflows on top for cases where the execution path should be explicit rather than left to a model. For a .NET team this is the first-party answer that previously did not exist. The cost is timing: it is young, and code already running on either predecessor faces a migration rather than an upgrade.
It covers the parts of a spoken agent that sit outside the model and usually get assembled by hand: detecting when someone is speaking, deciding when a turn has ended, separating speakers, connecting to the phone network, driving a lip-synced avatar, even reaching an embedded device. Read the licence before building on it. Apache 2.0 is qualified by additional conditions from Agora that prohibit hosting the framework on end-user devices, mobile terminals included, and prohibit deploying it in a way that competes with Agora's own offerings.
You hand XAgent a goal in one sentence and it takes over from there. An outer loop plans, breaking the goal into subtasks; an inner loop acts, working each subtask with a shell, a Python notebook, a web browser and a file editor. The project calls this its dual-loop mechanism and names three parts: a dispatcher that picks the agent for each stage, a planner and an actor. The tools live in Docker containers the stack starts for you, so the agent works in a container's own workspace and the files it makes are pulled back afterwards, and between subtasks the planner can reopen the plan and split, add or delete the steps still ahead. The catch is that this is a research demonstrator from late 2023 rather than a maintained product: every model call is written in OpenAI's function-calling format, the shipped configuration still names a model OpenAI has shut down, and feature work stopped in early 2024. Read it as a clearly documented reference implementation of the planner-and-actor split, and expect repair work before it runs at all.
Swarms pairs one Agent class with a large stock of ready-made ways to run several agents together: a chain, a parallel fan-out, a directed graph, a director handing work to specialists, a panel that votes or debates. One router reaches fifteen of those shapes by name, so trying a different form of collaboration is usually an edit rather than a rewrite — though several also want a setting of their own, and one wants its three agents in a fixed order. Breadth is the cost as well as the appeal: in the 14.0.0 release the Agent constructor alone takes 89 named options, and the reference documentation has already drifted from the code on whether an agent's memory survives a restart. Read the defaults out of the source before the first production run, not out of the page that describes them.
When AutoGen went into maintenance mode it left two ways forward, and this is the community one: the same volunteers, the same conversation-driven idea of several agents talking a task through, carried on outside Microsoft. Anyone arriving from AutoGen needs to know what version 1.0 did, though. The classic classes people actually wrote code against — ConversableAgent, GroupChat, the `autogen` import name — have moved to a separate repository maintained alongside this one, and `pip install ag2` now gives you a newer protocol-driven framework instead. Existing code keeps working; it just no longer lives here.