← Back to all projects

DeerFlow

ByteDance's self-hosted general agent — a finished application, no longer the deep-research framework it began as

OfficialMIT
Stars
81.1k
Forks
11.2k
Open issues
916
Last commit
28 Aug 2026

What is DeerFlow?

Version 2.0 is a complete application behind one port — web UI, gateway backend and all — where a lead agent spawns sub-agents, keeps persistent memory, executes work in sandboxes from local Docker to Kubernetes and E2B, and produces reports, slide decks, web pages and video through Markdown-defined skills. Models wire up through LangChain provider classes, search works keyless via DuckDuckGo and swaps for six alternatives, and the agent answers from Telegram, Slack, Feishu and other chat channels. Two facts should precede adoption: the celebrated v1 deep-research pipeline is frozen on a separate branch and is not what installs today, and the README scopes the design to loopback-only trusted use — gateway admin is equivalent to code execution on the host.

What can you do with DeerFlow?

  • One make command stands the whole thing up — make setup interviews you and writes the configuration, make doctor verifies it, and make docker-start serves the web UI on a single port; a TUI and an embeddable Python client cover the exceptions. The sizing table is honest — it starts at 4 vCPUs and 8 GB of RAM and recommends double for Docker development.
  • Sub-agents work inside sandboxes you choose — The lead agent delegates to sub-agents whose execution lands in a sandbox — local process, Docker, Kubernetes or E2B — so letting an agent run code becomes a deployment decision rather than an act of faith.
  • Skills define what a run can produce — Research reports, slide decks, web pages, images and video come from Markdown-defined skills that load progressively, activate with a slash command, and install as archives. Model wiring reaches OpenAI-compatible gateways, vLLM, and CLI-backed providers such as Claude Code's own sign-in.
  • Reachable from the chat apps you already run — Connectors for Telegram, Slack, Discord, Feishu, WeChat, WeCom and DingTalk let the agent take work from a chat thread without the server ever being exposed to the internet.

Before you choose DeerFlow

  • The deep-research DeerFlow that made the project famous is v1, wrapped up on the main-1.x branch — installing today gets the 2.0 rewrite, which shares no code with it.Reported in#824
  • Security is scoped to trusted loopback-only deployment by the README's own account, and skill gating is behavioural rather than a hard boundary — a shell command can reach what a disabled skill could not.Reported in#4107
  • There has been exactly one release: users track the main branch of a rewrite that went GA in June 2026, moving at over a hundred commits in the first three weeks of August.

Star history

19 Aug to 28 Aug · +744

80.3k81.1k

Frequently asked questions

Is DeerFlow free for commercial use?

DeerFlow is released under the MIT licence — OSI-approved open source, which permits commercial use.

How can DeerFlow be deployed?

DeerFlow is available as Self-hosted / Runs locally.

Documentation

Reproduced from the bytedance/deer-flow README, published under MIT. Read the original ↗

🦌 DeerFlow - 2.0

English | 中文 | 日本語 | Français | Русский

On February 28th, 2026, DeerFlow claimed the 🏆 #1 spot on GitHub Trending following the launch of version 2. Thanks a million to our incredible community — you made this happen! 💪🔥

DeerFlow (Deep Exploration and Efficient Research Flow) is an open-source super agent harness that orchestrates sub-agents, memory, and sandboxes to do almost anything — powered by extensible skills.

https://github.com/user-attachments/assets/a8bcadc4-e040-4cf2-8fda-dd768b999c18

[!NOTE] DeerFlow 2.0 is a ground-up rewrite. It shares no code with v1. If you’re looking for the original Deep Research framework, it’s maintained on the 1.x branch — contributions there are still welcome. Active development has moved to 2.0.

Official Website

Learn more and see real demos on our official website. The landing-page case studies open as allowlisted, read-only showcases without requiring a sign-in.

Sister Projects

  • LLM Space - Meet our secret weapon behind DeerFlow — one desktop tool to prototype agent ideas, inspect each harness step, replay failures, and benchmark performance.

Coding Plan from ByteDance Volcengine

InfoQuest

DeerFlow has newly integrated the intelligent search and crawling toolset independently developed by BytePlus—InfoQuest (supports free online experience)


Table of Contents

One-Line Agent Setup

If you use Claude Code, Codex, Cursor, Windsurf, or another coding agent, you can hand it the setup instructions in one sentence:

Help me clone DeerFlow if needed, then bootstrap it for local development by following https://raw.githubusercontent.com/bytedance/deer-flow/main/Install.md

That prompt is intended for coding agents. It tells the agent to clone the repo if needed, choose Docker when available, and stop with the exact next command plus any missing config the user still needs to provide.

Quick Start

Configuration

  1. Clone the DeerFlow repository

    git clone https://github.com/bytedance/deer-flow.git
    cd deer-flow
  2. Run the setup wizard

    From the project root directory (deer-flow/), run:

    make setup

    This launches an interactive wizard that guides you through choosing an LLM provider, optional web search, and execution/safety preferences such as sandbox mode, bash access, and file-write tools. It generates a minimal config.yaml and writes your keys to .env. Takes about 2 minutes.

    The wizard also lets you configure an optional web search provider, or skip it for now.

    Run make doctor at any time to verify your setup and get actionable fix hints. If you are opening a GitHub issue about a local setup or runtime problem, run make support-bundle. The command prints reporter next steps, writes a *-issue-summary.md file to paste into the issue, a *-issue-draft.md file for AI-assisted issue filing, and an optional evidence zip under .deer-flow/support-bundles/. If an AI assistant files the issue, start from the draft and replace every REQUIRED placeholder instead of inventing missing facts. Attach the zip only if a maintainer asks for it, or if the summary alone is not enough. Maintainers and AI triage tools can start with triage.json; the bundle includes redacted diagnostics and file manifests only, and does not include .env, raw conversation messages, or user file contents.

    Advanced / manual configuration: If you prefer to edit config.yaml directly, run make config instead to copy the full template. See config.example.yaml for the complete reference including CLI-backed providers (Codex CLI, Claude Code OAuth), OpenRouter, Responses API, subagent runtime caps such as subagents.max_total_per_run, and more.

    Optional per-model pricing must use one currency across all priced models. DeerFlow disables Console cost estimates when currencies are mixed rather than presenting an invalid aggregate.

    models:
      - name: gpt-4o
        display_name: GPT-4o
        use: langchain_openai:ChatOpenAI
        model: gpt-4o
        api_key: $OPENAI_API_KEY
    
      - name: openrouter-gemini-2.5-flash
        display_name: Gemini 2.5 Flash (OpenRouter)
        use: langchain_openai:ChatOpenAI
        model: google/gemini-2.5-flash-preview
        api_key: $OPENROUTER_API_KEY
        base_url: https://openrouter.ai/api/v1
    
      - name: gpt-5-responses
        display_name: GPT-5 (Responses API)
        use: langchain_openai:ChatOpenAI
        model: gpt-5
        api_key: $OPENAI_API_KEY
        use_responses_api: true
        output_version: responses/v1
    
      - name: qwen3-32b-vllm
        display_name: Qwen3 32B (vLLM)
        use: deerflow.models.vllm_provider:VllmChatModel
        model: Qwen/Qwen3-32B
        api_key: $VLLM_API_KEY
        base_url: http://localhost:8000/v1
        supports_thinking: true
        when_thinking_enabled:
          extra_body:
            chat_template_kwargs:
              enable_thinking: true

    OpenRouter and similar OpenAI-compatible gateways should be configured with langchain_openai:ChatOpenAI plus base_url. If you prefer a provider-specific environment variable name, point api_key at that variable explicitly (for example api_key: $OPENROUTER_API_KEY).

    To route OpenAI models through /v1/responses, keep using langchain_openai:ChatOpenAI and set use_responses_api: true with output_version: responses/v1.

    For vLLM 0.19.0, use deerflow.models.vllm_provider:VllmChatModel. For Qwen-style reasoning models, DeerFlow toggles reasoning with extra_body.chat_template_kwargs.enable_thinking and preserves vLLM’s non-standard reasoning field across multi-turn tool-call conversations. Legacy thinking configs are normalized automatically for backward compatibility. If the endpoint reports a cumulative usage snapshot on every streaming chunk, set cumulative_stream_usage: true so DeerFlow converts those snapshots into per-chunk deltas; the option is disabled by default and leaves usage unchanged when a stable completion id is unavailable. Reasoning models may also require the server to be started with --reasoning-parser .... If your local vLLM deployment accepts any non-empty API key, you can still set VLLM_API_KEY to a placeholder value.

    CLI-backed provider examples:

    models:
      - name: gpt-5.4
        display_name: GPT-5.4 (Codex CLI)
        use: deerflow.models.openai_codex_provider:CodexChatModel
        model: gpt-5.4
        supports_thinking: true
        supports_reasoning_effort: true
    
      - name: claude-sonnet-4.6
        display_name: Claude Sonnet 4.6 (Claude Code OAuth)
        use: deerflow.models.claude_provider:ClaudeChatModel
        model: claude-sonnet-4-6
        max_tokens: 4096
        supports_thinking: true
    • Codex CLI reads ~/.codex/auth.json
    • Claude Code accepts CLAUDE_CODE_OAUTH_TOKEN, ANTHROPIC_AUTH_TOKEN, CLAUDE_CODE_CREDENTIALS_PATH, or ~/.claude/.credentials.json
    • ACP agent entries are separate from model providers — if you configure acp_agents.codex, point it at a Codex ACP adapter such as npx -y @zed-industries/codex-acp
    • MiniMax Code speaks ACP directly. Install and authenticate it, then add it as an ACP agent:
    npm install --global @minimax-ai/code
    mcode login
    acp_agents:
      mcode:
        command: mcode
        args: ["acp"]
        description: MiniMax Code for implementation, refactoring, debugging, and repository tasks
        auto_approve_permissions: false

    mcode must be on the Gateway process’s PATH; installing it only on the Docker host does not make it available inside the Gateway container. DeerFlow invokes it through invoke_acp_agent in a per-thread ACP workspace and forwards enabled MCP servers. Keep auto_approve_permissions: false for untrusted tasks; enable it only when MCode must edit files or run commands and you trust the task.

    • On macOS, export Claude Code auth explicitly if needed:
    eval "$(python3 scripts/export_claude_code_oauth.py --print-export)"

    API keys can also be set manually in .env (recommended) or exported in your shell:

    OPENAI_API_KEY=your-openai-api-key
    TAVILY_API_KEY=your-tavily-api-key

Running the Application

Deployment Sizing

Use the table below as a practical starting point when choosing how to run DeerFlow:

Deployment targetStarting pointRecommendedNotes
Local evaluation / make dev4 vCPU, 8 GB RAM, 20 GB free SSD8 vCPU, 16 GB RAMGood for one developer or one light session with hosted model APIs. 2 vCPU / 4 GB is usually not enough.
Docker development / make docker-start4 vCPU, 8 GB RAM, 25 GB free SSD8 vCPU, 16 GB RAMImage builds, bind mounts, and sandbox containers need more headroom than pure local dev.
Long-running server / make up8 vCPU, 16 GB RAM, 40 GB free SSD16 vCPU, 32 GB RAMPreferred for shared use, multi-agent runs, report generation, or heavier sandbox workloads.
  • These numbers cover DeerFlow itself. If you also host a local LLM, size that service separately.
  • Linux plus Docker is the recommended deployment target for a persistent server. macOS and Windows are best treated as development or evaluation environments.
  • If CPU or memory usage stays pinned, reduce concurrent runs first, then move to the next sizing tier.

Development (hot-reload, source mounts):

make docker-init    # Pull sandbox image (only once or when image updates)
make docker-start   # Start services (auto-detects sandbox mode from config.yaml)
make docker-logs    # View logs

make docker-start starts provisioner only when config.yaml uses provisioner mode (sandbox.use: deerflow.community.aio_sandbox:AioSandboxProvider with provisioner_url).

Docker builds use the upstream uv registry by default. If you need faster mirrors in restricted networks, export UV_INDEX_URL=https://pypi.tuna.tsinghua.edu.cn/simple and NPM_REGISTRY=https://registry.npmmirror.com before running make docker-init or make docker-start.

Local AIO sandbox control traffic is always direct: loopback/private addresses, single-label cluster hosts, and Docker/Podman internal hostnames do not inherit HTTP_PROXY or HTTPS_PROXY. External sandbox FQDNs and public IPs still honor environment proxy settings.

Backend processes automatically pick up config.yaml changes on the next config access, so model metadata updates do not require a manual restart during development. The checkpoint storage settings database.checkpoint_channel_mode and database.checkpoint_delta.snapshot_frequency (default 10) are exceptions: both are frozen when the process first builds an agent (including through DeerFlowClient) and require a process restart to change safely.

The optional database.checkpoint_cache section (delta channel mode only) caches materialized checkpoint histories: type is memory (default) or redis, and max_entries: 0 disables the cache. The redis backend is Gateway/async-only; the sync TUI/embedded path supports memory only. The cache is performance-only — results are identical with it disabled — so it is never frozen and workers sharing one checkpoint database may safely run different cache settings.

[!TIP] On Linux, if Docker-based commands fail with permission denied while trying to connect to the Docker daemon socket at unix:///var/run/docker.sock, add your user to the docker group and re-login before retrying. See CONTRIBUTING.md for the full fix.

Production (builds images locally, mounts runtime config and data):

make up     # Build images and start all production services
make down   # Stop and remove containers

Access: http://localhost:2026

make up waits for the Gateway /health endpoint before reporting success. If the Gateway does not become healthy within the startup window, deployment exits non-zero and prints the container status plus recent Gateway logs. The production image starts from its already-built environment and never resolves or installs Python dependencies at container startup.

For persistent deployments, configure database.backend as sqlite or postgres. The selected backend is shared by the LangGraph checkpointer, LangGraph Store, and DeerFlow application data. The deprecated checkpointer section, when present, overrides the first two for backward compatibility.

The unified nginx endpoint is same-origin by default and does not emit browser CORS headers. If you run a split-origin or port-forwarded browser client, set GATEWAY_CORS_ORIGINS to comma-separated exact origins such as http://localhost:3000; the Gateway then applies the CORS allowlist and matching CSRF origin checks.

Browser login uses HttpOnly session cookies. The login page offers a “keep me signed in” option that extends the browser session when the request is HTTPS (including trusted X-Forwarded-Proto: https) or localhost HTTP. The localhost exception uses the direct request Host and ignores forwarded host headers. Public HTTP deployments, including many temporary sandbox URLs, fall back to session cookies by default. DeerFlow never stores the password in browser storage; the UI may remember only the email address.

DeerFlow still uses Forwarded / X-Forwarded-* headers to recover the browser-facing scheme and origin behind a proxy. The bundled nginx sets X-Forwarded-Proto, but preserves an upstream HTTPS value and does not overwrite every forwarded header. Configure the outer trusted proxy to replace or strip client-supplied forwarding headers before traffic reaches DeerFlow.

[!IMPORTANT] The Gateway still owns active run tasks in process, so production defaults to a single Gateway worker (GATEWAY_WORKERS=1). Multi-worker deployments require Postgres, the Redis stream bridge (stream_bridge.type: redis), run_ownership.heartbeat_enabled: true, and run_events.backend: db; process-local memory/JSONL event stores cannot enforce singleton delivery receipts across workers. The bridge shares SSE delivery and bounded Last-Event-ID replay across workers. When a valid reconnect cursor has been trimmed, or a subscriber that already established an empty-stream wait falls behind before its first delivery, Memory and Redis emit a machine-readable SSE gap event instead of silently returning a partial replay; the Web UI reloads durable thread/event state and resumes from the retained tail. Lease reconciliation marks runs from dead workers as errors, persists their delivery receipts, publishes the terminal stream marker, schedules retained-stream cleanup, and updates the affected thread status. SSE and /wait consumers also refresh durable status on heartbeats as a fallback if terminal publication fails. Malformed Redis reconnect IDs live-tail new events instead of replaying the retained buffer, and the rolling retained-buffer TTL (stream_ttl_seconds) remains a cleanup safety net rather than a run timeout. IM channel state and other process-local services still need their own multi-worker coordination.

Run cancellation may land on any Gateway worker. A non-owning worker now persists the interrupt or rollback request for the live owner, which observes it during lease renewal and performs the normal cancellation flow; load-balancer routing alone no longer produces a 409. The first accepted action wins even if a retry lands on the owner, and accepted cancellation competes atomically with owner completion. Dead owners still follow lease takeover and orphan recovery. Cancellation latency is therefore bounded by the lease heartbeat interval.

With lease heartbeat enabled, a transient RunStore renewal error is retried only until the last confirmed lease expires; the stale worker then cancels local execution and suppresses checkpoint, completion-hook, delivery-receipt, and thread-status finalization. A remote tool side effect already in flight may still be outside local cancellation.

Reconciliation uses an atomic takeover claim that re-checks the lease after candidate selection, so a successful owner renewal wins over orphan recovery and only one reconciler can report a run as recovered. When multiple Gateway workers share the Docker/AIO or E2B sandbox backend, also configure sandbox.ownership.type: redis; E2B uses the leases during background startup and periodic reconciliation so duplicate/orphan cleanup cannot terminate a live peer’s sandbox.

See CONTRIBUTING.md for detailed Docker development guide.

Option 2: Local Development

If you prefer running services locally:

Prerequisite: complete the “Configuration” steps above first (make setup). make dev requires a valid config.yaml in the project root. Set DEER_FLOW_PROJECT_ROOT to define that root explicitly, or DEER_FLOW_CONFIG_PATH to point at a specific config file. Runtime state defaults to .deer-flow under the project root and can be moved with DEER_FLOW_HOME; skills default to skills/ under the project root and can be moved with DEER_FLOW_SKILLS_PATH. Run make doctor to verify your setup before starting. On Windows, run the local development flow from Git Bash. Native cmd.exe and PowerShell shells are not supported for the bash-based service scripts, and WSL is not guaranteed because some scripts rely on Git for Windows utilities such as cygpath.

  1. Check prerequisites:

    make check  # Verifies Node.js 22+, pnpm, uv, nginx

    The local make check, make install, make dev, and make start entry points use a direct pnpm/pnpm.cmd executable when available and otherwise fall back to corepack pnpm. The shared runner and diagnostics resolve repository paths absolutely, so these checks work regardless of the caller’s current directory. Corepack runs from frontend/, so it honors the packageManager version pinned in frontend/package.json; enabling a global pnpm shim is not required.

  2. Install dependencies:

    make install  # Install backend + frontend dependencies + pre-commit hooks
  3. (Optional) Pre-pull sandbox image:

    # Recommended if using Docker/Container-based sandbox
    make setup-sandbox
  4. (Optional) Load sample memory data for local review:

    python scripts/load_memory_sample.py

    This copies the sample fixture into the default local runtime memory file so reviewers can immediately test Settings > Memory. See backend/docs/MEMORY_SETTINGS_REVIEW.md for the shortest review flow.

  5. Start services:

    make dev
  6. Access: http://localhost:2026

Local services always use their internal ports (8001, 3000, and 2026). The root .env variable PORT configures only the published Docker ingress; it does not change the Next.js port used by make dev.

This README has been shortened. The full version is on GitHub. Read the original ↗