← プロジェクト一覧に戻る

Agent Lightning

すでに動いているエージェントを書き換えず、モデルとの間に入って、その中のモデルを学習させる

公式MIT
スター
17.9k
フォーク
1.6k
オープンIssue
152
最終コミット
2026年8月27日

Agent Lightningとは

エージェントに強化学習を使うという話でいつも問題になるのは、学習環境として作り直すことになり、その時点で本番のものを改善しているとは言えなくなる点です。Agent Lightningは、モデルの窓口があった場所に中継役を置くことでこれを避けます。エージェントは自分のツール、プロンプト、制御の流れ、環境をそのまま保ち、中継役を通るやり取りがそのまま学習データになります。学習側が方策を更新し、制御役がエージェントを手元またはKubernetesのジョブとして動かします。Microsoftは、90億パラメータのモデルを6,000件の学習データでSWE-bench Verifiedの41.8%から56.4%へ引き上げたと報告し、その一連の手順を公開しています。これはアプリケーションに足すライブラリではなく、学習基盤を伴うGPUの作業です。

Agent Lightningで何ができますか?

  • すでに動かしているエージェントを学習させる — エージェント側の変更は不要です。ツールもプロンプトも制御の流れもそのまま経路に残るため、良くなるのは実際に配置している仕組みそのものです。
  • 通常のやり取りから学習データを集める — モデルの手前に置いた中継役が、各やり取りを学習用の事例に変えます。人工的な環境ではなく実際の実行からデータが得られます。
  • 必要なら全体を読み通す — コード量はおよそ3,500行だと開発元が示しています。学習の挙動がおかしいときに追える大きさです。
  • 経験収集をKubernetesのジョブとして動かす — エージェントをクラスタ上のジョブとして直接起動できるため、間に別のサンドボックスサービスを挟まずに規模を広げられます。
  • 完結した手順をそのまま辿る — コーディングエージェントの例では、データの整備、報酬の悪用を防ぐ工夫、学習スクリプトまでが一式で公開されています。断片から組み直す必要がありません。

Agent Lightningを選ぶ前に

  • これはモデルの学習です。GPUと、CUDAの版を合わせた強化学習の実行環境が前提になります。アプリケーションにライブラリを足すのとは負担がまったく違います。
  • 1.0で全面的に作り直されており、それ以前の版は別のブランチに置かれています。1.0より前に書かれた解説は別の仕組みについての説明です。

よくある質問

Agent Lightningは商用利用できますか?

Agent LightningはMITライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。

Agent Lightningはどの形で使えますか?

Agent Lightningはセルフホストの形で利用できます。

ドキュメント

microsoft/agent-lightning のREADMEより転載(MIT)。 原文を読む ↗

Agent Lightning was completely refactored in v1.0. For legacy releases earlier than v1.0, see this branch.

⚡ Key Features

  • 🪶 ~3,500 lines of code: We treat simplicity as the first principle.
  • 🧩 Train with real agent harnesses: Agents interact with the model through the Agent Lightning v1.0 proxy with ZERO changes, while keeping tools, context, control flow, and environments in the loop.
  • ☸️ Native Kubernetes support: Run agents directly as Kubernetes Jobs without relying on external sandbox services.
  • 💻 Full coding agent training example: Using only 6K training samples, an end-to-end Qwen3.5-9B workflow improves SWE-bench Verified from 41.8% to 56.4%, a gain of 14.6 percentage points. We release the full pipeline, including data cleaning, reward-hacking prevention, and training scripts.

⚡ Installation

The following is an example installation on a CUDA 13.0 machine:

cd <this-repo>
uv sync
bash scripts/setup_verl.sh 0.8.0 cu130

See the Installation Guide for details.

⚡ Architecture

Agent Lightning v1.0 keeps the training architecture simple with three lightweight components:

  • Trainer: Runs verl and vLLM, builds training samples, and updates the policy.
  • API Gateway: Proxies model requests and captures training data.
  • Rollout Controller: Runs agents locally or as Kubernetes Jobs.

The Trainer creates rollouts, the Controller launches agents, and the Gateway turns interactions into training data, while agents continue to run with their real harnesses.

⚡ Results

We evaluate Agent Lightning v1.0 across several practical training domains, including Search R1, LLM-in-Sandbox, and Coding Agent. Pure RL delivers substantial improvements across all three domains, as shown below.

⚡ Documentation

SectionContent
InstallationBase environment and verl GPU stack
Quick StartLocal first run and end-to-end flow
BasicsComponents, rollouts, events, and trajectories
Trainer Configurationverl integration and trace aggregation
API Gateway ConfigurationGateway and model proxy settings
Controller ConfigurationLocal and Kubernetes runners
Asynchronous TrainingCollocated async collection and pause/drain

⚡ Examples

ExampleDescription
Calc-XPOC math reasoning example with AutoGen and MCP calculator tools, requiring only one GPU.
GSM8KPOC grade-school math reasoning example.
ScienceWorldInteractive science tasks in a text-based environment.
Search-R1Multi-turn retrieval and reasoning agent.
LLM-in-SandboxGeneral agent with computer and code execution tools.
Coding AgentCoding agent trained with repository tests.

⚡ Articles

⚡ Community Projects

⚡ Citation

If you use Agent Lightning v1.0 in your research or projects, please cite the technical report:

@misc{he2026agentlightningv10harnessed,
  title={Agent Lightning v1.0: Towards Harnessed Agentic RL},
  author={Zhiyuan He and Siwei Zhang and Zhiwen Zhou and Yuqing Yang and Yu Kang and Yuge Zhang and Luna K. Qiu and Tin Yan Tsui and Jiahang Xu and Chong Luo},
  year={2026},
  eprint={2608.17528},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2608.17528},
}

For the original Agent Lightning paper, please use:

@misc{luo2025agentlightningtrainai,
      title={Agent Lightning: Train ANY AI Agents with Reinforcement Learning},
      author={Xufang Luo and Yuge Zhang and Zhiyuan He and Zilong Wang and Siyun Zhao and Dongsheng Li and Luna K. Qiu and Yuqing Yang},
      year={2025},
      eprint={2508.03680},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2508.03680},
}

⚡ Contributing

This project welcomes contributions and suggestions. Start by reading the Contributing Guide for recommended contribution points, environment setup, branching conventions, and pull request expectations. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.opensource.microsoft.com.

When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

⚡ Trademarks

This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft’s Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party’s policies.

⚡ Responsible AI

This project has been evaluated and certified to comply with the Microsoft Responsible AI Standard. The team will continue to monitor and maintain the repository, addressing any severe issues, including potential harms, if they arise.

⚡ License

Agent Lightning v1.0 is released under the MIT License.

Agent Lightning
AIに聞く
GitHub