
Midscene.js
UIテストを文章で書き、DOMではなく画面のピクセルから操作する。改修のたびにセレクタを直す作業がなくなる
Midscene.jsとは
UIテストは、製品と関係のない理由で壊れます。クラス名が変わった、要素が1階層深くなった、コンポーネントが移動した、といった理由です。Midsceneはセレクタそのものを無くします。手順は文章で書き、要素の位置はモデルが画面の画像から特定し、検証も「その行が強調されている」「レイアウトが崩れていない」のように人が見て分かる内容で書けます。DOMからは届かない場所にも手が届きます。canvasの描画、アイコンだけのボタン、モバイルのネイティブアプリ、別ドメインのフレームなどです。引き換えになるのは1手順あたりの費用と再現性です。各手順が参照ではなくモデルの呼び出しになるため、セレクタなら数秒で終わる一式が同じ時間では終わりません。
Midscene.jsで何ができますか?
- 要素を特定せず手順を書く — セレクタの代わりに文章を書きます。クラス名の変更やコンポーネントの移動といった改修のたびにテストを直し歩く作業がなくなります。
- 人が見て分かる内容で検証する — 色や強調、レイアウトについて検証できます。要素の存在だけを確かめるテストでは通ってしまう種類の不具合が対象です。
- DOMで表現できない部分に届く — canvasの描画、アイコンだけの操作部品、ネイティブアプリ、別ドメインのフレームも、ここではすべて同じ画素として扱えます。
- 同じ記述をスマートフォンとブラウザに使う — 1つのAPIでWeb、Android、iOS、HarmonyOS、デスクトップを扱えるため、サイトで検証した流れを対応するアプリにも適用できます。
- 画面から構造化されたデータを取り出す — 問い合わせを書くと、画面の内容が文章ではなくデータとして返ります。APIの無い画面からの取得手段としても使えます。
- コードを書く前にブラウザで試す — Chrome拡張から自然言語の指示をページに対して実行できるため、テストファイルを1つも書かずに方式を評価できます。
Midscene.jsを選ぶ前に
- 1手順ごとにモデルを呼ぶため、実行は遅く、手順の数だけ費用がかかり、同じ一式でも結果が毎回同じとは限りません。全ての統合を止めるテストとして使っている場合は影響が出ます。
- 画面上の位置特定に強い多モーダルモデルが必要です。動作確認済みのモデルは明示されており、その一覧に無い汎用モデルではクリック位置がずれます。
よくある質問
Midscene.jsは商用利用できますか?
Midscene.jsはMITライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。
Midscene.jsはどの形で使えますか?
Midscene.jsはローカル実行・セルフホストの形で利用できます。
ドキュメント
web-infra-dev/midscene のREADMEより転載(MIT)。 原文を読む ↗
English | 简体中文
Official Website: https://midscenejs.com/
📣 Midscene Skills is here!
Use Midscene Skills with OpenClaw to test and automate web, mobile, and desktop interfaces.
Showcases
- Web Automation - Automatically register the GitHub form in a web browser and pass all field validations
- iOS Automation - Meituan coffee order
- iOS Automation - Auto-like the first @midscene_ai tweet
- Android Automation - DCar: Xiaomi SU7 specs
- Android Automation - Booking a hotel for Christmas
- robotic arm + vision + voice for in-vehicle testing
💡 Why Midscene
Most UI automation — including AI tools that read the DOM or the accessibility tree — depends on page structure. That structure is fragile and incomplete: selectors break on every refactor, elements without semantic markup (icon-only buttons, custom controls, <canvas>) are invisible to it, native apps and cross-origin iframes are out of reach, and it cannot tell whether something actually looks right. Midscene works from the screenshot alone, and you describe each step in natural language:
- Less maintenance — no selectors to chase when the UI changes.
- Reach every element and surface — if a human can see it, Midscene can target it, even with no semantic annotations, on
<canvas>, native apps, and cross-origin iframes. - Assert what users actually see — verify colors, highlights, layout, and rendered state, not just whether a DOM node exists.
- Two ways to test — add Midscene to your Playwright / Vitest suite, or let an AI agent test autonomously via Skills.
Midscene is built for UI testing first, but the same vision-driven engine handles any UI automation task.
💡 What you can automate
Midscene works anywhere you can take a screenshot — web browsers, Android, iOS, HarmonyOS, desktop apps, and any custom interface — all through one API. Write automation with the JavaScript SDK or in YAML, hand it to AI agents via Skills, and look up every method (aiAct, aiQuery, aiAssert, and more) in the API reference.
🚀 Get started
- Try Midscene in Chrome — use the Quick start to configure a model, install the Chrome extension, and run your first natural-language instruction.
- Write your first script — create an Agent and run a complete browser script with Playwright or Puppeteer.
- Other platforms — getting-started guides for Android, iOS, HarmonyOS, and desktop.
✨ Driven by Multimodal Models
Midscene is all-in on pure vision for UI actions: element localization is based on screenshots only. It runs on multimodal models with strong UI localization, such as Qwen3.x, Doubao-Seed-2.1, GLM-4.6V, gemini-3.5-flash, and UI-TARS, including open-source options you can self-host. For data extraction and page understanding, you can still opt in to include DOM when needed.
Read more about Model Strategy.
📄 Resources
- Documentation: https://midscenejs.com
- Sample projects: midscene-example
- API reference: https://midscenejs.com/reference/#common
🤝 Community
🌟 Awesome Midscene
Community projects that extend Midscene.js capabilities:
- midscene-ios - iOS Mirror automation support for Midscene
- midscene-pc - PC operation device for Windows, macOS, and Linux
- midscene-pc-docker - Docker image with Midscene-PC server pre-installed
- Midscene-Python - Python SDK for Midscene automation
- midscene-java by @Master-Frank - Java SDK for Midscene automation
- midscene-java by @alstafeev - Java SDK for Midscene automation
📝 Credits
We would like to thank the following projects:
- Rsbuild and Rslib for the build tools.
- UI-TARS for the open-source agent model UI-TARS.
- Qwen-VL for the open-source multimodal model Qwen-VL.
- scrcpy and yume-chan for browser-based Android device control.
- appium-adb for its JavaScript bridge to ADB.
- appium-webdriveragent for controlling XCTest from JavaScript.
- YADB for improving text input performance.
- libnut-core for cross-platform native keyboard and mouse control.
- Puppeteer for browser automation and control.
- Playwright for browser automation, control, and testing.
📖 Citation
If you use Midscene.js in your research or project, please cite:
@software{Midscene.js,
author = {Xiao Zhou, Tao Yu, YiBing Lin},
title = {Midscene.js: GUI Agent for E2E Testing.},
year = {2025},
publisher = {GitHub},
url = {https://github.com/web-infra-dev/midscene}
}
✨ Star History
📝 License
Midscene.js is MIT licensed.