
Midscene.js
Writes UI tests as sentences and drives the screen from the pixels, so there are no selectors to fix after a refactor
What is Midscene.js?
UI tests break for reasons that have nothing to do with the product: a class name changed, a wrapper was added, a component moved. Midscene removes the selector entirely. Steps are written as sentences, the model locates elements from the screenshot, and assertions can be about what a person would see — the row is highlighted, the layout has not collapsed — rather than about whether a node exists. That also reaches surfaces the DOM cannot: canvas drawings, icon-only buttons, native mobile applications, cross-origin frames. The trade is per-step cost and determinism, because each step is now a model call rather than a lookup, so a suite that runs in seconds with selectors will not run in seconds here.
What can you do with Midscene.js?
- Describe the step instead of locating the element — A sentence replaces the selector, so a refactor that renames classes or moves a component does not send you back through the test suite.
- Assert what a person would see — Checks can be about colour, highlighting and layout — the kinds of breakage that pass a test asserting only that an element exists.
- Reach the parts the DOM cannot describe — Canvas drawings, icon-only controls, native applications and cross-origin frames are all just pixels here, so they are as reachable as anything else.
- Use the same script on phone and browser — One API covers web, Android, iOS, HarmonyOS and desktop, so a flow tested on the site can be run against the app it mirrors.
- Pull structured data off a page — A query returns what is on screen as data rather than text, which turns a test tool into a way of extracting from interfaces with no API.
- Start in the browser before writing code — A Chrome extension runs natural-language instructions against a page, so the approach can be evaluated before any test file exists.
Before you choose Midscene.js
- Every step is a model call, so runs are slower and cost money per step, and two runs of the same suite can differ — which is a real change if your tests currently gate every merge.
- It needs a multimodal model that is good at locating things on screen; the project names the ones it works with, and a general-purpose model that is not on that list will place clicks badly.
Frequently asked questions
Is Midscene.js free for commercial use?
Midscene.js is released under the MIT licence — OSI-approved open source, which permits commercial use.
How can Midscene.js be deployed?
Midscene.js is available as Runs locally / Self-hosted.
Documentation
Reproduced from the web-infra-dev/midscene README, published under MIT. Read the original ↗
English | 简体中文
Official Website: https://midscenejs.com/
📣 Midscene Skills is here!
Use Midscene Skills with OpenClaw to test and automate web, mobile, and desktop interfaces.
Showcases
- Web Automation - Automatically register the GitHub form in a web browser and pass all field validations
- iOS Automation - Meituan coffee order
- iOS Automation - Auto-like the first @midscene_ai tweet
- Android Automation - DCar: Xiaomi SU7 specs
- Android Automation - Booking a hotel for Christmas
- robotic arm + vision + voice for in-vehicle testing
💡 Why Midscene
Most UI automation — including AI tools that read the DOM or the accessibility tree — depends on page structure. That structure is fragile and incomplete: selectors break on every refactor, elements without semantic markup (icon-only buttons, custom controls, <canvas>) are invisible to it, native apps and cross-origin iframes are out of reach, and it cannot tell whether something actually looks right. Midscene works from the screenshot alone, and you describe each step in natural language:
- Less maintenance — no selectors to chase when the UI changes.
- Reach every element and surface — if a human can see it, Midscene can target it, even with no semantic annotations, on
<canvas>, native apps, and cross-origin iframes. - Assert what users actually see — verify colors, highlights, layout, and rendered state, not just whether a DOM node exists.
- Two ways to test — add Midscene to your Playwright / Vitest suite, or let an AI agent test autonomously via Skills.
Midscene is built for UI testing first, but the same vision-driven engine handles any UI automation task.
💡 What you can automate
Midscene works anywhere you can take a screenshot — web browsers, Android, iOS, HarmonyOS, desktop apps, and any custom interface — all through one API. Write automation with the JavaScript SDK or in YAML, hand it to AI agents via Skills, and look up every method (aiAct, aiQuery, aiAssert, and more) in the API reference.
🚀 Get started
- Try Midscene in Chrome — use the Quick start to configure a model, install the Chrome extension, and run your first natural-language instruction.
- Write your first script — create an Agent and run a complete browser script with Playwright or Puppeteer.
- Other platforms — getting-started guides for Android, iOS, HarmonyOS, and desktop.
✨ Driven by Multimodal Models
Midscene is all-in on pure vision for UI actions: element localization is based on screenshots only. It runs on multimodal models with strong UI localization, such as Qwen3.x, Doubao-Seed-2.1, GLM-4.6V, gemini-3.5-flash, and UI-TARS, including open-source options you can self-host. For data extraction and page understanding, you can still opt in to include DOM when needed.
Read more about Model Strategy.
📄 Resources
- Documentation: https://midscenejs.com
- Sample projects: midscene-example
- API reference: https://midscenejs.com/reference/#common
🤝 Community
🌟 Awesome Midscene
Community projects that extend Midscene.js capabilities:
- midscene-ios - iOS Mirror automation support for Midscene
- midscene-pc - PC operation device for Windows, macOS, and Linux
- midscene-pc-docker - Docker image with Midscene-PC server pre-installed
- Midscene-Python - Python SDK for Midscene automation
- midscene-java by @Master-Frank - Java SDK for Midscene automation
- midscene-java by @alstafeev - Java SDK for Midscene automation
📝 Credits
We would like to thank the following projects:
- Rsbuild and Rslib for the build tools.
- UI-TARS for the open-source agent model UI-TARS.
- Qwen-VL for the open-source multimodal model Qwen-VL.
- scrcpy and yume-chan for browser-based Android device control.
- appium-adb for its JavaScript bridge to ADB.
- appium-webdriveragent for controlling XCTest from JavaScript.
- YADB for improving text input performance.
- libnut-core for cross-platform native keyboard and mouse control.
- Puppeteer for browser automation and control.
- Playwright for browser automation, control, and testing.
📖 Citation
If you use Midscene.js in your research or project, please cite:
@software{Midscene.js,
author = {Xiao Zhou, Tao Yu, YiBing Lin},
title = {Midscene.js: GUI Agent for E2E Testing.},
year = {2025},
publisher = {GitHub},
url = {https://github.com/web-infra-dev/midscene}
}
✨ Star History
📝 License
Midscene.js is MIT licensed.