Computer & Browser Use
Driving a browser or desktop on the user's behalf. The dividing line is whether the agent reasons over structured elements or raw pixels — the former is far more reliable wherever a usable DOM exists.
12 projects
Extracts the interactive elements of a page and hands the model a structured list, instead of asking it to guess click coordinates from a screenshot. That makes it markedly more reliable for form flows and scraping. It degrades on canvas-heavy applications and on sites with aggressive bot protection, where there is no clean DOM to read.
Type `i` in a project directory and it reads your files, edits them and runs commands inside a sandbox you configure. What sets it apart is `/harness`: it swaps the whole model-facing surface — system prompt, tool schema, message conversion — for the one a given vendor tuned its model on, which is how it aims to get usable work out of inexpensive models like DeepSeek or Kimi. The trade-off is what you are actually adopting: this is a fork of OpenAI's Codex CLI, rewritten in Rust and still on the 0.0.x releases it started in June 2026, sharing little more than a name and a repository with the Python Open Interpreter most people have heard of.
Type an instruction and the app screenshots your screen, feeds it to the UI-TARS vision-language model, and issues the clicks, keystrokes and scrolling to carry it out — across the whole desktop or a browser, which is what separates it from browser-only agents. The model is the catch: you bring an endpoint serving the UI-TARS or Doubao family — self-hosting the open 7B weights via vLLM is the documented local route — and frontier models like GPT or Claude cannot drive the desktop app. Weigh the project's pulse too: desktop binaries have not shipped since August 2025, the hosted remote operators were discontinued, and 2026 activity slowed to a trickle as the repository pivoted toward its Agent TARS stack.
Playwright MCP puts a real Chrome, Firefox, WebKit or Edge behind an MCP server and returns each page as an accessibility snapshot: a text outline of roles, names and reference ids that the agent clicks and types into, with no vision model and no coordinate guessing. Past plain navigation it can mock network requests, read and write cookies and web storage, record a trace or a video, and emit Playwright locators and assertions, which makes it a practical way to turn a session you drove by hand into a test you can check in. The cost is context: Microsoft's own README points coding agents at the Playwright CLI with skills instead, because the tool definitions and the snapshots compete for the same window as your codebase. It earns its keep in long-running loops, where holding one browser open across many turns is worth paying that.
Lets you write deterministic Playwright for the steps you already know and drop into natural language only where the page is unpredictable. That mix is the pragmatic middle ground: fully AI-driven automation is slow and flaky, fully scripted automation breaks on redesigns.
Combines vision with DOM analysis so a flow keeps working when a site is redesigned — the failure mode that kills conventional selector-based scripts. Note the AGPL licence: if you intend to offer it as a network service, read the terms before you build on it.
Computer use runs into the same wall every time — the agent needs a machine, and it cannot be the one you are working on. Cua is four answers to that, which is worth knowing before you install it, because you almost certainly want one of them rather than all four. The drivers let an agent click and type in native applications in the background, without taking the cursor or the focus away from you. The sandbox gives Linux, macOS, Windows or Android behind a single API, locally under QEMU or in the maintainers' cloud. Cua-Bench runs OSWorld and similar evaluations and exports trajectories for training. Lume creates macOS virtual machines on Apple Silicon using Apple's own virtualisation.
You write a character file — a name, a system prompt, a few bio lines and the list of plugins to load — and the runtime turns it into an agent that stores each message, composes context from the providers the turn actually routes to, lets the model pick an action, and files what it learned through evaluators that run after the reply. First-party plugins ship in the same repository for the chat platforms people already use, MCP servers, browser and desktop control, calendar and inbox assistants, non-custodial EVM and Solana wallets, and an on-device model path that answers with the network off. Two things to weigh before adopting it: the repository is a whole product stack — desktop and mobile apps, a hosted cloud, native device bridges — so you take on far more than an agent library, and the stable release on npm is still the older 1.x line while the stack described here ships only under beta and alpha tags.
UI tests break for reasons that have nothing to do with the product: a class name changed, a wrapper was added, a component moved. Midscene removes the selector entirely. Steps are written as sentences, the model locates elements from the screenshot, and assertions can be about what a person would see — the row is highlighted, the layout has not collapsed — rather than about whether a node exists. That also reaches surfaces the DOM cannot: canvas drawings, icon-only buttons, native mobile applications, cross-origin frames. The trade is per-step cost and determinism, because each step is now a model call rather than a lookup, so a suite that runs in seconds with selectors will not run in seconds here.
Agent S separates two jobs that a single model does badly at once: deciding what to do next, and working out where on the screen to click. A general model handles the planning; a smaller vision model, run separately, converts "the Save button" into coordinates. That split is why it scores as it does — Simular reports 72.60% on OSWorld with its best-of-several-attempts setting, against roughly 72% for humans. It is also why setup is more work than a pip install: you supply the grounding model as its own endpoint and tell the agent what resolution it reports coordinates in. By default it drives your actual desktop, one monitor, with no sandbox in between.
It is the common yardstick the computer-use projects cite, and the reason it carries weight is that the environment is genuine: Ubuntu or Windows in a VM, running real applications, reachable through VMware, VirtualBox, Docker with KVM, or AWS, Azure and Modal for parallel runs. That realism is also the cost — adopting it means hypervisors, VM images and cleanup after interrupted runs. Results are version-sensitive too: the Verified update changed tasks and signals, and the project asks that comparisons be made against the re-run numbers.
Almost every desktop-agent benchmark runs on Linux, and almost every desktop that matters commercially runs Windows. This is Microsoft's answer to that gap — a reproducible Windows environment with tasks across the applications people actually use, scored on whether the task was completed. Its useful trick is parallelism: runs can be spread across cloud machines so a few hundred tasks finish in minutes instead of over a day, which is what makes it usable during development rather than only at publication time. Read the last commit date before planning around it, though — this is a research artefact tied to a 2024 paper, and it has moved slowly since.