Trust is earned, not given

A different perspective

2025-05-15 · Projects

AI Frontiers, part 15: Computer use and GUI agents

Part 15from the AI Frontiers series · 65 parts in all

On 22 October 2024 Anthropic released a beta capability it called computer use: a model that takes a screenshot, moves a virtual cursor, clicks, types, and takes another screenshot to see what happened. OpenAI followed in January 2025 with Operator, a system that drives a browser the same way (OpenAI, "Introducing Operator"). Google announced Project Mariner in December 2024 with similar ambitions (Google DeepMind).

The demos are compelling and slightly unnerving, and the temptation is to read them as the arrival of a general-purpose digital employee. The measurements say something much narrower and, I think, much more useful: GUI agents work, they are slow, they are expensive, and their failure rate is dominated by how many steps a task requires rather than how hard it is. This entry is about the architecture behind them, the reason the benchmarks are still so low, and where I think this actually gets deployed first.

Why the screen is the tempting fallback

Start with the argument for it, because the argument is strong. Part 12 made the case that the world should be exposed to models as typed, permissioned tool calls. That case assumes the software in question has an API. A great deal of it does not, and some of it never will: legacy desktop applications, internal line-of-business systems written before anyone thought about integration, thick clients that are really a database with a window, mainframe terminals, and the long tail of SaaS products whose API covers the twenty percent of functionality their customers asked for and none of the rest.

For all of that, the graphical interface is the only interface that exists, and it is a remarkably consistent one. Buttons, text fields, menus, tables. A human can be trained on one application and transfer intuitions to the next, because the conventions are shared. That is exactly the property that makes GUI control a plausible model capability: if the screen is a universal API, then one agent architecture works everywhere — no integration, no vendor cooperation, no API keys. The cost of that universality is that the agent must work at the speed and precision of a person using a mouse, and that any small change to a layout breaks it.

The loop: perceive, ground, act, repeat

Every one of these systems is a control loop over screenshots.

Perception is a screenshot, resized to whatever resolution the model consumes. This is where the token arithmetic from part 10 bites hardest: a 1920×1080 desktop screenshot is a lot of image tokens, and the agent needs a new one after every action. A twenty-step task means twenty screenshots in context, which is why context management rather than model intelligence is often the binding constraint on long sessions.

Grounding is the hard subproblem. Knowing that you should click the Save button is one capability; knowing that the Save button is at pixel (1447, 892) is another. Research on GUI grounding — SeeClick (Cheng et al.) and successors — is specifically about this translation, and it is where accuracy falls off. A range of tricks help: prompting the model with set-of-mark overlays that label candidate elements before it chooses (Yang et al.), or providing the accessibility tree so the model receives semantic element names instead of pixels. Anthropic's implementation and most browser agents use some combination: screenshot for layout, structured tree for labels, coordinates for the action.

Action is a small vocabulary — click at a coordinate, type text, press a key, scroll. Occasionally a screenshot becomes an explicit scroll-and-re-screenshot because the target is below the fold. The interesting design work is in what the model is not allowed to do: no arbitrary code, no clipboard access to credentials, no unattended navigation outside an allowlist.

The benchmarks are humbling on purpose

OSWorld (Xie et al.) asks an agent to complete multi-step tasks in a real Ubuntu desktop across 369 tasks — editing documents, managing files, configuring applications. Human performance is around 72%. When Anthropic released computer use, Claude 3.5 Sonnet scored 14.9%, roughly tripling the previous state of the art and still missing five out of six tasks. OpenAI's Operator, benchmarked in early 2025, was in a similar band on desktop-style tasks, with better numbers on the narrower web subset.

The web benchmarks tell a more nuanced story. WebArena (Zhou et al.) stood up self-hosted clones of common web applications to test agents end to end; VisualWebArena (Koh et al.) added visually grounded tasks; WebVoyager (He et al.) ran against live websites; Mind2Web (Deng et al.) tested generalization to unseen sites; AndroidWorld (Rawles et al.) did the same for a mobile operating system. Performance is best where the task is single-site and the markup is clean, and degrades sharply when the task spans multiple applications or requires remembering state across pages.

Reliability compounds, and that is the whole story

Here is the arithmetic that explains the gap between demo and deployment. Suppose a per-step success rate of 95%, which would be excellent. A ten-step task succeeds 60% of the time. A twenty-step task succeeds 36% of the time. A thirty-step task, which is not an unusual length for real work like "reconcile this invoice against the purchase order and file the discrepancy," succeeds 21% of the time. And each step costs a screenshot, a model call and one to three seconds of wall-clock latency, so the thirty-step task takes a minute or more even when it works.

Worse, the failure mode is not a clean stop. An agent that misreads a table can click through four more screens of a workflow before anything looks wrong, and then it has to recover — which is a genuinely harder capability than forward progress, because it requires noticing that what happened does not match what was expected. Every credible design therefore puts deterministic checkpoints in the loop: read the value back, verify the confirmation number, assert the record exists. That is not AI; it is the same idempotency-and-verification discipline you would apply to a payment integration, and it is what turns a 60% agent into a usable one.

This is also the counterweight to the "just throw inference-time compute at it" instinct from part 11. More thinking per step raises per-step accuracy, which is worth a lot when errors compound. But you cannot buy your way out of a fundamentally unstable control loop with a bigger budget for each iteration; you fix it with checkpoints, retries and a narrower task.

The worst-case tool surface

I want to be blunt about the security posture of GUI agents, because it is qualitatively different from a tool-calling agent. A tool-calling agent can be tricked into calling a tool; a computer-use agent can be tricked into calling a tool, then copying the result into a web form, and then clicking submit. Indirect prompt injection (Greshake et al.) is a published, well-understood attack: instructions embedded in content the model reads. On a desktop, "the content the model reads" is everything on the screen, and "the action it takes" is every action a user can take. AgentDojo (Debenedetti et al.) demonstrates attacks and defenses in exactly this tool-integrated setting, and the general finding is that defenses reduce rather than eliminate the problem.

The mitigations are structural, not clever. Run the agent in a virtual machine with no credentials it does not need. Keep a human in the loop for irreversible actions — payments, deletions, sending mail to an external address. Restrict navigation to an allowlist of domains or applications. Log every action with the screenshot that motivated it, so that a post-incident review can reconstruct what the model saw. Anthropic's own guidance for computer use reads like this, and it reads that way because there is no version of this technology where the model decides its own permissions wisely.

The economics per task

It is worth putting a number on the unit economics, because they determine the adoption order far more than capability does. A thirty-step task means roughly thirty screenshots and thirty model calls, at one to three seconds each, with a high-resolution image in context every time. Call it a minute and a half of wall-clock time and, depending on the model and the image encoding, somewhere between ten cents and a dollar of API spend. Add a supervisor who reads the outcome, and the honest comparison is not "agent versus free" — it is "agent plus a few minutes of human verification versus a human doing the whole thing."

That arithmetic produces three useful rules of thumb. Ascending value of a task makes generic agents more attractive; high volume makes a deterministic script more attractive. Short tasks of two or three steps, where the compounding math does not bite, are where agents win today — and where reliability is high enough that the human check can be a sample rather than a review of every run. Long multi-application workflows are where the failure rate makes the cost of supervision exceed the cost of the work.

Durability matters as much as cost. A pixel-coordinate agent breaks when the layout changes; an accessibility-tree agent breaks when the labels change but survives a re-theme; a script driving a documented API breaks only when the API changes, and tells you loudly when it does. The ordering of fragility is the ordering of preference, which is why the sensible long-term architecture is a hierarchy: structured API where one exists, accessibility or DOM signals where the semantics are available, and raw pixels only when the screen is all there is. The agent's intelligence should be spent on judgment, not on finding a button.

Where it lands

My prediction is unromantic and I think correct: computer use gets adopted first as a better RPA — robotic process automation, the industry that has been scripting GUI clicks since the 1990s. The pitch against traditional RPA is strong. Classic screen-scraping automation breaks whenever a button moves, and maintenance is the entire cost of ownership. A model-driven agent that reads the screen degrades more gracefully, and when it fails it fails in a way a human can inspect. The pitch against a human is weaker: the agent is slower, and for a short repetitive task the person is simply faster.

So the adoption pattern will look like the adoption pattern of every automation wave before it: high-volume, repetitive, verifiable, previously manual workflows where the alternative is a person doing the same clicks forty times a day. Claims processing. Order entry. QA passes on configurations. Screenshot-driven regression testing, where the agent does not have to be right, only useful at finding things that changed. Accessibility, where "operate this application for me" is a well-defined and valuable outcome regardless of speed.

There is also a version of this that is not about replacing anyone. For a person who cannot use a mouse reliably, or who can see a screen but cannot type into it, "describe the goal and let the agent operate it" is not a cost optimization — it is the difference between being able to use software and not. That framing is worth keeping in view when the adoption debates turn into arguments about labor, because it is the one case where the agent being slower than a human costs nothing at all.

And the long run, if per-step reliability keeps climbing, is a slow inversion: as more software exposes structured tools, the GUI becomes the fallback for the shrinking set of systems that cannot. The interesting thing about computer use is not that it replaces APIs. It is that for the next several years, it is the only thing that works on everything that does not have one — and in enterprise software, that is a great deal of everything.

Works Cited

Anthropic. "Computer Use (Beta)." Anthropic Documentation, docs.anthropic.com/en/docs/agents-and-tools/computer-use. Accessed 15 May 2025.

Cheng, Kanzhi, et al. "SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents." arXiv, 2024, arxiv.org/abs/2401.10935. Accessed 15 May 2025.

Debenedetti, Edoardo, et al. "AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents." arXiv, 2024, arxiv.org/abs/2406.13352. Accessed 15 May 2025.

Deng, Xiang, et al. "Mind2Web: Towards a Generalist Agent for the Web." arXiv, 2023, arxiv.org/abs/2306.06070. Accessed 15 May 2025.

Google DeepMind. "Project Mariner." Google DeepMind, Dec. 2024, deepmind.google/discover/blog/project-mariner/. Accessed 15 May 2025.

Greshake, Kai, et al. "Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection." arXiv, 2023, arxiv.org/abs/2302.12173. Accessed 15 May 2025.

He, Hongliang, et al. "WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models." arXiv, 2024, arxiv.org/abs/2401.13919. Accessed 15 May 2025.

Koh, Jing Yu, et al. "VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks." arXiv, 2024, arxiv.org/abs/2401.13649. Accessed 15 May 2025.

OpenAI. "Introducing Operator." OpenAI, 23 Jan. 2025, openai.com/index/introducing-operator/. Accessed 15 May 2025.

Rawles, Christopher, et al. "AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents." arXiv, 2024, arxiv.org/abs/2405.14573. Accessed 15 May 2025.

Xie, Tianbao, et al. "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments." arXiv, 2024, arxiv.org/abs/2404.07972. Accessed 15 May 2025.

Yang, Jianwei, et al. "Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V." arXiv, 2023, arxiv.org/abs/2310.11441. Accessed 15 May 2025.

Zhang, Chaoyun, et al. "Large Language Model-Brained GUI Agents: A Survey." arXiv, 2024, arxiv.org/abs/2411.18279. Accessed 15 May 2025.

Zhou, Shuyan, et al. "WebArena: A Realistic Web Environment for Building Autonomous Agents." arXiv, 2023, arxiv.org/abs/2307.13854. Accessed 15 May 2025.