Computer use

Screenshot in, click out. The fallback for software that has no API, and the most brittle tool an agent has.

Part of the Agents track on lAItest.

Some software has no API and never will. The only interface left is the one built for hands and eyes.

The same loop, with pixels

The observation is a screenshot. The actions are move, click, type, scroll, and take another screenshot. The model has to locate a target inside an image, emit coordinates, then read the next screenshot to work out whether the click did what it expected. Same loop as any other agent; the tools are just unusually low-level.

A common misconception

Commonly believed: If a model can drive a browser, it can do anything a person can do on a computer.

Actually: It is the slowest and most brittle way to give an agent a capability. Every step costs a screenshot, layouts move, a modal appears and the plan is already stale, and there is no undo. Where an API exists, the API wins on every dimension. Computer use is the fallback, not the goal.

What it changes about risk

A screen agent inherits everything the logged-in session can do. It can click Send, Confirm and Delete, and it cannot reliably tell a real system dialog from one drawn by a hostile page. The design question is blast radius: which account, which machine, which permissions, and what is reachable from that desktop.

In one sentence

Computer use is the fallback for software with no API, and it should be scoped like handing a stranger your unlocked laptop.