What Is Computer Use in AI? How Agents Drive a Screen
Computer use lets an AI model drive an ordinary screen with clicks and keystrokes. Here is the loop, the cost, and the safety setup the vendor documentation recommends.
Computer use is a capability that lets an AI model operate software the way a person does: it looks at a screenshot, decides what to do, and sends a click, a keystroke or a scroll. Nothing in the target program has to cooperate. There is no API to integrate, because the model works through the same screen a human would. That is what makes it useful for old or closed software, and what makes it risky.
How computer use works
The mechanism is a loop. Anthropic's computer use documentation lays it out in five steps:
You give the model a task and the computer toolset.
The model replies with tool calls, such as take a screenshot, click at these coordinates, or type this text.
Your application carries out each action in an environment you control.
You send the results back, usually a fresh screenshot.
The model looks at the result and asks for more actions until the task is done.
The docs call the repetition of steps 3 and 4 without user input the agent loop. The model never touches the machine itself. Your code is the hands, and the model is the eyes and the decision-maker.
The toolset covers the basics of using a computer: taking and zooming into screenshots, left, right and double clicks, dragging, moving the mouse, typing, pressing keys, scrolling and waiting. The docs add that coordinates are in screenshot pixel space, so if you shrink screenshots you must scale the coordinates back up before acting.
Computer use versus an API call
Approach | How the model acts | Best for | Main cost |
|---|---|---|---|
API or tool call | Sends structured data to a defined function | Anything with a real API | Needs an integration to exist |
Browser-only tool | Drives web pages | Tasks that stay inside a browser | Limited to the web |
Computer use | Looks at the screen, clicks and types | Software with no API, desktop apps, mixed workflows | Slower, more tokens, more ways to fail |
The rule that follows: use an API when one exists, because it is faster and more predictable. Reach for computer use when the alternative is a person sitting at that screen. Our explainer on tool calling in AI agents covers the structured route.
What it costs
Each screenshot the model looks at is an image it has to read. The docs put it at roughly 1,000 to 1,800 tokens per screenshot and warn that long loops accumulate tokens quickly, suggesting you prune or resize history. A task that takes forty steps means forty images in play. That is why computer use tends to be slower and pricier than a direct integration for the same job.
Safety, straight from the docs
Because the model can click anything on the screen, the documentation is direct about precautions. Its recommendations:
Use a dedicated virtual machine or container with minimal privileges, so a mistake cannot reach your real system.
Keep sensitive data away from the model, especially account logins.
Limit internet access to an allowlist of domains.
Ask a human to confirm decisions with real-world consequences, such as accepting cookies, completing a financial transaction or agreeing to terms.
The docs also warn that the model can be hit by prompt injection from webpage content or images, meaning text on a page that tries to give the agent new orders. Anthropic describes classifiers that scan what the tools return and steer the model to check whether an instruction really came from you. It notes this protection may not suit every use case, for example setups with no human in the loop.
The practical takeaway is to treat the agent's machine as disposable. Our guide to how to sandbox an AI agent walks through that setup, and why AI agents need their own browser explains why sharing your everyday browser profile is the wrong default.
Where you will meet it
The idea is not tied to one vendor. At DevDay, OpenAI said its Agents API supports computer use, with agents navigating a hosted browser, as Runtime Wire's roundup lists. Expect the capability to keep spreading, and expect the safety advice above to apply wherever it appears. For the broader model background, start with how AI models work.
FAQ
Is computer use the same as browser automation?
Not quite. Browser automation usually scripts fixed steps against a page. Computer use lets a model decide each step from what it sees, and it can work in desktop software, not only a browser.
Does computer use need an API?
No, and that is the point. The model works through the screen, so the target software does not have to expose anything. When a proper API exists, it is usually the better choice.
Is computer use safe to run on my own laptop?
The vendor documentation recommends a dedicated virtual machine or container with minimal privileges instead, and keeping sensitive data such as logins away from the model.
Why is computer use slower than an API call?
Each step involves a screenshot the model must read and a fresh decision, and long tasks repeat that many times. A direct API call skips the looking and the clicking.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


