AppFlow
← Blog

Claude Computer Use: Automate Any Workflow by Letting AI Control Your Computer (2026)

Most automation guides assume you have an API. You often do not. The government procurement portal your team submits vendor registrations through has no API. The 1998-era ERP your finance department uses to close the month has no API. The SaaS competitor pricing portal that blocks curl and headless browsers definitely has no API. That is exactly the gap Claude Computer Use fills.

Claude Computer Use is a beta capability from Anthropic that lets Claude control a computer directly: it takes a screenshot, reasons about what it sees, then returns actions like left_click, type, or key. Your code executes those actions on an actual display, captures another screenshot, and feeds it back. I have spent months building Claude-powered agents across client projects, and I will tell you plainly where the beta still shows: how to implement the agent loop, three workflows that work well today, where it consistently breaks, what it costs per task, and an honest comparison against UiPath and Power Automate.

If you are an operations manager with repetitive work stuck inside systems nobody can touch via API, or a developer asked to automate something that previously required a human with a mouse, this guide is for you. I include specific model IDs, API parameter names, and token cost calculations, not vague estimates.

How Claude Computer Use Works: The Agent Loop

The computer use tool is declared in your API request alongside your messages. The current tool type for Claude Fable 5, Opus 4.8, Sonnet 5, and Sonnet 4.6 is computer_20251124. Older models (Haiku 4.5, Sonnet 4.5) use computer_20250124, which lacks the newer zoom action. The beta header anthropic-beta: computer-use-2025-11-24 is required on every request. Skip it and the tool is simply not found, without a clear error telling you why.

A minimal tool definition looks like this:

{
  "type": "computer_20251124",
  "name": "computer",
  "display_width_px": 1280,
  "display_height_px": 800,
  "enable_zoom": true
}

The agent loop is straightforward. You call the API, Claude responds with a tool_use block (for example, a screenshot action). You execute that action, capture the result, send it back as a tool_result, and repeat. Claude signals it is done by responding with a text block instead of another tool call. The reference implementation at github.com/anthropics/anthropic-quickstarts/tree/main/computer-use-demo wraps this in a sampling_loop() function with a max_iterations guard, defaulting to 10. That guard is not optional. Without it, an agent that hits an unexpected state will retry indefinitely.

For production deployments, Anthropic recommends running inside a Docker container with an Xvfb virtual display (X11). This isolation is not just a best practice for sensitive screens; it also gives you a clean, reproducible environment independent of your host OS. One fixed overhead to know for cost math: every API call carries 466 to 499 system prompt tokens plus 735 tool definition tokens for the computer use tool on Claude 4.x models. This baseline fires before any screenshot pixels or task text are counted.

For a deeper look at how agent loops work and when to use Computer Use as part of a larger agentic architecture, see Claude AI agents: real business impact vs. the hype.

Three Workflows That Actually Work With Computer Use

After building Computer Use across several client projects, these are the three use cases that produce reliable return at current beta reliability levels. I am not listing theoretical applications; I am listing things I have seen work in production.

1. Repetitive Web Form Submission on Locked Portals

Operations teams routinely submit vendor registration forms on procurement portals: eight screens, mandatory fields, conditional logic (VAT fields appear only for EU vendors), and a file attachment step. No API exists. Previously, a person sat at a screen for two hours a day.

Claude navigates each screen, fills the correct fields based on a structured JSON object you pass in the system prompt, handles the conditional branches, and takes a confirmation screenshot as the audit record. It reads field labels rather than XPath selectors, which means it adapts to minor UI changes without re-recording the flow.

What breaks: Cloudflare CAPTCHAs pause the agent. Claude can attempt simple image-select challenges visible in a screenshot, but it cannot reliably pass reCAPTCHA v3, audio challenges, or Arkose Labs FunCaptcha. The working solution is to detect the CAPTCHA in the current screenshot, exit the Claude loop, call a third-party CAPTCHA API (CapSolver or 2Captcha) externally, inject the solution token, and resume. The second failure mode is silent session expiry. Many portals drop session cookies after 30 minutes of inactivity. If a form submission runs longer than that, Claude does not notice the session has reset. You need to check for the login screen in each screenshot and re-authenticate proactively.

2. Legacy ERP Data Extraction

A finance team using a 1998-era ERP with no API used Computer Use to navigate the green-screen interface, extract general ledger data across fiscal periods, and paste it into a modern analytics tool. The task that previously required a dedicated analyst over 3 to 4 months completed in 3 to 4 weeks with Claude supervising the agent loop. The binding constraint: the session ran with a human on standby because the ERP threw modal error dialogs requiring contextual judgment, not just a reflexive click.

This is the strongest Computer Use argument against RPA specifically. A traditional RPA recorder cannot handle a character-mode terminal reliably. Coordinate-based clicking on green-screen is brittle because the layout is dynamic. Claude reads the semantic content of each screen and navigates by what the text says, not by pixel position, so it handles minor screen variations cleanly.

3. ERP Report Generation for Monthly Close

Accounting teams trigger Claude to log into their ERP, navigate to Financial Reporting, select cost centers from a multi-select dropdown, and export to CSV. The reason Computer Use beats RPA here specifically: the dropdown order changes quarterly when new accounts are added. A UiPath bot recorded against a specific dropdown order breaks every quarter and requires a developer to re-record it. Claude selects by reading the account name label, not by position, so quarterly account additions are transparent to it. Finance teams using this pattern for GL reconciliation saw 25 to 35 percent acceleration in monthly close cycles. Customer support data aggregation tasks went from 12 to 15 minutes per ticket down to 90 to 180 seconds execution time.

For a full model comparison with cost-per-query breakdowns across Haiku 4.5, Sonnet 4.5, and Opus 4.8, see the custom Claude assistant cost guide.

Have repetitive tasks on software with no API? Let's see if Computer Use can automate them.

Book a free call →

Claude Computer Use vs RPA: An Honest Comparison

The comparison developers and operations managers actually need before choosing a path:

Criteria Claude Computer Use UiPath (Unattended) Power Automate (Process)
Setup for a new workflow Hours (write a prompt, wire the loop) Days to weeks (record sequences, map selectors) Days (record sequences, configure connectors)
Handles UI changes Good, reads labels semantically Fragile, breaks on selector changes Fragile, breaks on layout changes
Fixed licensing cost None (pay per token used) ~$1,680/robot/year ($140/month) $150/month
Speed per action 1 to 3 seconds (API round-trip) Sub-second (native automation) Sub-second
Green-screen / legacy terminals Works Requires special extension Very limited
Supervision required Medium (error dialogs, CAPTCHAs) Low for stable flows Low for stable flows
High-volume runs (2,000+/day) Expensive at scale Cost-effective Cost-effective with connectors

The pattern I see repeatedly: Computer Use wins during discovery and prototyping, and on workflows where the UI changes often enough to keep a dedicated RPA maintainer busy. UiPath wins when you have a stable, high-volume process and can absorb the Year 1 cost, which typically runs 2 to 3 times the license price once professional services are included. Power Automate wins when you are already inside the Microsoft 365 ecosystem and the workflow maps cleanly to available connectors.

What It Actually Costs Per Task

Here is the token math for a realistic 30-action task on a 1024x768 virtual display, built from Anthropic's documented pricing:

  • Each screenshot: approximately 1,036 visual tokens, computed from the 28x28 pixel patch formula: ceil(1024/28) × ceil(768/28) = 37 × 28 = 1,036 tokens
  • 30 screenshots: approximately 31,080 visual tokens
  • Fixed overhead per API call: 466 to 499 system prompt tokens plus 735 tool definition tokens, so roughly 1,200 tokens × 30 calls = 36,000 tokens
  • Task text and Claude reasoning across the session: approximately 5,000 tokens
  • Output tokens (action JSON, reasoning): approximately 3,000 tokens

At Claude Sonnet 5 intro pricing ($2/MTok input, $10/MTok output, valid through August 31, 2026): approximately $0.19 per 30-action run. At Claude Opus 4.8 ($5/$25 per MTok): approximately $0.49 per run.

At 50 monthly runs: roughly $9.50/month on Sonnet 5, versus $150/month for Power Automate Process or $140/month for a UiPath unattended bot. Computer Use wins clearly at low volume. At 2,000 monthly runs: Sonnet 5 costs roughly $380/month, which crosses both RPA fixed-cost thresholds. At that scale, and assuming the workflow is stable enough for RPA to handle reliably, the RPA economics start to make sense.

Model choice is not a cost decision alone. Claude Sonnet 4.6 is documented to be more mechanically precise at clicking for standard structured tasks. Claude Opus 4.8 handles complex multi-step reasoning and scores 83.4% on the OSWorld-Verified benchmark (July 2026). But on OSWorld 2.0, which tests long-horizon, multi-application workflows, Opus 4.8 completes only 20.6% of tasks successfully. That gap is critical for production planning: single-application, bounded tasks work well. Chaining across five different applications in one session is still unreliable at production quality.

For long sessions (30 minutes or more), Claude Managed Agents adds $0.08 per session-hour on top of token costs. A 2-hour monthly close session adds $0.16 per run in runtime overhead. Prompt caching reduces the per-call overhead substantially for stable system prompts: cache reads cost 10% of the base input price, so caching the 735-token tool definition and 499-token system prompt across a 30-turn loop saves meaningfully at volume.

For a complete breakdown of Claude pricing across models with worked examples at different query volumes, the custom Claude assistant guide covers Haiku, Sonnet, and Opus 4.8 economics side by side.

The Coordinate Drift Trap: What Actually Goes Wrong

This is the most common silent failure I have seen in Computer Use deployments, and it is invisible until you watch an agent confidently click the wrong element for twenty minutes straight without any error being thrown.

macOS Retina displays capture screenshots at a device pixel ratio of 2. A 1440x900 logical screen produces a 2880x1800 pixel image. When a developer passes this full-resolution image to the API with display_width_px: 2880 and display_height_px: 1800 set to match, Claude processes the image in that coordinate space. The problem is in how those coordinates are consumed: Claude correctly identifies the Submit button and calls left_click with coordinate [1200, 850], but the actual system click lands at [600, 425]. The click is off by a factor of 2 in both dimensions, placing it on a completely different element.

The agent takes a new screenshot, sees the wrong state, attempts a corrective action, and spirals toward the max_iterations ceiling, making the situation worse with each step. The root cause: display_width_px and display_height_px must exactly match the pixel dimensions of the image you actually send to the API. Any mismatch silently misdirects every single click in the session. The fix is explicitly documented in the Anthropic platform docs under the Retina display section. Two options:

  1. Capture at native resolution (2880x1800), downscale the image to 1440x900 before sending, and set display_width_px: 1440 and display_height_px: 900. Claude returns logical coordinates directly, matching what your click function expects.
  2. Capture at native resolution, send at full resolution with display_width_px: 2880 and display_height_px: 1800, then divide every coordinate Claude returns by the scale factor (2) before executing the system click.

The reference implementation on GitHub handles this via a scale_factor variable computed at initialization. Option 1 is simpler to reason about when building from scratch. Either way, the rule holds: the declared display dimensions and the image dimensions sent to the API must match exactly.

Beyond coordinate drift, seven recurring failure patterns appear across Computer Use implementations. Every production build needs an explicit plan for each: misclicked elements (add bounding box checks and retry with a more specific description), infinite retry loops (always set max_iterations, never skip it), hallucinated UI components (Claude occasionally describes elements it cannot actually see in the screenshot), silent session drift (session expires mid-loop without Claude detecting the state change), unguarded destructive actions (add confirmation gates before any action that submits, deletes, or logs out), credential exposure in screenshots (blur or mask credential fields before passing screenshots to the API), and prompt injection from on-screen text (Anthropic's classifiers scan screenshots for injected instructions by default; this can be disabled for headless deployments by contacting support).

Frequently Asked Questions

Is Claude Computer Use ready for production use in 2026?

As of July 2026, it remains in beta. The beta header anthropic-beta: computer-use-2025-11-24 is required on every request, and Anthropic has not announced a general availability date. It is production-viable for bounded, single-application tasks with a human on standby for error dialogs and CAPTCHAs. For long-horizon, multi-application workflows, Claude Opus 4.8 completes only 20.6% of tasks on the OSWorld 2.0 benchmark, so plan for significant failure-handling engineering if you are building multi-step pipelines that cross multiple systems.

How does Claude Computer Use handle CAPTCHAs?

It cannot reliably pass reCAPTCHA v3, audio challenges, or Arkose Labs FunCaptcha. Claude can attempt simple image-select challenges visible in a screenshot. For production, the working pattern is: detect the CAPTCHA in the screenshot via a visual check in your agent loop, pause Claude, call a third-party CAPTCHA service (CapSolver or 2Captcha) externally, inject the solution token into the browser session, and resume the Claude loop. This adds latency and a small per-solve cost, but it works consistently for standard image CAPTCHAs.

What does Claude Computer Use cost compared to UiPath or Power Automate?

A 30-action task costs approximately $0.19 at Claude Sonnet 5 intro pricing ($2/MTok input through August 31, 2026) or $0.49 at Opus 4.8 ($5/MTok). UiPath unattended bots cost approximately $1,680 per robot per year, with total Year 1 cost typically 2 to 3 times the license after professional services. Power Automate Process is $150 per month. Computer Use is cheaper for low-to-medium volumes. At 2,000 or more monthly runs the per-token cost scales past the fixed RPA license, making RPA more economical at high, stable volumes.

Which Claude model should I use for Computer Use tasks?

For most structured, single-application tasks, Claude Sonnet 5 ($2/MTok input) gives the best balance of cost and accuracy. Claude Sonnet 4.6 is documented to be more mechanically precise at clicking for standard UI interactions. Claude Opus 4.8 ($5/MTok) handles complex multi-step reasoning and reaches 83.4% on OSWorld-Verified benchmarks, but is overkill for straightforward form-filling. Claude Haiku 4.5 supports only the older computer_20250124 tool type with a 200k context limit, which makes it unsuitable for complex, long-running agent loops.

Can Claude Computer Use automate a legacy Windows application or green-screen ERP?

Yes, and this is one of its clearest advantages over RPA. QA teams use Claude to click through legacy Windows applications over VNC, verifying labels, error messages, and navigation flows by reading screen content semantically, without needing maintained XPath selectors. For green-screen ERPs, coordinate-based RPA automation is extremely brittle because the character-mode layout is dynamic. Claude reads the text and navigates by meaning, making it resilient to minor content changes that would break a recorded bot.

How do I prevent Claude from taking unintended destructive actions?

Three layers work together: first, state explicitly in the system prompt what Claude is not permitted to do, including submit final records without confirmation, delete, or log out. Second, add a checkpoint in your agent loop code that inspects each action before executing it; flag any action whose text includes "delete", "confirm final", "submit", or "logout" and route it to a human approval queue. Third, run inside a Docker container with a sandboxed Xvfb virtual display as Anthropic recommends, so that even an unintended action is isolated from your production systems.

What is coordinate drift and how do I fix it?

Coordinate drift is a mismatch between the pixel dimensions of the screenshot you send to the API and the coordinate space your click function uses. On macOS Retina displays, screenshots capture at 2x pixel density, producing a 2880x1800 image from a 1440x900 logical screen. Without correcting for this, every click Claude returns lands at the wrong position by a factor of 2. The fix: either downscale the screenshot to your logical display resolution before sending it and set display_width_px and display_height_px to match, or send the full-resolution image and divide every coordinate Claude returns by the scale factor before issuing the system click. The display_width_px and display_height_px parameters must always match the actual pixel dimensions of the image you send.

Computer Use works anywhere legacy software blocks automation. If your business is in retail and commerce, construction, or professional services, see the sector pages for automation scenarios specific to your context.

Aurélien Migeot is a freelance AI developer based in France. He has spent 8 years building Claude-powered tools and agents for small businesses. Founder of AppFlow Solutions. Book a free discovery call

Let's discuss your project. Free 30-minute discovery call.

Book a call
Claude Computer UseAnthropic Computer Use 2026browser automation AIworkflow automationRPA vs AI automationlegacy software automationcomputer use tutorial