We use cookies to enhance user experience, personalize content, and analyze traffic. Cookie Policy

← Back to all articles

Computer Use Agents: How They Browse the Web and Why IPs Matter

How a computer use agent browses with screenshots and clicks, what websites see from its sandbox, and how to route its browser through a proxy safely.

by Unknown Proxies

14 min read

October 2, 2026

Computer Use Agents: How They Browse the Web and Why IPs Matter

A computer use agent is an AI model that operates a computer the way a person does. It looks at a screenshot, decides where to click or what to type, and sends mouse and keyboard actions to a desktop. A small program on your side carries out each action and returns a fresh screenshot. The loop repeats until the task is done.

When a computer use agent browses the web, it isn't calling a special API. It is clicking around a normal browser running on a virtual desktop, usually a container or VM you provide. Websites see that browser's traffic, coming from whatever network the sandbox sits on. In practice that usually means a cloud server's IP address. That is why IPs matter: the site never sees the model, only the network and browser you gave it.

This guide explains how computer use agents browse, what a website can observe about them, and how to put a proxy in the right place in the sandbox. Done right, the browser gets a deliberate network identity without leaking credentials to the model or sending your screenshots through a metered proxy.

What Is a Computer Use Agent?

The model APIs that power computer use agents work the same way at the core. Anthropic's computer use tool, OpenAI's computer use tool in the Responses API, and Google's Gemini computer use all take screenshots in and send UI actions out. In all three, you supply the environment: the display, the browser, and the code that turns "click at (412, 230)" into a real input event.

That separates computer use agents from the other ways an AI touches the web:

Computer use agent DOM-based browser agent Fetch or search tool
What the model sees Screenshots DOM or accessibility tree, sometimes plus screenshots Page text or HTML
What it sends back Clicks and keystrokes at pixel coordinates Actions on element references A URL to fetch
What it can control Anything on the screen: browser, PDFs, desktop apps One browser Nothing interactive
Typical harness Virtual display (Xvfb) driven by xdotool, or a browser driven by Playwright Playwright or the Chrome DevTools Protocol An HTTP client
Where the proxy goes The sandbox's browser or network The browser launch or context options The HTTP client

Pixel-level control is slower and costs more per step than reading the DOM, but it works on canvas apps, embedded PDFs, and native desktop software that DOM agents can't reach. Developers often combine the two by giving a single agent both a computer tool and DOM or shell tools.

Consumer products such as the cloud browser in ChatGPT Work run the same kind of loop on the vendor's own cloud machines. You can't change their network, so this guide focuses on agents you run yourself.

How Computer Use Agents Browse the Web

Anthropic's reference implementation is a good concrete example because most self-hosted setups look like it. It is a Docker image with Ubuntu, an Xvfb virtual display, the Mutter window manager, Firefox ESR, and xdotool. The agent loop is a Python app that runs inside the same container. When the model asks for a click, the loop runs xdotool mousemove and click against the virtual display. When the model asks for a screenshot, the loop captures the display and sends the image back to the API.

Computer use agent sandbox with two separate network paths, one from the agent loop to the model API carrying screenshots and actions, and one from the sandbox browser through a proxy to websites

That gives a computer use agent two network paths that are easy to mix up:

  1. The model path. The agent loop sends screenshots to the model API and receives actions back. Websites never see this traffic.
  2. The browsing path. The browser on the virtual desktop loads pages, scripts, images, and API calls from the sites the agent visits. This is the only traffic a website sees.

The proxy belongs on the browsing path only. If you send the model path through a residential proxy, every screenshot upload counts against your proxy bandwidth and adds latency to each step. Nothing improves on the website side.

Two other properties shape how these agents look from the outside:

What a Website Sees From a Computer Use Agent

A site has three kinds of evidence: the network connection, the browser, and the behavior. Computer use agents look unusual in a different mix from other agents.

The network. The connection comes from the sandbox host. On AWS, Google Cloud, Azure, or a hosted sandbox provider, that is a hosting ASN. Many bot management rules score hosting ranges as high-risk because so little human browsing comes from them. This is the signal that most often decides whether the agent gets a page or a 403 Forbidden on its first request.

The browser. This depends on the harness. In a desktop-style setup like Anthropic's demo, Firefox is a normal, headed browser with no WebDriver or DevTools session attached, because input arrives as X11 events from xdotool. We loaded a test page in the reference image's Firefox ESR 140, and navigator.webdriver reported false. Harnesses that drive the browser through Playwright or a DevTools connection, which is how many browser-only computer use examples work, usually report true and carry the usual automation tells.

Even the desktop-style browser has quirks. It runs at a small virtual screen size (Anthropic's demo defaults to 1024×768), and renders without a GPU. Its timezone and language default to whatever the container has, usually UTC and US English.

The behavior. The pointer jumps straight to its target instead of moving across the screen, because xdotool mousemove teleports it. Actions come in bursts separated by multi-second pauses while the model thinks. When something fails, an agent may click the same button again right away.

Put together, a desktop-style computer use agent often has a cleaner browser fingerprint than a headless scraper and a worse network identity than a home user. That makes the IP address the signal you most directly control.

Why IPs Matter for Computer Use Agents

Cloud IP reputation. An agent that works from your laptop and fails as soon as it runs in the cloud is usually hitting an IP problem. The sandbox moved from a residential connection to a hosting range. For the background on that gap, see datacenter vs residential proxies.

Geo and locale. A computer use agent reads the rendered page and reports what it sees. If the exit IP is in Virginia but the task is about German prices, the agent will confidently report the wrong currency, stock, or search results. The proxy country, the browser's Accept-Language, and the timezone should all agree.

Session length. Because every step takes seconds, computer use sessions are long. A task with a login, a search, and a few pages of results can take 10 minutes or more. If the exit IP changes partway through, the site sees an authenticated session move between networks. That often means a logout or a new challenge. Use one stable exit per task, not per-request rotation. The sticky vs rotating proxies guide covers session design. Unknown Proxies sticky residential sessions can hold an IP for up to 2 hours, which covers most agent tasks.

Parallel sandboxes. Ten containers on one VM share one public IP unless you give each its own exit. To the target, that looks like one very busy visitor. Plan for one sandbox, one exit IP.

How to Route a Computer Use Agent Through a Proxy

There are three places to set a proxy in a computer use sandbox. They differ in what they cover and what the agent can undo.

Comparison of three places to set a proxy for a computer use agent sandbox, browser policy, container environment variables, and an isolated network with an egress gateway

Placement Covers Agent can bypass it? Use it when
Browser policy (locked) The sandbox browser only Not from the browser UI, but yes through a shell tool The agent only uses the browser
HTTP_PROXY / HTTPS_PROXY in the container CLI tools and HTTP libraries that read them Yes, it can unset them You also need curl or scripts in the sandbox proxied
Isolated network with an egress gateway Everything in the sandbox No The agent has shell access or you need fail-closed routing

For the common case, the browser policy plus a credential-holding forwarder is the right default. The walkthrough below uses Anthropic's reference image, but the same approach works for any Firefox-based sandbox.

Step 1: Run a local forwarder that holds the credentials

Don't put proxy credentials anywhere the agent can see them. A password typed into a Firefox prompt ends up in a screenshot. An environment variable in the sandbox can be printed by a shell tool. Instead, run a small forwarder in its own container. It authenticates to the upstream proxy and exposes an unauthenticated listener on a private Docker network.

This example uses GOST, a single-binary proxy tunnel:

docker network create cua-net

docker run -d --name egress --network cua-net gogost/gost:3 \
  -L "http://:3128" \
  -F "http://${PROXY_USER}:${PROXY_PASS}@${PROXY_HOST}:${PROXY_PORT}"

The listener on egress:3128 needs no password, so don't publish that port to the host or the internet. Only containers on cua-net can reach it. If your proxy plan supports source-IP allowlisting, you can allowlist the sandbox host's IP and skip the forwarder.

Step 2: Point the sandbox browser at it with a locked policy

Firefox reads enterprise policies from /etc/firefox/policies/policies.json on Linux. Save this as policies.json on the host:

{
  "policies": {
    "Proxy": {
      "Mode": "manual",
      "Locked": true,
      "HTTPProxy": "egress:3128",
      "UseHTTPProxyForAllProtocols": true,
      "Passthrough": "localhost, 127.0.0.1"
    },
    "Preferences": {
      "intl.accept_languages": { "Value": "de-DE, de, en", "Status": "locked" }
    },
    "DisableTelemetry": true,
    "DisableAppUpdate": true
  }
}

Locked: true greys out the connection settings, so the agent can't switch the proxy off from Firefox's settings page. That matters because a computer use agent has full access to the desktop UI, and a prompt-injected page could tell it to do exactly that. The Preferences block sets Accept-Language to match a German exit. DisableTelemetry cuts Firefox's telemetry pings so they don't use your proxy bandwidth. DisableAppUpdate is harmless here (the packaged ESR build doesn't self-update) and keeps the policy portable to Mozilla's own builds.

Step 3: Start the sandbox on the private network

docker run -d --name cua --network cua-net \
  -e ANTHROPIC_API_KEY="$ANTHROPIC_API_KEY" \
  -e TZ=Europe/Berlin \
  -v "$PWD/policies.json:/etc/firefox/policies/policies.json:ro" \
  -v "$HOME/.anthropic:/home/computeruse/.anthropic" \
  -p 127.0.0.1:5900:5900 -p 127.0.0.1:8501:8501 \
  -p 127.0.0.1:6080:6080 -p 127.0.0.1:8080:8080 \
  ghcr.io/anthropics/anthropic-quickstarts:computer-use-demo-latest

Bind the published ports to 127.0.0.1. The demo's VNC server has no password, so an open port gives anyone who finds it the agent's desktop and its chat UI, which spends your API key and browses through your proxy. On a remote host, reach the demo through an SSH tunnel such as ssh -L 8080:localhost:8080 -L 6080:localhost:6080 -L 8501:localhost:8501 user@host, never an open port.

Mount the policy read-only. The demo user has passwordless sudo, so a writable file could be edited by the agent's shell tool. TZ sets the timezone Firefox reports to pages. Change it, the Accept-Language value, and the proxy country together.

We tested the forwarder and this policy with the reference image's Firefox ESR 140. Page loads went through the forwarder to an authenticated upstream proxy. A test page reported Europe/Berlin as the timezone, de-DE,de,en as navigator.languages, and navigator.webdriver as false.

The model path stays direct. The agent loop's API calls leave the container normally and never touch the proxy, because the proxy is set only in Firefox. If you add HTTPS_PROXY to the container for command-line tools, also set NO_PROXY to the model API host, such as api.anthropic.com. Otherwise most Python HTTP clients will send API calls and screenshots through the proxy too.

Step 4: Verify the exit before the agent starts

Check from the forwarder's side first. The curl image's entrypoint passes any arguments that start with - straight to curl:

docker run --rm --network cua-net curlimages/curl -s -x http://egress:3128 https://api.ipify.org

Then open the demo's combined interface at http://localhost:8080 (or the desktop-only view at http://localhost:6080/vnc.html), load https://api.ipify.org in the sandbox Firefox, and confirm it shows the same address. Visiting about:policies shows whether Firefox picked up the policy file. If the forwarder prints an IP but Firefox shows the host's IP, the policy didn't load. If the forwarder check fails, test the upstream proxy directly with curl -x "http://$PROXY_USER:$PROXY_PASS@$PROXY_HOST:$PROXY_PORT" https://api.ipify.org. A 407 Proxy Authentication Required there means the credentials are wrong.

When the agent has a shell, isolate the network

A browser policy only controls the browser. Anthropic's reference agent also has a bash tool, and anything the agent runs there, such as curl or pip, uses the container's direct route. If the sandbox must never reach the internet except through the proxy, take away the direct route.

Attach the desktop container only to a Docker network created with docker network create --internal, and connect the forwarder to both that network and a normal one. Run the agent loop in its own container with outside access so it can still reach the model API. This takes more work than the reference demo, because the loop then has to control the desktop over the network. The payoff is that any connection that skips the proxy fails instead of leaking your host IP.

Anthropic's own guidance points the same way. It recommends a dedicated VM or container with minimal privileges and limiting internet access to an allowlist of domains. GOST can enforce that allowlist at the forwarder with a whitelist bypass rule, marked by the ~ prefix. Put it on the listener, not the upstream hop:

docker run -d --name egress --network cua-net gogost/gost:3 \
  -L "http://:3128?bypass=~example.com,.example.com" \
  -F "http://${PROXY_USER}:${PROXY_PASS}@${PROXY_HOST}:${PROXY_PORT}"

On the listener, a request for a host outside the list is refused, and the browser gets a 403 on CONNECT. On the -F hop, the same rule means "skip this hop," so non-matching traffic leaves directly from the forwarder's own IP. We tested both. With the rule on the hop, an off-list request succeeded without ever reaching the upstream proxy.

Choosing the Exit IP for a Computer Use Agent

Pick the proxy type by the shape of the task. The ISP vs residential proxies guide covers the tradeoffs in depth.

Budget for bandwidth. A computer use agent loads full pages with images, fonts, and scripts, and it often revisits pages while reasoning. Measure one real task through the forwarder, then size a plan with the data usage calculator.

What a Better IP Won't Fix

A clean exit IP removes the hosting-range penalty. It doesn't change the rest of the picture:

FAQ

What is a computer use agent?

A computer use agent is an AI model that controls a computer through screenshots and mouse and keyboard actions. It works on whatever is on the screen, including browsers, documents, and desktop apps, inside an environment the developer provides.

What is the difference between a computer use agent and a browser agent?

A browser agent reads the page's DOM or accessibility tree and acts on elements, so it only works inside a browser. A computer use agent works from pixels, so it can operate any application. It is usually slower and costs more per step.

Does a computer use agent have its own IP address?

No. It uses the IP of the machine or container its browser runs on. For a self-hosted sandbox on a cloud server, that is the server's datacenter IP unless you route the browser through a proxy. Vendor-hosted agents use the vendor's network.

Can websites detect computer use agents?

Sometimes. A desktop-style agent driving a normal Firefox through OS input doesn't set navigator.webdriver, so the browser looks fairly ordinary. Sites can still see a cloud IP, a small virtual screen, software rendering, and machine-like timing. Agents built on Playwright also expose the standard automation flags.

Should I set the proxy in Docker or in the browser?

Set it in the browser when only the browser needs it. That keeps model API traffic and screenshots off the proxy. Add container environment variables only for command-line tools, with NO_PROXY set for the model API host. Use network isolation when the agent has shell access and must never go direct.

Do computer use agents need residential or ISP proxies?

Only when the default network is the problem, such as a cloud IP being blocked, the wrong country, or several sandboxes sharing one address. Use ISP proxies for recurring logged-in work and sticky residential sessions for one-off tasks in a specific country.

Conclusion

A computer use agent browses the web through a real browser on a virtual desktop, so websites judge it by that browser's network, fingerprint, and behavior. The model never touches the site. Of those three, the IP is the one you control most directly, and on a cloud sandbox it is usually the weakest.

Put the proxy on the browsing path only: a locked browser policy pointing at a forwarder that keeps the credentials off the screen. Verify the exit before each task, keep one exit per sandbox, and isolate the network when the agent has a shell. When you're ready to give your sandboxes a deliberate network identity, compare residential proxies for country-targeted sticky sessions or ISP plans for dedicated static IPs.

About the Author

Unknown Proxies

Proxy Infrastructure Team

Stay Unknown

High-performance dedicated proxies optimized for speed and reliability. Get uncompromising quality, 99.9% uptime, and unmatched support. Stay Unknown.

Explore Plans