---
source_url: "https://github.com/vercel-labs/agent-browser"
title: "GitHub - vercel-labs/agent-browser: Browser automation CLI for AI agents · GitHub"
mirrored_at: 2026-08-04T01:02:47.984Z
host: github.com
cited_in_42a: true
mirror_canonical: "https://index.42a.ai/github.com/vercel-labs/agent-browser"
---

> **Original source:** https://github.com/vercel-labs/agent-browser

Browser automation CLI for AI agents. Fast native Rust CLI.

## Installation

### Global Installation (recommended)

Installs the native Rust binary:

npm install -g agent-browser
agent-browser install  # Download Chrome from Chrome for Testing (first time only)

### Project Installation (local dependency)

For projects that want to pin the version in `package.json`:

npm install agent-browser
agent-browser install

Then use via `package.json` scripts or by invoking `agent-browser` directly.

### Homebrew (macOS)

brew install agent-browser
agent-browser install  # Download Chrome from Chrome for Testing (first time only)

### Cargo (Rust)

cargo install agent-browser
agent-browser install  # Download Chrome from Chrome for Testing (first time only)

### From Source

Requires Node.js 24+, pnpm 11+, and Rust.

git clone https://github.com/vercel-labs/agent-browser
cd agent-browser
pnpm install
pnpm build
pnpm build:native   # Requires Rust (https://rustup.rs)
pnpm link --global  # Makes agent-browser available globally
agent-browser install

### Linux Dependencies

On Linux, install system dependencies:

agent-browser install --with-deps

This exits nonzero if the package manager cannot install every required browser library.

### Updating

Upgrade to the latest version:

agent-browser upgrade

Detects your installation method (npm, Homebrew, or Cargo) and runs the appropriate update command automatically.

### Requirements

-   **Chrome** - Run `agent-browser install` to download Chrome from [Chrome for Testing](https://developer.chrome.com/blog/chrome-for-testing/) (Google's official automation channel). Existing Chrome, Brave, Playwright, and Puppeteer installations are detected automatically. No Playwright or Node.js required for the daemon.
-   **Node.js 24+ and pnpm 11+** - Only needed when building from source.
-   **Rust** - Only needed when building from source (see From Source above).

## Quick Start

agent-browser open example.com
agent-browser snapshot                    # Get accessibility tree with refs
agent-browser click @e2                   # Click by ref from snapshot
agent-browser fill @e3 "test@example.com" # Fill by ref
agent-browser get text @e1                # Get text by ref
agent-browser screenshot page.png
agent-browser close

Clicks fail early when another element covers the target's click point, for example a consent banner or modal. Dismiss or interact with the reported covering element, then take a fresh snapshot before retrying the original ref.

Headless Chromium screenshots hide native scrollbars for consistent image output. Pass `--hide-scrollbars false` when launching to keep native scrollbars visible.

### Traditional Selectors (also supported)

agent-browser click "#submit"
agent-browser fill "#email" "test@example.com"
agent-browser find role button click --name "Submit"

## Commands

### Core Commands

agent-browser open                    # Launch browser (no navigation); stays on about:blank
agent-browser open <url\>              # Launch + navigate to URL (aliases: goto, navigate)
agent-browser read \[url\]              # Fetch agent-readable text, or read rendered active-tab DOM
agent-browser click <sel\>             # Click element (--new-tab to open in new tab)
agent-browser dblclick <sel\>          # Double-click element
agent-browser focus <sel\>             # Focus element
agent-browser type <sel\> <text\>       # Type into element
agent-browser fill <sel\> <text\>       # Clear and fill
agent-browser press <key\>             # Press key (Enter, Tab, Control+a) (alias: key)
agent-browser keyboard type <text\>    # Type with real keystrokes (no selector, current focus)
agent-browser keyboard inserttext <text\>  # Insert text without key events (no selector)
agent-browser keydown <key\>           # Hold key down
agent-browser keyup <key\>             # Release key
agent-browser hover <sel\>             # Hover element
agent-browser select <sel> <val\>      # Select dropdown option
agent-browser check <sel\>             # Check checkbox
agent-browser uncheck <sel\>           # Uncheck checkbox
agent-browser scroll <dir\> \[px\]       # Scroll (up/down/left/right, --selector <sel>)
agent-browser scrollintoview <sel\>    # Scroll element into view (alias: scrollinto)
agent-browser drag <src\> <tgt\>        # Drag and drop
agent-browser upload <sel\> <files\>    # Upload files
agent-browser screenshot \[path\]       # Take screenshot (--full for full page, saves to a temporary directory if no path)
agent-browser screenshot --annotate   # Annotated screenshot with numbered element labels
agent-browser screenshot --screenshot-dir ./shots    # Save to custom directory
agent-browser screenshot --screenshot-format jpeg --screenshot-quality 80
agent-browser pdf <path\>              # Save as PDF
agent-browser snapshot                # Accessibility tree with refs (best for AI)
agent-browser eval <js\>               # Run JavaScript (-b for base64, --stdin for piped input)
agent-browser connect <port\>          # Connect to browser via CDP
agent-browser stream enable \[--port <port\>\]  # Start runtime WebSocket streaming
agent-browser stream status           # Show runtime streaming state and bound port
agent-browser stream disable          # Stop runtime WebSocket streaming
agent-browser close                   # Close browser (aliases: quit, exit)
agent-browser close --all             # Close all active sessions
agent-browser chat "<instruction>"    # AI chat: natural language browser control (single-shot)
agent-browser chat                    # AI chat: interactive REPL mode

### Get Info

agent-browser get text <sel\>          # Get text content
agent-browser get html <sel\>          # Get innerHTML
agent-browser get value <sel\>         # Get input value
agent-browser get attr <sel\> <attr\>   # Get attribute
agent-browser get title               # Get page title
agent-browser get url                 # Get current URL
agent-browser get cdp-url             # Get CDP WebSocket URL (for DevTools, debugging)
agent-browser get count <sel\>         # Count matching elements
agent-browser get box <sel\>           # Get bounding box
agent-browser get styles <sel\>        # Get computed styles

### Read Agent-Friendly Text

agent-browser read
agent-browser read https://example.com/article
agent-browser read https://example.com/article --filter overview
agent-browser read https://example.com/article --outline
agent-browser read https://docs.example.com --llms index --filter auth
agent-browser read https://docs.example.com --llms full --filter auth
agent-browser read example.com/article --require-md
agent-browser read https://example.com/article --json

`read` fetches a URL without launching Chrome. Omit the URL to read the rendered DOM of the active tab in the current browser session, including browser auth state and client-side updates. Explicit URL reads send `Accept: text/markdown` by default, try the same URL with `.md` appended when the first response is not markdown, walk ancestor paths toward `/` to find the nearest `llms.txt` for a matching docs link, print markdown or plain text when available, and fall back to readable text extracted from HTML. `--llms` and `--require-md` with no URL use the active tab URL because they depend on HTTP resources. `read` does not read `llms-full.txt` unless you ask for it.

Options: `--raw` prints the response body without HTML extraction, `--require-md` fails unless the server returns `Content-Type: text/markdown`, `--outline` prints a compact heading outline for one page, `--llms index` prints a compact nearest-ancestor `llms.txt` link list, `--llms full` reads the nearest-ancestor `llms-full.txt`, `--filter <text>` narrows page sections, llms links/sections, or outline headings, and `--timeout <ms>` changes the request timeout. Global safeguards such as `--allowed-domains`, `--content-boundaries`, and `--max-output` also apply to read fetches and output.

### Check State

agent-browser is visible <sel\>        # Check if visible
agent-browser is enabled <sel\>        # Check if enabled
agent-browser is checked <sel\>        # Check if checked

### Find Elements (Semantic Locators)

agent-browser find role <role\> <action\> \[value\]       # By ARIA role
agent-browser find text <text\> <action\> \[value\]       # By text content
agent-browser find label <label\> <action\> \[value\]     # By label
agent-browser find placeholder <ph\> <action\> \[value\]  # By placeholder
agent-browser find alt <text\> <action\> \[value\]        # By alt text
agent-browser find title <text\> <action\> \[value\]      # By title attr
agent-browser find testid <id\> <action\> \[value\]       # By data-testid
agent-browser find first <sel\> <action\> \[value\]       # First match
agent-browser find last <sel\> <action\> \[value\]        # Last match
agent-browser find nth <n\> <sel\> <action\> \[value\]     # Nth match

**Actions:** `click`, `fill`, `check`, `hover`, `text`

**Options:** `--name <name>` (filter role by accessible name), `--exact` (exact, case-sensitive match; for `role` it applies to the accessible name, whose default is a case-insensitive substring)

**Examples:**

agent-browser find role button click --name "Submit"
agent-browser find role heading text --name "Skills"     # implicit roles work: <h2>=heading, <ul>=list, top-level <header>=banner
agent-browser find text "Sign In" click
agent-browser find label "Email" fill "test@test.com"
agent-browser find first ".item" click
agent-browser find nth 2 "a" text

### Wait

agent-browser wait <selector\>         # Wait for element to be visible
agent-browser wait <ms\>               # Wait for time (milliseconds)
agent-browser wait --text "Welcome"   # Wait for text to appear (substring match)
agent-browser wait --url "\*\*/dash"    # Wait for URL pattern
agent-browser wait --load networkidle # Wait for load state
agent-browser wait --fn "window.ready === true"  # Wait for JS condition

# Wait for text/element to disappear
agent-browser wait --fn "!document.body.innerText.includes('Loading...')"
agent-browser wait "#spinner" --state hidden

**Load states:** `load`, `domcontentloaded`, `networkidle`

### Batch Execution

Execute multiple commands in a single invocation. Commands can be passed as quoted arguments or piped as JSON via stdin. This avoids per-command process startup overhead when running multi-step workflows.

# Argument mode: each quoted argument is a full command
agent-browser batch "open https://example.com" "snapshot -i" "screenshot"

# With --bail to stop on first error
agent-browser batch --bail "open https://example.com" "click @e1" "screenshot"

# Stdin mode: pipe commands as JSON
echo '\[
  \["open", "https://example.com"\],
  \["snapshot", "-i"\],
  \["click", "@e1"\],
  \["screenshot", "result.png"\]
\]' | agent-browser batch --json

### Clipboard

agent-browser clipboard read                      # Read text from clipboard
agent-browser clipboard write "Hello, World!"     # Write text to clipboard
agent-browser clipboard copy                      # Copy current selection (Ctrl+C)
agent-browser clipboard paste                     # Paste from clipboard (Ctrl+V)

### Mouse Control

agent-browser mouse move <x\> <y\>      # Move mouse
agent-browser mouse down \[button\]     # Press button (left/right/middle)
agent-browser mouse up \[button\]       # Release button
agent-browser mouse wheel <dy\> \[dx\]   # Scroll wheel

### Browser Settings

agent-browser set viewport <w\> <h\> \[scale\]  # Set viewport size (scale for retina, e.g. 2)
agent-browser set device <name\>       # Emulate device ("iPhone 14")
agent-browser set geo <lat\> <lng\>     # Set geolocation
agent-browser set offline \[on|off\]    # Toggle offline mode
agent-browser set headers <json\>      # Extra HTTP headers
agent-browser set credentials <u\> <p\> # HTTP basic auth
agent-browser set media \[dark|light\]  # Emulate color scheme

### Cookies & Storage

agent-browser cookies                 # Get all cookies
agent-browser cookies set <name\> <val\> # Set cookie
agent-browser cookies set --curl <file\> # Import cookies from a Copy-as-cURL dump,
                                        # JSON array, or bare Cookie header (auto-detected)
agent-browser cookies clear           # Clear cookies

agent-browser storage local           # Get all localStorage
agent-browser storage local <key\>     # Get specific key
agent-browser storage local set <k\> <v\>  # Set value
agent-browser storage local clear     # Clear all

agent-browser storage session         # Same for sessionStorage

### Network

agent-browser network route <url\>              # Intercept requests
agent-browser network route <url\> --abort      # Block requests
agent-browser network route <url\> --body <json\>  # Mock response
agent-browser network route '\*' --abort --resource-type script  # Block scripts only
agent-browser network unroute \[url\]            # Remove routes
agent-browser network requests                 # View tracked requests
agent-browser network requests --filter api    # Filter requests
agent-browser network requests --type xhr,fetch  # Filter by resource type
agent-browser network requests --method POST   # Filter by HTTP method
agent-browser network requests --status 2xx    # Filter by status (200, 2xx, 400-499)
agent-browser network request <requestId\>      # View full request/response detail
agent-browser network har start                # Start HAR recording (embeds text response bodies)
agent-browser network har start --content all  # Embed all response bodies (binary as base64)
agent-browser network har start --content none # Metadata only, no bodies
agent-browser network har stop \[output.har\]    # Stop and save HAR (temp path if omitted)

### Tabs & Windows

agent-browser tab                              # List tabs (shows \`tabId\` and optional label)
agent-browser tab new \[url\]                    # New tab (optionally with URL)
agent-browser tab new --label docs \[url\]       # New tab with a user-assigned label
agent-browser tab <t<N\>|label\>                 # Switch to a tab by id or label
agent-browser tab close \[t<N\>|label\]           # Close a tab (defaults to active)
agent-browser window new                       # New window

Tab ids are stable strings of the form `t1`, `t2`, `t3`. They're never reused within a session, so scripts and agents can keep referring to the same tab even after other tabs are opened or closed. Positional integers like `tab 2` are **not** accepted; the `t` prefix disambiguates handles from indices and mirrors the `@e1` convention used for element refs.

You can also assign a memorable label (`docs`, `app`, `admin`) and use it interchangeably with the id. Labels are never auto-generated and never rewritten on navigation — they're yours to name and keep:

agent-browser tab new --label docs https://docs.example.com
agent-browser tab docs               # switch to the docs tab
agent-browser snapshot               # populate refs for docs
agent-browser click @e3              # click uses docs's refs
agent-browser tab close docs         # close by label

Switching to a tab discarded by Chrome's Memory Saver reactivates it, since a discarded tab has no renderer to drive. Reactivation reloads the discarded page and resets its unsaved state, and the switch result reports `"revived": true`. A tab whose page is paused by a JavaScript dialog is alive rather than discarded, so the switch leaves it untouched and reports `"dialogBlocked": true`; resolve the dialog with `dialog accept` or `dialog dismiss` before interacting. Closing the active tab onto a discarded successor revives it the same way and reports `"activeTabRevived": true`.

### Frames

agent-browser frame <sel\>             # Switch to iframe
agent-browser frame main              # Back to main frame

### Dialogs

agent-browser dialog accept \[text\]    # Accept (with optional prompt text)
agent-browser dialog dismiss          # Dismiss
agent-browser dialog status           # Check if a dialog is currently open

By default, `alert` and `beforeunload` dialogs are automatically accepted so they never block the agent. `confirm` and `prompt` dialogs still require explicit handling. Use `--no-auto-dialog` (or `AGENT_BROWSER_NO_AUTO_DIALOG=1`) to disable automatic handling.

When a JavaScript dialog is pending, all command responses include a `warning` field with the dialog type and message.

### Diff

agent-browser diff snapshot                              # Compare current vs last snapshot
agent-browser diff snapshot --baseline before.txt        # Compare current vs saved snapshot file
agent-browser diff snapshot --selector "#main" --compact # Scoped snapshot diff
agent-browser diff screenshot --baseline before.png      # Visual pixel diff against baseline
agent-browser diff screenshot --baseline b.png -o d.png  # Save diff image to custom path
agent-browser diff screenshot --baseline b.png -t 0.2    # Adjust color threshold (0-1)
agent-browser diff url https://v1.com https://v2.com     # Compare two URLs (snapshot diff)
agent-browser diff url https://v1.com https://v2.com --screenshot  # Also visual diff
agent-browser diff url https://v1.com https://v2.com --wait-until networkidle  # Custom wait strategy
agent-browser diff url https://v1.com https://v2.com --selector "#main"  # Scope to element

### Debug

agent-browser trace start             # Start recording trace
agent-browser trace stop \[path\]       # Stop and save trace
agent-browser profiler start          # Start Chrome DevTools profiling
agent-browser profiler stop \[path\]    # Stop and save profile (.json)
agent-browser console                 # View console messages (log, error, warn, info)
agent-browser console --json          # JSON output with raw CDP args for programmatic access
agent-browser console --clear         # Clear console
agent-browser errors                  # View page errors (uncaught JavaScript exceptions)
agent-browser errors --clear          # Clear errors
agent-browser highlight <sel\>         # Highlight element
agent-browser inspect                 # Open Chrome DevTools for the active page
agent-browser state save <path\>       # Save auth state
agent-browser state load <path\>       # Load auth state
agent-browser state list              # List saved state files
agent-browser state show <file\>       # Show state summary
agent-browser state rename <old\> <new\> # Rename state file
agent-browser state clear \[name\]      # Clear states for session
agent-browser state clear --all       # Clear all saved states
agent-browser state clean --older-than <days\>  # Delete old states

### Navigation

agent-browser back                    # Go back
agent-browser forward                 # Go forward
agent-browser reload                  # Reload page
agent-browser pushstate <url\>         # SPA client-side nav; auto-detects window.next.router.push,
                                      # falls back to history.pushState + popstate

### Pre-navigation setup

Some flows (SSR debug, auth cookies for protected origins, init scripts) need state set up _before_ the first navigation. Use `open` with no URL to launch the browser, then stage cookies / routes / init scripts, then navigate. `batch` sends it all in one CLI call:

agent-browser batch \\
  '\["open"\]' \\
  '\["network","route","\*","--abort","--resource-type","script"\]' \\
  '\["cookies","set","--curl","cookies.curl","--domain","localhost"\]' \\
  '\["navigate","http://localhost:3000/target"\]'

Without `batch` the same sequence is three commands that all reuse the same daemon (fast, but not one turn).

### React / Web Vitals

Agent-browser ships with first-class React introspection and universal Web Vitals metrics. The React commands need the React DevTools hook installed at launch; Web Vitals and pushstate are framework-agnostic.

agent-browser open --enable react-devtools <url\>   # Launch with React hook installed
agent-browser react tree                           # Full component tree
agent-browser react inspect <fiberId\>              # props, hooks, state, source
agent-browser react renders start                  # Begin fiber render recording
agent-browser react renders stop \[--json\]          # Stop and print profile (--json for raw data)
agent-browser react suspense \[--only-dynamic\] \[--json\]  # Suspense boundaries + classifier
                                                         # --only-dynamic hides the "static" list
agent-browser vitals \[url\] \[--json\]                # LCP/CLS/TTFB/FCP/INP + hydration summary

Each `react ...` subcommand requires `--enable react-devtools` to have been passed at launch (the React DevTools `installHook.js` is embedded in the binary). Without it the commands error with \`React DevTools hook not installed

-   relaunch with --enable react-devtools\`.

Works on any React app — Next.js, Remix, Vite+React, CRA, TanStack Start, React Native Web, etc. `vitals` and `pushstate` are framework-agnostic. `vitals` prints a summary by default; pass `--json` for the full structured payload.

### Accessibility audits

Run an [axe-core](https://github.com/dequelabs/axe-core) accessibility audit against the current page or a URL. The axe-core engine is embedded in the binary, so it works offline and under strict CSP. It runs private partial audits across the page's frame tree and merges serialized results without page messaging, so page-provided `window.axe` values remain intact and iframe violations retain their frame selector paths. Accessibility audits require a CDP browser and are not available with Safari or iOS WebDriver sessions.

agent-browser a11y                                 # Audit the current page
agent-browser a11y https://example.com             # Navigate, then audit
agent-browser a11y --tags wcag2a,wcag2aa           # Only rules with these axe tags
agent-browser a11y --selector "#main"              # Scope the audit to a subtree
agent-browser a11y example.com --json              # Full structured results

The default output lists each violation with its impact, rule id, fix guidance URL, and the CSS selectors of failing nodes:

```
url: https://example.com/
axe-core: 4.12.1  violations: 2  incomplete: 0  passes: 24

[critical] image-alt: Images must have alternative text (3 nodes)
  https://dequeuniversity.com/rules/axe/4.12/image-alt
  - img.hero
  - #logo > img
  - footer img
[serious] color-contrast: Elements must meet minimum color contrast ratio thresholds (1 node)
  https://dequeuniversity.com/rules/axe/4.12/color-contrast
  - .nav a.muted
```

`--json` returns the same data structured for automation (`counts`, `violations`, `incomplete`, each violation's `nodes` with `target`, `html`, and `failureSummary`). Each `target` preserves axe's selector path arrays, including nested arrays for shadow DOM boundaries. Rules that axe could not evaluate automatically are reported under `incomplete` for manual review.

### Init scripts

agent-browser open --init-script <path\>           # Register page init script before first navigation
                                                  # (repeatable; also AGENT\_BROWSER\_INIT\_SCRIPTS env)
agent-browser addinitscript <js\>                  # Register at runtime (returns identifier)
agent-browser removeinitscript <identifier\>       # Remove a previously registered init script

### Setup

agent-browser install                 # Download Chrome from Chrome for Testing (Google's official automation channel)
agent-browser install --with-deps     # Also install system deps (Linux)
agent-browser upgrade                 # Upgrade agent-browser to the latest version
agent-browser doctor                  # Diagnose the install and auto-clean stale daemon files
agent-browser doctor --fix            # Also run destructive repairs (reinstall Chrome, purge old state, ...)
agent-browser doctor --offline --quick  # Skip network probes and the live launch test
agent-browser mcp                     # Start an MCP stdio server

`doctor` checks your environment, Chrome install, daemon state, config files, encryption key, providers, network reachability, and runs a live headless browser launch test. Stale socket/pid sidecar files are auto-cleaned. Output is also available as `--json` for agents.

### Skills

agent-browser skills                  # List available skills
agent-browser skills list             # Same as above
agent-browser skills get <name\>       # Output a skill's full content
agent-browser skills get <name\> --full  # Include references and templates
agent-browser skills get --all        # Output every skill
agent-browser skills path \[name\]      # Print skill directory path

Serves bundled skill content that always matches the installed CLI version. AI agents use this to get current instructions rather than relying on cached copies. Set `AGENT_BROWSER_SKILLS_DIR` to override the skills directory path.

### MCP Server

agent-browser mcp
agent-browser mcp --tools all
agent-browser mcp --tools core,network,react

Starts a Model Context Protocol server over stdio. MCP clients launch this command as a subprocess and exchange newline-delimited JSON-RPC on stdin and stdout. The server defaults to MCP protocol 2025-11-25 and accepts older supported client protocol versions during initialization.

The default tools profile is `core`, which keeps MCP context small for everyday browser automation. Use `--tools all` for the full typed CLI parity surface, or combine profiles with commas, such as `--tools core,network,react`.

Profiles:

-   `core` — Default. Navigation, snapshots, interaction, waits, reads, screenshots, JavaScript eval, close, tab basics, and profile discovery
-   `network` — Network routes, request inspection, HAR, headers, credentials, offline
-   `state` — Cookies, storage, auth, saved state, sessions, profiles, skills
-   `debug` — Console/errors, tracing, profiling, recording, a11y audit, clipboard, plugins, doctor, dashboard, install, upgrade, chat, diff, batch, confirm/deny
-   `tabs` — Back/forward/reload, tabs, windows, frames, dialogs
-   `react` — React tree/inspect/renders/suspense, vitals, pushstate
-   `mobile` — Viewport/device/geolocation/media, touch, swipe, mouse, keyboard
-   `all` — Every MCP tool, including the full typed CLI parity surface

Common tools include:

-   `agent_browser_tools_profiles`
-   `agent_browser_open`
-   `agent_browser_snapshot`
-   `agent_browser_click`
-   `agent_browser_fill`
-   `agent_browser_type`
-   `agent_browser_press`
-   `agent_browser_wait_for_selector`
-   `agent_browser_screenshot`
-   `agent_browser_get_url`
-   `agent_browser_eval`
-   `agent_browser_close`

Each tool has typed fields such as `url`, `selector`, `text`, `key`, `session`, and `allowedDomains`, so MCP clients show meaningful approval prompts instead of raw command arrays. The common `allowedDomains` array maps to `--allowed-domains` and activates the same WebRTC containment and launch-mode restrictions. Each tool also accepts `extraArgs` for advanced CLI flags and exact CLI parity. Tool discovery is paginated and includes read-only/open-world annotations so modern MCP clients can load the large typed surface incrementally.

Example MCP client config:

{
  "mcpServers": {
    "agent-browser": {
      "command": "agent-browser",
      "args": \["mcp"\]
    }
  }
}

Full parity MCP client config:

{
  "mcpServers": {
    "agent-browser": {
      "command": "agent-browser",
      "args": \["mcp", "\--tools", "all"\]
    }
  }
}

Tool invocations use the same config files and environment variables as the CLI. Use `session` in the tool arguments, or set `AGENT_BROWSER_SESSION`, to isolate browser state.

## Authentication

agent-browser provides multiple ways to persist login sessions so you don't re-authenticate every run.

### Quick summary

Approach

Best for

Flag / Env

**Chrome profile reuse**

Reuse your existing Chrome login state (cookies, sessions) with zero setup

`--profile <name>` / `AGENT_BROWSER_PROFILE`

**Persistent profile**

Full browser state (cookies, IndexedDB, service workers, cache) across restarts

`--profile <path>` / `AGENT_BROWSER_PROFILE`

**Session persistence**

Auto-save/restore cookies + localStorage from a stable session key

`--session <id> --restore` / `AGENT_BROWSER_RESTORE`

**Import from your browser**

Grab auth from a Chrome session you already logged into

`--auto-connect` + `state save`

**State file**

Load a previously saved state JSON on launch

`--state <path>` / `AGENT_BROWSER_STATE`

**Auth vault**

Store credentials locally (encrypted), login by name

`auth save` / `auth login`

### Import auth from your browser

If you are already logged in to a site in Chrome, you can grab that auth state and reuse it:

# 1. Launch Chrome with remote debugging enabled
#    macOS:
"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" --remote-debugging-port=9222
#    Or use --auto-connect to discover an already-running Chrome

# 2. Connect and save the authenticated state
agent-browser --auto-connect state save ./my-auth.json

# 3. Use the saved auth in future sessions
agent-browser --state ./my-auth.json open https://app.example.com/dashboard

# 4. Or use --restore for automatic persistence
SESSION="$(agent-browser session id --scope worktree --prefix myapp)"
agent-browser --session "$SESSION" --restore --state ./my-auth.json open https://app.example.com/dashboard
# From now on, --session "$SESSION" --restore auto-saves/restores this state

> **Security notes:**
> 
> -   `--remote-debugging-port` exposes full browser control on localhost. Any local process can connect. Only use on trusted machines and close Chrome when done.
> -   State files contain session tokens in plaintext. Add them to `.gitignore` and delete when no longer needed. For encryption at rest, set `AGENT_BROWSER_ENCRYPTION_KEY` (see [State Encryption](#state-encryption)).

For full details on login flows, OAuth, 2FA, cookie-based auth, and the auth vault, see the [Authentication](https://github.com/vercel-labs/agent-browser/blob/main/docs/src/app/sessions/page.mdx) docs.

## Sessions

Run multiple isolated browser instances:

# Different sessions
agent-browser --session agent1 open site-a.com
agent-browser --session agent2 open site-b.com

# Or via environment variable
AGENT\_BROWSER\_SESSION=agent1 agent-browser click "#btn"

# List active sessions
agent-browser session list
# Output:
# Active sessions:
# -> default
#    agent1

# Show current session
agent-browser session

# Generate a stable worktree-scoped session id
agent-browser session id --scope worktree --prefix next-dev-loop

# Inspect daemon, launch, and restore status
agent-browser session info --json

Each session has its own:

-   Browser instance
-   Cookies and storage
-   Navigation history
-   Authentication state

## Chrome Profile Reuse

The fastest way to use your existing login state: pass a Chrome profile name to `--profile`:

# List available Chrome profiles
agent-browser profiles

# Reuse your default Chrome profile's login state
agent-browser --profile Default open https://gmail.com

# Use a named profile (by display name or directory name)
agent-browser --profile "Work" open https://app.example.com

# Or via environment variable
AGENT\_BROWSER\_PROFILE=Default agent-browser open https://gmail.com

This copies your Chrome profile to a temp directory (read-only snapshot, no changes to your original profile), so the browser launches with your existing cookies and sessions.

> **Note:** On Windows, close Chrome before using `--profile <name>` if Chrome is running, as some profile files may be locked.

## Persistent Profiles

For a persistent custom profile directory that stores state across browser restarts, pass a path to `--profile`:

# Use a persistent profile directory
agent-browser --profile ~/.myapp-profile open myapp.com

# Login once, then reuse the authenticated session
agent-browser --profile ~/.myapp-profile open myapp.com/dashboard

# Or via environment variable
AGENT\_BROWSER\_PROFILE=~/.myapp-profile agent-browser open myapp.com

The profile directory stores:

-   Cookies and localStorage
-   IndexedDB data
-   Service workers
-   Browser cache
-   Login sessions

**Tip**: Use different profile paths for different projects to keep their browser state isolated.

## Session Persistence

Use `--restore` with a stable `--session` to automatically save and restore cookies and localStorage across browser restarts:

# Generate a stable id for this worktree and auto-save/load state
SESSION="$(agent-browser session id --scope worktree --prefix twitter)"
agent-browser --session "$SESSION" --restore open twitter.com

# Login once, then state persists automatically
# State files stored in ~/.agent-browser/sessions/

# Optional: validate restored state before auto-saving again
agent-browser --session "$SESSION" --restore --restore-check-text Dashboard open twitter.com

State is saved when the browser closes (explicit `close`, idle timeout, or daemon shutdown) and also periodically while the browser is open, so a browser window you close by hand still leaves a recent save behind. Periodic autosave waits for commands to settle, then saves at most once per `AGENT_BROWSER_AUTOSAVE_INTERVAL_MS` (default 30000; set to `0` to save only on close). Idle sessions keep saving on the same interval, so changes the page makes on its own (token refreshes, background requests) are captured too. It respects the `--restore-save` policy.

### State Encryption

Encrypt saved session data at rest with AES-256-GCM:

# Generate key: openssl rand -hex 32
export AGENT\_BROWSER\_ENCRYPTION\_KEY=<64-char-hex-key\>

# State files are now encrypted automatically
agent-browser --session secure --restore open example.com

Variable

Description

`AGENT_BROWSER_RESTORE`

Auto-save/load state persistence name

`AGENT_BROWSER_RESTORE_SAVE`

Restore save policy: `auto`, `always`, or `never`

`AGENT_BROWSER_AUTOSAVE_INTERVAL_MS`

Min ms between periodic autosaves (default: 30000, 0 disables)

`AGENT_BROWSER_NAMESPACE`

Namespace for daemon sockets and restore state

`AGENT_BROWSER_SESSION_NAME`

Legacy auto-save/load state persistence name

`AGENT_BROWSER_ENCRYPTION_KEY`

64-char hex key for AES-256-GCM encryption

`AGENT_BROWSER_STATE_EXPIRE_DAYS`

Auto-delete states older than N days (default: 30)

## Security

agent-browser includes security features for safe AI agent deployments. All features are opt-in, and existing workflows are unaffected until you explicitly enable a feature:

-   **Authentication Vault**: Store credentials locally (always encrypted), reference by name. The LLM never sees passwords. `auth login` navigates with `load` and then waits for login form selectors to appear (SPA-friendly, timeout follows the default action timeout). A key is auto-generated at `~/.agent-browser/.encryption-key` if `AGENT_BROWSER_ENCRYPTION_KEY` is not set: `echo "pass" | agent-browser auth save github --url https://github.com/login --username user --password-stdin` then `agent-browser auth login github`
-   **Plugin System**: Extend agent-browser with external executable plugins. Plugins run out-of-process over the `agent-browser.plugin.v1` stdio JSON protocol and declare capabilities such as `credential.read`, `browser.provider`, `launch.mutate`, or `command.run`.
-   **Content Boundary Markers**: Wrap page output in delimiters so LLMs can distinguish tool output from untrusted content: `--content-boundaries`
-   **Domain Allowlist**: Restrict navigation to trusted domains (wildcards like `*.example.com` also match the bare domain): `--allowed-domains "example.com,*.example.com"`. Sub-resource requests (scripts, images, fetch), WebSocket/EventSource connections, and `sendBeacon` calls to non-allowed domains are blocked. WebRTC peer connections are disabled in supported Chromium sessions while the allowlist is active to prevent STUN, TURN, and DNS traffic from bypassing HTTP interception. Dedicated and shared workers are guarded with a bootstrap wrapper; if a page CSP forbids that wrapper, the worker fails closed rather than running without the allowlist guard. Pre-existing CDP sessions, auto-connect, Chrome profiles, direct-page provider plugins, agent-browser restore or state-file replay, raw Chrome args that select profiles, restore sessions, or open startup pages, iOS, and Safari reject this option because agent-browser cannot install equivalent containment before page scripts run. Include any CDN domains your target pages depend on (e.g., `*.cdn.example.com`).
-   **Action Policy**: Gate destructive actions with a static policy file: `--action-policy ./policy.json`
-   **Action Confirmation**: Require explicit approval for sensitive action categories: `--confirm-actions eval,download`
-   **Output Length Limits**: Prevent context flooding: `--max-output 50000`

Variable

Description

`AGENT_BROWSER_CONTENT_BOUNDARIES`

Wrap page output in boundary markers

`AGENT_BROWSER_MAX_OUTPUT`

Max characters for page output

`AGENT_BROWSER_ALLOWED_DOMAINS`

Comma-separated allowed domain patterns; requires a fresh controllable browser context without profile/session startup args, restore/state replay, or direct-page provider plugins

`AGENT_BROWSER_ACTION_POLICY`

Path to action policy JSON file

`AGENT_BROWSER_CONFIRM_ACTIONS`

Action categories requiring confirmation

`AGENT_BROWSER_CONFIRM_INTERACTIVE`

Enable interactive confirmation prompts

`AGENT_BROWSER_PLUGINS`

JSON plugin registry override

See [Security documentation](https://agent-browser.dev/security) for details.

### Plugin System

Plugins let third-party tools integrate without becoming built-in agent-browser dependencies. Add a plugin from npm or GitHub:

agent-browser plugin add agent-browser-plugin-captcha
agent-browser plugin add @company/agent-browser-plugin-vault --name vault
agent-browser plugin add org/agent-browser-plugin-cloud-browser

References are resolved by shape: `name` uses npm, `@scope/name` uses npm, and `owner/repo` uses GitHub. `plugin add` writes `./agent-browser.json` by default; use `--global` for `~/.agent-browser/config.json`.

Plugin packages should support `plugin.manifest` so `plugin add` can discover their name and capabilities automatically. If a plugin does not support manifests, pass `--capability <name>` during add.

Plugins can also be configured manually in `agent-browser.json`:

{
  "plugins": \[
    {
      "name": "vault",
      "command": "agent-browser-plugin-vault",
      "capabilities": \["credential.read"\]
    },
    {
      "name": "cloud-browser",
      "command": "agent-browser-plugin-cloud-browser",
      "capabilities": \["browser.provider"\]
    },
    {
      "name": "stealth",
      "command": "agent-browser-plugin-stealth",
      "capabilities": \["launch.mutate"\]
    },
    {
      "name": "captcha",
      "command": "agent-browser-plugin-captcha",
      "capabilities": \["command.run", "captcha.solve"\]
    }
  \]
}

Inspect configured plugins:

agent-browser plugin list
agent-browser plugin show vault

Use a credential provider plugin for one login:

agent-browser auth login my-app --credential-provider vault --item "My App"
agent-browser auth login my-app --credential-provider vault --item "My App" --url https://app.example.com/login --username-selector "#email" --password-selector "#password" --submit-selector "button\[type=submit\]"

Use a browser provider plugin:

agent-browser --provider cloud-browser open https://example.com

Use a launch mutator plugin for stealth or local launch customization. The plugin can append Chrome args, extensions, and init scripts before the browser starts:

agent-browser open https://example.com

Use a generic plugin command for domain-specific tools such as CAPTCHA solvers:

agent-browser plugin run captcha captcha.solve --payload '{"siteKey":"...","url":"https://example.com"}'

The protocol request always includes `protocol`, `type`, `capability`, and `request`. A credential plugin receives `credential.resolve`, a browser provider receives `browser.launch`, a launch mutator receives `launch.mutate`, and generic commands receive the supplied request type. `plugin run` is for `command.run` and custom capabilities; core capabilities and protocol request types use their dedicated command paths. agent-browser keeps browser automation, redaction-sensitive output, and policy enforcement in core.

Gate plugin access by capability action:

agent-browser --confirm-actions plugin:vault:credential.read auth login my-app --credential-provider vault --item "My App"
agent-browser --confirm-actions plugin:cloud-browser:browser.provider --provider cloud-browser open https://example.com
agent-browser --confirm-actions plugin:stealth:launch.mutate open https://example.com

Do not put vault tokens or passwords in plugin command args. Use the vault vendor's own login/session mechanism or environment outside agent-browser config.

## Snapshot Options

The `snapshot` command supports filtering to reduce output size:

agent-browser snapshot                    # Full accessibility tree
agent-browser snapshot -i                 # Interactive elements only (buttons, inputs, links)
agent-browser snapshot -i --urls          # Interactive elements with link URLs
agent-browser snapshot -c                 # Compact (remove empty structural elements)
agent-browser snapshot -d 3               # Limit depth to 3 levels
agent-browser snapshot -s "#main"         # Scope to CSS selector
agent-browser snapshot -i -c -d 5         # Combine options

Option

Description

`-i, --interactive`

Only show interactive elements (buttons, links, inputs)

`-u, --urls`

Include href URLs for link elements

`-c, --compact`

Remove empty structural elements

`-d, --depth <n>`

Limit tree depth

`-s, --selector <sel>`

Scope to CSS selector

## Annotated Screenshots

The `--annotate` flag overlays numbered labels on interactive elements in the screenshot. Each label `[N]` corresponds to ref `@eN`, so the same refs work for both visual and text-based workflows.

Annotated screenshots are supported on the CDP-backed browser path (Chrome/Lightpanda). The Safari/WebDriver backend does not yet support `--annotate`.

agent-browser screenshot --annotate
# -> Screenshot saved to /tmp/screenshot-2026-02-17T12-00-00-abc123.png
#    \[1\] @e1 button "Submit"
#    \[2\] @e2 link "Home"
#    \[3\] @e3 textbox "Email"

After an annotated screenshot, refs are cached so you can immediately interact with elements:

agent-browser screenshot --annotate ./page.png
agent-browser click @e2     # Click the "Home" link labeled \[2\]

This is useful for multimodal AI models that can reason about visual layout, unlabeled icon buttons, canvas elements, or visual state that the text accessibility tree cannot capture.

## Options

Option

Description

`--session <name>`

Use isolated session (or `AGENT_BROWSER_SESSION` env)

`--restore [name]`

Auto-save/restore session state. Bare `--restore` uses `--session` as the key

`--restore-save <policy>`

Restore save policy: `auto`, `always`, or `never`

`--restore-check-url <glob>`

Validate restored state against a URL pattern

`--restore-check-text <text>`

Validate restored state against page text

`--restore-check-fn <js>`

Validate restored state against a truthy JavaScript expression

`--namespace <name>`

Isolate daemon sockets and restore-state directories

`--session-name <name>`

Legacy alias for restore persistence key

`--profile <name|path>`

Chrome profile name or persistent directory path (or `AGENT_BROWSER_PROFILE` env)

`--state <path>`

Load storage state from JSON file (or `AGENT_BROWSER_STATE` env)

`--headers <json>`

Set HTTP headers scoped to the URL's origin

`--executable-path <path>`

Custom browser executable (or `AGENT_BROWSER_EXECUTABLE_PATH` env)

`--extension <path>`

Load browser extension (repeatable; or `AGENT_BROWSER_EXTENSIONS` env)

`--init-script <path>`

Register a page init script before the first navigation (repeatable; or `AGENT_BROWSER_INIT_SCRIPTS` env)

`--enable <feature>`

Built-in init scripts: `react-devtools` (repeatable or comma-list; or `AGENT_BROWSER_ENABLE` env)

`--args <args>`

Browser launch args, comma or newline separated (or `AGENT_BROWSER_ARGS` env)

`--user-agent <ua>`

Custom User-Agent string (or `AGENT_BROWSER_USER_AGENT` env)

`--proxy <url>`

Proxy server URL with optional auth (or `AGENT_BROWSER_PROXY` env)

`--proxy-bypass <hosts>`

Hosts to bypass proxy (or `AGENT_BROWSER_PROXY_BYPASS` env)

`--ignore-https-errors`

Ignore HTTPS certificate errors (useful for self-signed certs)

`--allow-file-access`

Allow file:// URLs to access local files (Chromium only)

`--hide-scrollbars <bool>`

Hide native scrollbars in headless Chromium screenshots, enabled by default (or `AGENT_BROWSER_HIDE_SCROLLBARS` env)

`-p, --provider <name>`

Browser provider, including configured `browser.provider` plugins (or `AGENT_BROWSER_PROVIDER` env)

`--device <name>`

iOS device name, e.g. "iPhone 15 Pro" (or `AGENT_BROWSER_IOS_DEVICE` env)

`--json`

JSON output (for agents)

`--annotate`

Annotated screenshot with numbered element labels (or `AGENT_BROWSER_ANNOTATE` env)

`--screenshot-dir <path>`

Default screenshot output directory (or `AGENT_BROWSER_SCREENSHOT_DIR` env)

`--screenshot-quality <n>`

JPEG quality 0-100 (or `AGENT_BROWSER_SCREENSHOT_QUALITY` env)

`--screenshot-format <fmt>`

Screenshot format: `png`, `jpeg` (or `AGENT_BROWSER_SCREENSHOT_FORMAT` env)

`--headed`

Show browser window (not headless) (or `AGENT_BROWSER_HEADED` env)

`--webgpu`

Enable WebGPU; SwiftShader software Vulkan on Linux, no GPU required (or `AGENT_BROWSER_WEBGPU` env)

`--cdp <port|url>`

Connect via Chrome DevTools Protocol (port or WebSocket URL)

`--auto-connect`

Auto-discover and connect to running Chrome (or `AGENT_BROWSER_AUTO_CONNECT` env)

`--color-scheme <scheme>`

Color scheme: `dark`, `light`, `no-preference` (or `AGENT_BROWSER_COLOR_SCHEME` env)

`--download-path <path>`

Default download directory (or `AGENT_BROWSER_DOWNLOAD_PATH` env)

`--content-boundaries`

Wrap page output in boundary markers for LLM safety (or `AGENT_BROWSER_CONTENT_BOUNDARIES` env)

`--max-output <chars>`

Truncate page output to N characters (or `AGENT_BROWSER_MAX_OUTPUT` env)

`--allowed-domains <list>`

Comma-separated allowed domain patterns; also disables WebRTC peer connections in supported Chromium sessions and rejects CDP, auto-connect, Chrome profiles, restore/state replay, direct-page provider plugins, unsafe startup `--args`, iOS, and Safari (or `AGENT_BROWSER_ALLOWED_DOMAINS` env)

`--action-policy <path>`

Path to action policy JSON file (or `AGENT_BROWSER_ACTION_POLICY` env)

`--confirm-actions <list>`

Action categories requiring confirmation (or `AGENT_BROWSER_CONFIRM_ACTIONS` env)

`--confirm-interactive`

Interactive confirmation prompts; auto-denies if stdin is not a TTY (or `AGENT_BROWSER_CONFIRM_INTERACTIVE` env)

`--engine <name>`

Browser engine: `chrome` (default), `lightpanda` (or `AGENT_BROWSER_ENGINE` env)

`--idle-timeout <time>`

Shut down the daemon after inactivity (`10s`, `3m`, `1h`, or raw ms). Defaults to `1h`; use `0` to disable (or `AGENT_BROWSER_IDLE_TIMEOUT_MS` env)

`--no-auto-dialog`

Disable automatic dismissal of `alert`/`beforeunload` dialogs (or `AGENT_BROWSER_NO_AUTO_DIALOG` env)

`--model <name>`

AI model for chat command (or `AI_GATEWAY_MODEL` env)

`-v`, `--verbose`

Show tool commands and their raw output (chat)

`-q`, `--quiet`

Show only AI text responses, hide tool calls (chat)

`--config <path>`

Use a custom config file (or `AGENT_BROWSER_CONFIG` env)

`--debug`

Debug output

## Observability Dashboard

Monitor agent-browser sessions in real time with a local web dashboard showing a live viewport and command activity feed.

# Start the dashboard server (runs in background on port 4848)
agent-browser dashboard start
agent-browser dashboard start --port 8080   # Custom port

# All sessions are automatically visible in the dashboard
agent-browser open example.com

# Stop the dashboard
agent-browser dashboard stop

The dashboard runs as a standalone background process on port 4848, independent of browser sessions. It stays available even when no sessions are running, and it works from `http://localhost:4848` or a proxied/forwarded URL that reaches the dashboard server, such as `https://dashboard.agent-browser.localhost` or a Coder workspace URL. The browser stays on the dashboard origin; session-specific tabs, status, and stream traffic are proxied internally, so session ports do not need to be exposed.

The dashboard displays:

-   **Live viewport**: real-time JPEG frames from the browser
-   **Activity feed**: chronological command/result stream with timing and expandable details
-   **Console output**: browser console messages (log, warn, error)
-   **Session creation**: create new sessions from the UI with local engines (Chrome, Lightpanda) or cloud providers (AgentCore, Browserbase, Browserless, Browser Use, Kernel)
-   **AI Chat**: chat with an AI assistant directly in the dashboard (requires Vercel AI Gateway configuration)

### AI Chat

The dashboard includes an optional AI chat panel powered by the Vercel AI Gateway. The same functionality is available directly from the CLI via the `chat` command. Set these environment variables to enable AI chat:

export AI\_GATEWAY\_API\_KEY=gw\_your\_key\_here
export AI\_GATEWAY\_MODEL=anthropic/claude-sonnet-4.6           # optional, this is the default
export AI\_GATEWAY\_URL=https://ai-gateway.vercel.sh           # optional, this is the default

**CLI usage:**

agent-browser chat "open google.com and search for cats"     # Single-shot
agent-browser chat                                           # Interactive REPL
agent-browser -q chat "summarize this page"                  # Quiet mode (text only)
agent-browser -v chat "fill in the login form"               # Verbose (show command output)
agent-browser --model openai/gpt-4o chat "take a screenshot" # Override model

The `chat` command translates natural language instructions into agent-browser commands, executes them, and streams the AI response. In interactive mode, type `quit` to exit. Use `--json` for structured output suitable for agent consumption.

**Dashboard usage:**

The Chat tab is always visible in the dashboard. When `AI_GATEWAY_API_KEY` is set, the Rust server proxies requests to the gateway and streams responses back using the Vercel AI SDK's UI Message Stream protocol. Without the key, sending a message shows an error inline.

## Configuration

Create an `agent-browser.json` file to set persistent defaults instead of repeating flags on every command.

**Locations (lowest to highest priority):**

1.  `~/.agent-browser/config.json`: user-level defaults
2.  `./agent-browser.json`: project-level overrides (in working directory)
3.  `AGENT_BROWSER_*` environment variables override config file values
4.  CLI flags override everything

**Example `agent-browser.json`:**

{
  "headed": true,
  "proxy": "http://localhost:8080",
  "profile": "./browser-data",
  "userAgent": "my-agent/1.0",
  "hideScrollbars": false,
  "ignoreHttpsErrors": true,
  "plugins": \[
    {
      "name": "vault",
      "command": "agent-browser-plugin-vault",
      "capabilities": \["credential.read"\]
    }
  \]
}

Use `--config <path>` or `AGENT_BROWSER_CONFIG` to load a specific config file instead of the defaults:

agent-browser --config ./ci-config.json open example.com
AGENT\_BROWSER\_CONFIG=./ci-config.json agent-browser open example.com

All options from the table above can be set in the config file using camelCase keys (e.g., `--executable-path` becomes `"executablePath"`, `--proxy-bypass` becomes `"proxyBypass"`). Plugins are configured with the `"plugins"` array shown above. Unknown keys are ignored for forward compatibility.

A [JSON Schema](https://github.com/vercel-labs/agent-browser/blob/main/agent-browser.schema.json) is available for IDE autocomplete and validation. Add a `$schema` key to your config file to enable it:

{
  "$schema": "https://agent-browser.dev/schema.json",
  "headed": true
}

Boolean flags accept an optional `true`/`false` value to override config settings. For example, `--headed false` disables `"headed": true` from config. A bare `--headed` is equivalent to `--headed true`.

Auto-discovered config files that are missing are silently ignored. If `--config <path>` points to a missing or invalid file, agent-browser exits with an error. Extensions from user and project configs are merged (concatenated), not replaced.

> **Tip:** If your project-level `agent-browser.json` contains environment-specific values (paths, proxies), consider adding it to `.gitignore`.

## Default Timeout

The default timeout for standard operations (clicks, waits, fills, etc.) is 25 seconds. This is intentionally below the CLI's 30-second IPC read timeout so that the daemon returns a proper error instead of the CLI timing out with EAGAIN.

Override the default timeout via environment variable:

# Set a longer timeout for slow pages (in milliseconds)
export AGENT\_BROWSER\_DEFAULT\_TIMEOUT=45000

> **Note:** Setting this above 30000 (30s) may cause EAGAIN errors on slow operations because the CLI's read timeout will expire before the daemon responds. The CLI retries transient errors automatically, but response times will increase.

Variable

Description

`AGENT_BROWSER_DEFAULT_TIMEOUT`

Default operation timeout in ms (default: 25000)

## Selectors

### Refs (Recommended for AI)

Refs provide deterministic element selection from snapshots:

# 1. Get snapshot with refs
agent-browser snapshot
# Output:
# - heading "Example Domain" \[ref=e1\] \[level=1\]
# - button "Submit" \[ref=e2\]
# - textbox "Email" \[ref=e3\]
# - link "Learn more" \[ref=e4\]

# 2. Use refs to interact
agent-browser click @e2                   # Click the button
agent-browser fill @e3 "test@example.com" # Fill the textbox
agent-browser get text @e1                # Get heading text
agent-browser hover @e4                   # Hover the link

When a ref click is blocked by an overlay, the error includes the covering element, such as `covered by <div#consent-banner>`. Click the banner or dialog control first, then run `snapshot` again before reusing refs.

**Why use refs?**

-   **Deterministic**: Ref points to exact element from snapshot
-   **Fast**: No DOM re-query needed
-   **AI-friendly**: Snapshot + ref workflow is optimal for LLMs

### CSS Selectors

agent-browser click "#id"
agent-browser click ".class"
agent-browser click "div > button"

### Text & XPath

agent-browser click "text=Submit"
agent-browser click "xpath=//button"

### Semantic Locators

agent-browser find role button click --name "Submit"
agent-browser find label "Email" fill "test@test.com"

## Agent Mode

Use `--json` for machine-readable output:

agent-browser snapshot --json
# Returns: {"success":true,"data":{"snapshot":"...","refs":{"e1":{"role":"heading","name":"Title"},...}}}

agent-browser get text @e1 --json
agent-browser is visible @e2 --json

### Optimal AI Workflow

# 1. Navigate and get snapshot
agent-browser open example.com
agent-browser snapshot -i --json   # AI parses tree and refs

# 2. AI identifies target refs from snapshot
# 3. Execute actions using refs
agent-browser click @e2
agent-browser fill @e3 "input text"

# 4. Get new snapshot if page changed
agent-browser snapshot -i --json

### Command Chaining

Commands can be chained with `&&` in a single shell invocation. The browser persists via a background daemon, so chaining is safe and more efficient:

# Open, wait for load, and snapshot in one call
agent-browser open example.com && agent-browser wait --load networkidle && agent-browser snapshot -i

# Chain multiple interactions
agent-browser fill @e1 "user@example.com" && agent-browser fill @e2 "pass" && agent-browser click @e3

# Navigate and screenshot
agent-browser open example.com && agent-browser wait --load networkidle && agent-browser screenshot page.png

Use `&&` when you don't need intermediate output. Run commands separately when you need to parse output first (e.g., snapshot to discover refs before interacting).

## Headed Mode

Show the browser window for debugging:

agent-browser open example.com --headed

This opens a visible browser window instead of running headless.

On Linux hosts with no display (servers, containers), `--headed` still works: when `DISPLAY` is unset and Xvfb is installed, agent-browser starts a private virtual display for the browser and cleans it up on close (opt out with `AGENT_BROWSER_NO_XVFB=1`). Needed for [WebGPU screenshots](#webgpu), and useful for extensions that misbehave headless.

> **Note:** Browser extensions work in both headed and headless mode (Chrome's `--headless=new`).

## WebGPU

Headless Chrome does not expose WebGPU by default, so pages using it (three.js `WebGPURenderer`, Babylon.js, etc.) silently render black. The `--webgpu` flag enables a launch preset that makes WebGPU work, including in GPU-less containers and CI:

agent-browser --webgpu open https://my-webgpu-app.example.com
agent-browser screenshot app.png

On macOS and Windows this uses the hardware Metal/D3D backend. On Linux it routes WebGPU through SwiftShader's software Vulkan (no GPU needed), which requires the system Vulkan loader and Mesa ICD:

apt-get install -y libvulkan1 mesa-vulkan-drivers

One upstream caveat: headless Chrome cannot capture WebGPU canvas presentation in screenshots on Windows and Linux (rendering and in-page readbacks work; the capture is black). Screenshots of WebGPU pages work headless on macOS; on Windows run `--headed` in a logged-in desktop session; on Linux just add `--headed` — when no `DISPLAY` is set and Xvfb is installed, agent-browser starts a private virtual display automatically (opt out with `AGENT_BROWSER_NO_XVFB=1`).

Verify the full pipeline (adapter, render pass, and screenshot capture) with:

agent-browser doctor --webgpu

Notes for WebGPU pages:

-   WebGPU only exists in secure contexts (`https://`, `http://localhost`, or `file://`).
-   three.js `WebGPURenderer` initializes asynchronously and silently falls back to WebGL2 when no adapter is available — wait for the app to render its first frame before taking a screenshot.
-   To prefer a real GPU on Linux instead of SwiftShader, override both the Vulkan driver and the adapter with `--args "--use-vulkan=native,--use-webgpu-adapter=default"` (user args win over the preset; `--use-webgpu-adapter` alone still enumerates only SwiftShader).

See the [WebGPU docs page](https://agent-browser.dev/webgpu) for the full platform matrix and container recipe.

## Authenticated Sessions

Use `--headers` to set HTTP headers for a specific origin, enabling authentication without login flows:

# Headers are scoped to api.example.com only
agent-browser open api.example.com --headers '{"Authorization": "Bearer <token>"}'

# Requests to api.example.com include the auth header
agent-browser snapshot -i --json
agent-browser click @e2

# Navigate to another domain - headers are NOT sent (safe!)
agent-browser open other-site.com

This is useful for:

-   **Skipping login flows** - Authenticate via headers instead of UI
-   **Switching users** - Start new sessions with different auth tokens
-   **API testing** - Access protected endpoints directly
-   **Security** - Headers are scoped to the origin, not leaked to other domains

To set headers for multiple origins, use `--headers` with each `open` command:

agent-browser open api.example.com --headers '{"Authorization": "Bearer token1"}'
agent-browser open api.acme.com --headers '{"Authorization": "Bearer token2"}'

For global headers (all domains), use `set headers`:

agent-browser set headers '{"X-Custom-Header": "value"}'

## Custom Browser Executable

Use a custom browser executable instead of the bundled Chromium. This is useful for:

-   **Serverless deployment**: Use lightweight Chromium builds like `@sparticuz/chromium` (~50MB vs ~684MB)
-   **System browsers**: Use an existing Chrome/Chromium installation
-   **Custom builds**: Use modified browser builds

### CLI Usage

# Via flag
agent-browser --executable-path /path/to/chromium open example.com

# Via environment variable
AGENT\_BROWSER\_EXECUTABLE\_PATH=/path/to/chromium agent-browser open example.com

### Serverless (Vercel)

Run agent-browser + Chrome in an ephemeral Vercel Sandbox microVM. No external server needed:

import { runAgentBrowserCommand, withAgentBrowserSandbox } from "@agent-browser/sandbox/vercel";

const result \= await withAgentBrowserSandbox(async (sandbox) \=> {
  await runAgentBrowserCommand(sandbox, \["open", "https://example.com"\]);
  return runAgentBrowserCommand(sandbox, \["screenshot"\]);
});

Install `@agent-browser/sandbox` and `@vercel/sandbox` in the consuming app. See the [sandbox helper example](https://github.com/vercel-labs/agent-browser/blob/main/examples/sandbox) for minimal Vercel Sandbox usage, or the [environments example](https://github.com/vercel-labs/agent-browser/blob/main/examples/environments) for a full UI demo with a deploy-to-Vercel button.

Fresh Vercel and eve sandboxes install Chromium system dependencies by default. Pass `installSystemDependencies: false` only when your sandbox image already includes those libraries.

### eve extension

Give an [eve](https://eve.dev/) agent the full browser tool set by mounting the [`@agent-browser/eve`](https://github.com/vercel-labs/agent-browser/blob/main/packages/@agent-browser/eve) extension:

// agent/extensions/browser.ts
import browser from "@agent-browser/eve";

export default browser({});

This composes ~20 namespaced tools into the agent — `browser__navigate`, `browser__snapshot`, `browser__click`, `browser__fill`, `browser__find`, `browser__screenshot`, and more — all running agent-browser inside the agent's sandbox. agent-browser installs automatically on first use; pre-install it in `agent/sandbox.ts` with the `@agent-browser/eve/sandbox` helpers to bake the cost into the sandbox template instead. Configuration (domain allowlists, output limits, session naming) and per-tool overrides are covered in the [package README](https://github.com/vercel-labs/agent-browser/blob/main/packages/@agent-browser/eve/README.md), and the [eve example](https://github.com/vercel-labs/agent-browser/blob/main/examples/eve) is a complete app with the extension mounted.

### Serverless (AWS Lambda)

import chromium from '@sparticuz/chromium';
import { execSync } from 'child\_process';

export async function handler() {
  const executablePath \= await chromium.executablePath();
  const result \= execSync(
    \`AGENT\_BROWSER\_EXECUTABLE\_PATH=${executablePath} agent-browser open https://example.com && agent-browser snapshot -i --json\`,
    { encoding: 'utf-8' }
  );
  return JSON.parse(result);
}

## Local Files

Open and interact with local files (PDFs, HTML, etc.) using `file://` URLs:

# Enable file access (required for JavaScript to access local files)
agent-browser --allow-file-access open file:///path/to/document.pdf
agent-browser --allow-file-access open file:///path/to/page.html

# Take screenshot of a local PDF
agent-browser --allow-file-access open file:///Users/me/report.pdf
agent-browser screenshot report.png

The `--allow-file-access` flag adds Chromium flags (`--allow-file-access-from-files`, `--allow-file-access`) that allow `file://` URLs to:

-   Load and render local files
-   Access other local files via JavaScript (XHR, fetch)
-   Load local resources (images, scripts, stylesheets)

**Note:** This flag only works with Chromium. For security, it's disabled by default.

## CDP Mode

Connect to an existing browser via Chrome DevTools Protocol:

# Start Chrome with: google-chrome --remote-debugging-port=9222

# Connect once, then run commands without --cdp
agent-browser connect 9222
agent-browser snapshot
agent-browser tab
agent-browser close

# Or pass --cdp on each command
agent-browser --cdp 9222 snapshot

# Connect to remote browser via WebSocket URL
agent-browser --cdp "wss://your-browser-service.com/cdp?token=..." snapshot

The `--cdp` flag accepts either:

-   A port number (e.g., `9222`) for local connections via `http://localhost:{port}`
-   A full WebSocket URL (e.g., `wss://...` or `ws://...`) for remote browser services

This enables control of:

-   Electron apps
-   Chrome/Chromium instances with remote debugging
-   WebView2 applications
-   Any browser exposing a CDP endpoint

### Auto-Connect

Use `--auto-connect` to automatically discover and connect to a running Chrome instance without specifying a port:

# Auto-discover running Chrome with remote debugging
agent-browser --auto-connect open example.com
agent-browser --auto-connect snapshot

# Or via environment variable
AGENT\_BROWSER\_AUTO\_CONNECT=1 agent-browser snapshot

Auto-connect discovers Chrome by:

1.  Reading Chrome's `DevToolsActivePort` file from the default user data directory
2.  Falling back to probing common debugging ports (9222, 9229)
3.  If HTTP-based discovery (`/json/version`, `/json/list`) fails, falling back to a direct WebSocket connection

This is useful when:

-   Chrome 144+ has remote debugging enabled via `chrome://inspect/#remote-debugging` (which uses a dynamic port)
-   You want a zero-configuration connection to your existing browser
-   You don't want to track which port Chrome is using

## Streaming (Browser Preview)

Stream the browser viewport via WebSocket for live preview or "pair browsing" where a human can watch and interact alongside an AI agent.

### Streaming

Every session automatically starts a WebSocket stream server on an OS-assigned port. Use `stream status` to see the bound port and connection state:

agent-browser stream status

To bind to a specific port, set `AGENT_BROWSER_STREAM_PORT`:

AGENT\_BROWSER\_STREAM\_PORT=9223 agent-browser open example.com

Frame encoding is daemon-wide:

Variable

Default

Description

`AGENT_BROWSER_STREAM_QUALITY`

`80`

JPEG quality, 0 to 100

`AGENT_BROWSER_STREAM_MAX_WIDTH`

the viewport

Caps frame width in pixels

`AGENT_BROWSER_STREAM_MAX_HEIGHT`

the viewport

Caps frame height in pixels

Width and height cap the encoded frame and leave the page size alone, so a portrait or HiDPI viewport keeps its resolution unless you cap it. The live stream requests jpeg. An explicit `screencast_start` reconfigures the same screencast, so a client can see the format change mid-stream. On a busy page at 1280x720, quality 80 costs about 54 KB per frame, quality 20 about 25 KB, and quality 20 at 640x360 about 9 KB.

# Cheaper frames for a constrained link
AGENT\_BROWSER\_STREAM\_QUALITY=20 \\
AGENT\_BROWSER\_STREAM\_MAX\_WIDTH=640 \\
AGENT\_BROWSER\_STREAM\_MAX\_HEIGHT=360 \\
agent-browser open example.com

You can also manage streaming at runtime with `stream enable`, `stream disable`, and `stream status`:

agent-browser stream enable --port 9223   # Re-enable on a specific port
agent-browser stream disable              # Stop streaming for the session

The WebSocket server streams the browser viewport and accepts input events.

### WebSocket Protocol

Connect to `ws://localhost:9223` to receive frames and send input:

**Receive frames:**

{
  "type": "frame",
  "seq": 41,
  "data": "<base64-encoded-jpeg>",
  "metadata": {
    "deviceWidth": 1280,
    "deviceHeight": 720,
    "pageScaleFactor": 1,
    "offsetTop": 0,
    "scrollOffsetX": 0,
    "scrollOffsetY": 0,
    "timestamp": 1785038682238
  }
}

`seq` is a monotonic frame id, echoed back in an `ack` message under ack pacing. `metadata.timestamp` is the capture time in epoch milliseconds, so a client can tell how old a frame is by the time it draws it.

**Send mouse events:**

{
  "type": "input\_mouse",
  "eventType": "mousePressed",
  "x": 100,
  "y": 200,
  "button": "left",
  "clickCount": 1
}

**Send keyboard events:**

{
  "type": "input\_keyboard",
  "eventType": "keyDown",
  "key": "Enter",
  "code": "Enter"
}

**Send touch events:**

{
  "type": "input\_touch",
  "eventType": "touchStart",
  "touchPoints": \[{ "x": 100, "y": 200 }\]
}

**Cap the frame rate (per client):**

{
  "type": "config",
  "maxFps": 10
}

Frames are delivered latest-first: the server picks the newest frame at send time, so frames produced while an earlier one is still being written are skipped rather than queued. `maxFps` (1 to 120, `0` = uncapped) limits delivery for that client only. A client that sends `{"type":"config","pacing":"ack"}` receives one frame at a time and acknowledges it with `{"type":"ack","seq":N}`, so nothing stale reaches the socket even if that client stalls; in the default push pacing, frames already handed to the transport are still delivered in order. Both settings can also be declared on the URL (`ws://127.0.0.1:<port>/?pacing=ack&maxFps=10`), which is the only way to cover the connection's opening frame. Input events are read on a dedicated task per connection, so clicks and keystrokes dispatch immediately even while frames are mid-write to a slow client. They are sent to the browser without waiting for its reply, so a click stays responsive behind a burst of mouse moves, and ordering is preserved.

## Architecture

agent-browser uses a client-daemon architecture:

1.  **Rust CLI** - Parses commands, communicates with daemon
2.  **Rust Daemon** - Pure Rust daemon using direct CDP, no Node.js required

The daemon starts automatically on first command and persists between commands for fast subsequent operations. After **1 hour** with no commands or dashboard input it saves configured restore state, closes the browser, and exits, so an integration that dies without calling `close` cannot leak the daemon and its browser indefinitely; the next command starts a fresh daemon and configured state restore works as usual. A session without `--restore` or another restore key does not save browser state, so its transient state and open tabs are discarded at shutdown. Set `--idle-timeout` to a duration such as `30s`, `5m`, or `1h`, or set `AGENT_BROWSER_IDLE_TIMEOUT_MS` to a value in milliseconds. Use `0` to disable idle shutdown entirely. The default never closes a headed browser, including Safari and iOS WebDriver sessions, or a user-attached browser because those may be in direct human use. Provider-owned cloud browsers remain eligible for cleanup. An explicitly set timeout applies to every browser.

**Browser Engine:** Uses Chrome (from Chrome for Testing) by default. The `--engine` flag selects between `chrome` and `lightpanda`. Supported browsers: Chromium/Chrome (via CDP) and Safari (via WebDriver for iOS).

## Platforms

Platform

Binary

macOS ARM64

Native Rust

macOS x64

Native Rust

Linux ARM64

Native Rust

Linux x64

Native Rust

Windows x64

Native Rust

## Usage with AI Agents

### Just ask the agent

The simplest approach is to tell your agent to use it:

```
Use agent-browser to test the login flow. Run agent-browser --help to see available commands.
```

The `--help` output is comprehensive and most agents can figure it out from there.

### AI Coding Assistants (recommended)

Add the skill to your AI coding assistant for richer context:

npx skills add vercel-labs/agent-browser

This works with Claude Code, Codex, Cursor, Gemini CLI, GitHub Copilot, Goose, OpenCode, and Windsurf. The skill is fetched from the repository, so it stays up to date automatically. Do not copy `SKILL.md` from `node_modules` as it will become stale.

### Claude Code

Install as a Claude Code skill:

npx skills add vercel-labs/agent-browser

This adds a thin discovery stub at `.claude/skills/agent-browser/SKILL.md`. The stub is intentionally minimal — it points Claude Code at `agent-browser skills get core` to load the actual workflow content at runtime. This way the instructions always match the installed CLI version instead of going stale between releases.

### AGENTS.md / CLAUDE.md

For more consistent results, add to your project or global instructions file:

\## Browser Automation

Use \`agent-browser\` for web automation. Run \`agent-browser --help\` for all commands.

Core workflow:

1. \`agent-browser open <url>\` - Navigate to page
2. \`agent-browser snapshot -i\` - Get interactive elements with refs (@e1, @e2)
3. \`agent-browser click @e1\` / \`fill @e2 "text"\` - Interact using refs
4. Re-snapshot after page changes

## Integrations

### iOS Simulator

Control real Mobile Safari in the iOS Simulator for authentic mobile web testing. Requires macOS with Xcode.

**Setup:**

# Install Appium and XCUITest driver
npm install -g appium
appium driver install xcuitest

**Usage:**

# List available iOS simulators
agent-browser device list

# Launch Safari on a specific device
agent-browser -p ios --device "iPhone 16 Pro" open https://example.com

# Same commands as desktop
agent-browser -p ios snapshot -i
agent-browser -p ios tap @e1
agent-browser -p ios fill @e2 "text"
agent-browser -p ios screenshot mobile.png

# Mobile-specific commands
agent-browser -p ios swipe up
agent-browser -p ios swipe down 500

# Close session
agent-browser -p ios close

Or use environment variables:

export AGENT\_BROWSER\_PROVIDER=ios
export AGENT\_BROWSER\_IOS\_DEVICE="iPhone 16 Pro"
agent-browser open https://example.com

Variable

Description

`AGENT_BROWSER_PROVIDER`

Set to `ios` to enable iOS mode

`AGENT_BROWSER_IOS_DEVICE`

Device name (e.g., "iPhone 16 Pro", "iPad Pro")

`AGENT_BROWSER_IOS_UDID`

Device UDID (alternative to device name)

**Supported devices:** All iOS Simulators available in Xcode (iPhones, iPads), plus real iOS devices.

**Note:** The iOS provider boots the simulator, starts Appium, and controls Safari. First launch takes ~30-60 seconds; subsequent commands are fast.

#### Real Device Support

Appium also supports real iOS devices connected via USB. This requires additional one-time setup:

**1\. Get your device UDID:**

xcrun xctrace list devices
# or
system\_profiler SPUSBDataType | grep -A 5 "iPhone\\|iPad"

**2\. Sign WebDriverAgent (one-time):**

# Open the WebDriverAgent Xcode project
cd ~/.appium/node\_modules/appium-xcuitest-driver/node\_modules/appium-webdriveragent
open WebDriverAgent.xcodeproj

In Xcode:

-   Select the `WebDriverAgentRunner` target
-   Go to Signing & Capabilities
-   Select your Team (requires Apple Developer account, free tier works)
-   Let Xcode manage signing automatically

**3\. Use with agent-browser:**

# Connect device via USB, then:
agent-browser -p ios --device "<DEVICE\_UDID>" open https://example.com

# Or use the device name if unique
agent-browser -p ios --device "John's iPhone" open https://example.com

**Real device notes:**

-   First run installs WebDriverAgent to the device (may require Trust prompt)
-   Device must be unlocked and connected via USB
-   Slightly slower initial connection than simulator
-   Tests against real Safari performance and behavior

### Browserless

[Browserless](https://browserless.io/) provides cloud browser infrastructure with a Sessions API. Use it when running agent-browser in environments where a local browser isn't available.

To enable Browserless, use the `-p` flag:

export BROWSERLESS\_API\_KEY="your-api-token"
agent-browser -p browserless open https://example.com

Or use environment variables for CI/scripts:

export AGENT\_BROWSER\_PROVIDER=browserless
export BROWSERLESS\_API\_KEY="your-api-token"
agent-browser open https://example.com

Optional configuration via environment variables:

Variable

Description

Default

`BROWSERLESS_API_URL`

Base API URL (for custom regions or self-hosted)

`https://production-sfo.browserless.io`

`BROWSERLESS_BROWSER_TYPE`

Type of browser to use (chromium or chrome)

chromium

`BROWSERLESS_TTL`

Session TTL in milliseconds

`300000`

`BROWSERLESS_STEALTH`

Enable stealth mode (`true`/`false`)

`true`

When enabled, agent-browser connects to a Browserless cloud session instead of launching a local browser. All commands work identically.

Get your API token from the [Browserless Dashboard](https://browserless.io/).

### Browserbase

[Browserbase](https://browserbase.com/) provides remote browser infrastructure to make deployment of agentic browsing agents easy. Use it when running the agent-browser CLI in an environment where a local browser isn't feasible.

To enable Browserbase, use the `-p` flag:

export BROWSERBASE\_API\_KEY="your-api-key"
agent-browser -p browserbase open https://example.com

Or use environment variables for CI/scripts:

export AGENT\_BROWSER\_PROVIDER=browserbase
export BROWSERBASE\_API\_KEY="your-api-key"
agent-browser open https://example.com

When enabled, agent-browser connects to a Browserbase session instead of launching a local browser. All commands work identically.

Get your API key from the [Browserbase Dashboard](https://browserbase.com/overview).

### Browser Use

[Browser Use](https://browser-use.com/) provides cloud browser infrastructure for AI agents. Use it when running agent-browser in environments where a local browser isn't available (serverless, CI/CD, etc.).

To enable Browser Use, use the `-p` flag:

export BROWSER\_USE\_API\_KEY="your-api-key"
agent-browser -p browseruse open https://example.com

Or use environment variables for CI/scripts:

export AGENT\_BROWSER\_PROVIDER=browseruse
export BROWSER\_USE\_API\_KEY="your-api-key"
agent-browser open https://example.com

When enabled, agent-browser connects to a Browser Use cloud session instead of launching a local browser. All commands work identically.

Get your API key from the [Browser Use Cloud Dashboard](https://cloud.browser-use.com/settings?tab=api-keys). Free credits are available to get started, with pay-as-you-go pricing after.

### Kernel

[Kernel](https://www.kernel.sh/) provides cloud browser infrastructure for AI agents with features like stealth mode and persistent profiles.

To enable Kernel, use the `-p` flag:

export KERNEL\_API\_KEY="your-api-key"
agent-browser -p kernel open https://example.com

Or use environment variables for CI/scripts:

export AGENT\_BROWSER\_PROVIDER=kernel
export KERNEL\_API\_KEY="your-api-key"
agent-browser open https://example.com

Optional configuration via environment variables:

Variable

Description

Default

`KERNEL_HEADLESS`

Run browser in headless mode (`true`/`false`)

`true`

`KERNEL_STEALTH`

Enable stealth mode to avoid bot detection (`true`/`false`)

`false`

`KERNEL_TIMEOUT_SECONDS`

Session timeout in seconds

`300`

`KERNEL_PROFILE_NAME`

Browser profile name for persistent cookies/logins (created if it doesn't exist)

(none)

When enabled, agent-browser connects to a Kernel cloud session instead of launching a local browser. All commands work identically.

**Profile Persistence:** When `KERNEL_PROFILE_NAME` is set, the profile will be created if it doesn't already exist. Cookies, logins, and session data are automatically saved back to the profile when the browser session ends, making them available for future sessions.

Get your API key from the [Kernel Dashboard](https://dashboard.onkernel.com/).

### AgentCore

[AWS Bedrock AgentCore](https://aws.amazon.com/bedrock/agentcore/) provides cloud browser sessions with SigV4 authentication.

To enable AgentCore, use the `-p` flag:

agent-browser -p agentcore open https://example.com

Or use environment variables for CI/scripts:

export AGENT\_BROWSER\_PROVIDER=agentcore
agent-browser open https://example.com

Credentials are automatically resolved from environment variables (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) or the AWS CLI (`aws configure export-credentials`), which supports SSO, profiles, and IAM roles.

Optional configuration via environment variables:

Variable

Description

Default

`AGENTCORE_REGION`

AWS region for the AgentCore endpoint

`us-east-1`

`AGENTCORE_BROWSER_ID`

Browser identifier

`aws.browser.v1`

`AGENTCORE_PROFILE_ID`

Browser profile for persistent state (cookies, localStorage)

(none)

`AGENTCORE_SESSION_TIMEOUT`

Session timeout in seconds

`3600`

`AWS_PROFILE`

AWS CLI profile for credential resolution

`default`

**Browser profiles:** When `AGENTCORE_PROFILE_ID` is set, browser state (cookies, localStorage) is persisted across sessions automatically.

When enabled, agent-browser connects to an AgentCore cloud browser session instead of launching a local browser. All commands work identically.

## License

Apache-2.0