docs

Browser Agent

The paid last-resort fallback for a page that needs clicking, a date picked, or a login walked through — the real start/status/send/stop shapes and the time limits.

bowmark.browser_agent is Bowmark's hosted browser agent — the fallback for when a page needs a form filled, a wizard driven, or a login walked through, and read.page (a plain fetch) can't do any of that. It is metered and billed, so reach for a typed provider or read.page first; use this when the task is genuinely an interaction, not as the next thing to try after a read fails.

One-off tasks only

It cannot run on a schedule or watch a page for changes — one task, a few minutes, then it stops. For a recurring job, call a capability or read.page from your own scheduled script (cron, Zapier, GitHub Actions) instead.

The shape: start, then poll from later calls

It runs through run() only (never session()), and it is asynchronous — start it in one run(), then check on it from separate, later run() calls. A single run() is capped at 90 seconds, so starting it and immediately polling in the same call hits that ceiling before the agent has had time to work.

// run 1 — start it, and show your user the watch link
const { result: started } = await run(
  `return bowmark.browser_agent.start({ task: "..." })`,
);
const { id, watchUrl } = started;

// run 2, 3, 4 … — one round trip each, from LATER calls
const { result: status } = await run(
  `return bowmark.browser_agent.status(${JSON.stringify(id)}, { waitMs: 45000 })`,
);
if (status.status === "needs_input") {
  // ask your user, then: bowmark.browser_agent.send(id, answer)
} else if (status.status === "idle") {
  await run(`return bowmark.browser_agent.stop(${JSON.stringify(id)})`);
}

start() takes task (a plain-language instruction) plus a few knobs (outputSchema, backend, model, maxCostUsd, proxyCountry, timeoutMs); status(id, { waitMs }) blocks up to waitMs (capped at 45s) waiting for progress; send(id, answer) answers a needs_input prompt in the same session; stop(id) closes it. list() shows every session you are holding.

The time limits

  • 90 seconds — the hard ceiling on any single run() call. This is why polling happens from separate runs, not a loop inside one.
  • A handful of poll cycles, roughly 1-3 minutes total is a normal task. The same step name repeating on back-to-back polls is not a stall — keep polling until idle or needs_input.
  • 5 minutes is the ceiling Bowmark itself enforces per turn — a turn stuck past that with no progress is cancelled automatically and reads failed. Set your own give-up deadline at 5 minutes or more; a shorter one abandons runs that were about to finish.
  • 20 idle minutes closes an open session automatically, but stop() it yourself — it costs money and counts against your concurrent-session limit until closed.

The login handoff

browser_agent has no stored credential and cannot do an unattended login — there is no secret field on start(), and interpolating bowmark.secret("…") into the task string sends the agent the literal placeholder text, never the password. When a task hits a login, a CAPTCHA, or any challenge only a person can solve, it returns needs_input with kind: "takeover" and waits for a person to act at watchUrl. This holds on every run, not just the first — nothing about a login is remembered between sessions, so a script that walked someone through signing in yesterday still needs a person today.

A site that blocks the browser outright (an Akamai/Cloudflare/PerimeterX/DataDome/Imperva "access denied" page) reports failed with an error starting blocked: , quickly — never needs_input, because nobody can click past a wall that already stopped the agent.

If the site has a typed provider instead, sign in once with providers.<site>.signIn() (using a stored bowmark.secret()) and make repeated calls inside that same session() — that reuses one cookie jar at typed-provider rates. read.page and browser_agent share no cookie jar with each other or with a provider's session, so neither can inherit the other's login.

What it costs

No flat per-run price — it is metered on the model turns the task takes plus the browser time it holds open, itemized on your billing dashboard as browser_agent.vendor. See Pricing § Hosted browser agents for the per-minute rates and a worked cost example. Your account can hold up to 3 open sessions at once and has a monthly spend cap (defaults to $2/run, configurable to $25).

The full reference

This page covers the shapes and limits; the worked examples with outputSchema, field-by-field write verification, and the "not a general fallback for a plain read" scoping are in Quickstart § Write actions and multi-step flows with browser_agent and Scripting § When nothing else works: the browser agent.

Reading this with an agent?

This page as plain text: /docs/browser-agent.md. The whole site as one file: https://bowmark.ai/llms-full.txt. Index of every page: https://bowmark.ai/llms.txt.