Jern Cloud docs

The browser.

Every attempt's machine carries a headless browser. The agent opens the application it started on loopback, or a host the baseline allows, reads pages as text, clicks and types, and takes screenshots that are kept with the attempt as evidence for the reviewer. The agent never sees an image; you do.

Turning it on

A workspace policy profile has a switch, Let the agent drive a browser. On, the baseline the profile writes allows ten tools on the run-scoped jern MCP server: browser_open, browser_read, browser_click, browser_type, browser_press, browser_wait, browser_evaluate, browser_screenshot, browser_console, and browser_close. A repository's own .jern/baseline.json names them under allow the same way, as mcp__jern__browser_open and so on, or all at once as mcp__jern__browser_*. The grant moves the policy digest like any other, and without it the tools are not offered and the browser never starts.

The application the agent drives is its own to start: a dev server under shell_allow, run on 127.0.0.1. A shell command that starts a server in the background and returns, one plain command rather than a chain, is the shape that works; the demo repository keeps one as python3 web/serve.py.

What the agent can do

ToolWhat it does
browser_openOpens an http or https page and reads it back: title, address, the interactive elements with a CSS selector for each, and the visible text, up to 12,000 characters. One page at a time.
browser_readThe current page again, plus how many console errors are waiting.
browser_clickClicks the first visible match of a selector, then reads the page back.
browser_typeClicks a field, clears it, types, optionally presses Enter, then reads the page back.
browser_pressEnter, Tab, Escape, Backspace, the arrows, PageUp, PageDown, Home, End.
browser_waitWaits for an element to appear, or for a time, at most thirty seconds.
browser_evaluateRuns JavaScript in the page and returns its value as JSON, up to 8,000 characters.
browser_screenshotTakes a JPEG of the page as evidence. At most twenty per attempt, two megabytes each.
browser_consoleConsole messages, uncaught exceptions, failed requests, and responses of 400 and above since the last call.
browser_closeCloses the page; the browser stays ready.

Selectors come from what the page read back: an element's id when it has one, its name, or a short path. An action that finds no element, or a page that does not answer within thirty seconds, comes back as a refusal the agent can act on, not a hang.

Where it can reach

The browser runs inside the agent's confinement, in the same private network namespace as the agent's shell. It reaches the server the agent started on loopback directly. It reaches a public host only by name and only through the same proxy the agent's other connections use, where the baseline's network_allow is enforced and every connection is counted on the receipt. It resolves no name of its own, and an address outside loopback fails in the browser exactly as it fails in a shell. Chromium's own sandbox is off because the runner's holds it; the browser runs as the agent's identity under the same filters, with two renderer processes at most.

What it leaves behind

After the agent phase the runner uploads the screenshots and a log of every page opened, console message, exception, failed request, and error response, beside the trace. Jern Cloud checks the archive, encrypts it like a workspace snapshot, keeps it once per run, and writes a summary on the run: the pages with their titles, each screenshot with the page it was taken on, console errors, failed requests. An upload that fails costs the reviewer the pictures and nothing else; the attempt goes on.

On the session, the attempt carries the summary as a fact, browser: 2 pages, 2 screenshots, and a strip of thumbnails; each opens the full image, served decrypted to a member who can see the session. The receipt check on the pull request carries a Browser row: the pages and the hosts they were on, the screenshots kept, the console errors and failed requests. The agent's summary is expected to say what each screenshot shows, since it cannot see them.

What a session looks like

On the demo repository, a converter page showed 212 °F as 194.22 °C. Told to serve the page, reproduce the bug in the browser, take a screenshot, fix the file, verify in the browser, and take another, the agent did exactly that in fifteen tool calls: one command to start the server, an open and a click that showed the wrong value, a before-fix screenshot, a one-line edit, an open, a click, and a typed value that showed 100.00 °C and 0.00 °C, an after-fix screenshot, and a pull request whose description says what each screenshot shows. The receipt read 2 pages on 127.0.0.1:8080; 2 screenshots kept as evidence on Jern Cloud; 0 console errors, 0 failed requests.

Bounds

Not yet

A screenshot named by an acceptance criterion, so a criterion can say which page after which action must look a certain way; and the screenshots on the pull request itself rather than on Jern Cloud alone.