Jern Cloud docs
The browser.
Every attempt's machine carries a headless browser. The agent opens the application it started on loopback, or a host the baseline allows, reads pages as text, clicks and types, and takes screenshots that are kept with the attempt as evidence for the reviewer. The agent never sees an image; you do.
- Getting started
- Sessions
- Policy
- Triggers
- Schedules
- Skills
- Environments
- Browser
- Memory
- Knowledge
- Receipts
- Spend
- Provider keys
- Settings
- MCP server
- Recipes
- FAQ and limits
Turning it on
A workspace policy profile has a switch, Let the agent drive a browser. On, the baseline the profile writes allows ten tools on the run-scoped jern MCP server: browser_open, browser_read, browser_click, browser_type, browser_press, browser_wait, browser_evaluate, browser_screenshot, browser_console, and browser_close. A repository's own .jern/baseline.json names them under allow the same way, as mcp__jern__browser_open and so on, or all at once as mcp__jern__browser_*. The grant moves the policy digest like any other, and without it the tools are not offered and the browser never starts.
The application the agent drives is its own to start: a dev server under shell_allow, run on 127.0.0.1. A shell command that starts a server in the background and returns, one plain command rather than a chain, is the shape that works; the demo repository keeps one as python3 web/serve.py.
What the agent can do
| Tool | What it does |
|---|---|
browser_open | Opens an http or https page and reads it back: title, address, the interactive elements with a CSS selector for each, and the visible text, up to 12,000 characters. One page at a time. |
browser_read | The current page again, plus how many console errors are waiting. |
browser_click | Clicks the first visible match of a selector, then reads the page back. |
browser_type | Clicks a field, clears it, types, optionally presses Enter, then reads the page back. |
browser_press | Enter, Tab, Escape, Backspace, the arrows, PageUp, PageDown, Home, End. |
browser_wait | Waits for an element to appear, or for a time, at most thirty seconds. |
browser_evaluate | Runs JavaScript in the page and returns its value as JSON, up to 8,000 characters. |
browser_screenshot | Takes a JPEG of the page as evidence. At most twenty per attempt, two megabytes each. |
browser_console | Console messages, uncaught exceptions, failed requests, and responses of 400 and above since the last call. |
browser_close | Closes the page; the browser stays ready. |
Selectors come from what the page read back: an element's id when it has one, its name, or a short path. An action that finds no element, or a page that does not answer within thirty seconds, comes back as a refusal the agent can act on, not a hang.
Where it can reach
The browser runs inside the agent's confinement, in the same private network namespace as the agent's shell. It reaches the server the agent started on loopback directly. It reaches a public host only by name and only through the same proxy the agent's other connections use, where the baseline's network_allow is enforced and every connection is counted on the receipt. It resolves no name of its own, and an address outside loopback fails in the browser exactly as it fails in a shell. Chromium's own sandbox is off because the runner's holds it; the browser runs as the agent's identity under the same filters, with two renderer processes at most.
What it leaves behind
After the agent phase the runner uploads the screenshots and a log of every page opened, console message, exception, failed request, and error response, beside the trace. Jern Cloud checks the archive, encrypts it like a workspace snapshot, keeps it once per run, and writes a summary on the run: the pages with their titles, each screenshot with the page it was taken on, console errors, failed requests. An upload that fails costs the reviewer the pictures and nothing else; the attempt goes on.
On the session, the attempt carries the summary as a fact, browser: 2 pages, 2 screenshots, and a strip of thumbnails; each opens the full image, served decrypted to a member who can see the session. The receipt check on the pull request carries a Browser row: the pages and the hosts they were on, the screenshots kept, the console errors and failed requests. The agent's summary is expected to say what each screenshot shows, since it cannot see them.
What a session looks like
On the demo repository, a converter page showed 212 °F as 194.22 °C. Told to serve the page, reproduce the bug in the browser, take a screenshot, fix the file, verify in the browser, and take another, the agent did exactly that in fifteen tool calls: one command to start the server, an open and a click that showed the wrong value, a before-fix screenshot, a one-line edit, an open, a click, and a typed value that showed 100.00 °C and 0.00 °C, an after-fix screenshot, and a pull request whose description says what each screenshot shows. The receipt read 2 pages on 127.0.0.1:8080; 2 screenshots kept as evidence on Jern Cloud; 0 console errors, 0 failed requests.
Bounds
- One browser and one page at a time, 1280 by 800, headless.
- Text read back is bounded: 12,000 characters of text, 80 elements, 8,000 characters from
browser_evaluate. - Twenty screenshots per attempt, two megabytes each; full-page captures stop at 8,000 pixels.
- The browser ends with the attempt, and every process of the agent's is sealed after the run either way.
Not yet
A screenshot named by an acceptance criterion, so a criterion can say which page after which action must look a certain way; and the screenshots on the pull request itself rather than on Jern Cloud alone.