guest@scrappylabs:~/blog$ less fara-wrangled.md

Fara1.5, wrangled — a computer-use agent on our own GPU

Fara is a vision-only browser agent: it looks at a screenshot of the page — no DOM, no accessibility tree — thinks in text, and answers with one concrete action (click these pixels, type this, scroll, visit this URL). Our harness executes that action in a real browser, takes a fresh screenshot, and loops until the model says terminate with a final answer. The model runs on one GPU in our rack; the browser can run anywhere — including attached to your own logged-in session.

53sHN task, end-to-end ✓
~3s/stepwarm action latency
16.8GBpinned VRAM slice
$0cloud / quota spend
9B FP886.6 webvoyager tier
80scold start → first answer

01

Where each piece runs

The model server and the browser are decoupled — an OpenAI-compatible HTTP seam connects them. Screenshots go up, actions come back.

THE ORCHESTRATOR — any box THE GPU BOX — one card THE BROWSER — pick one fara-cli harness microsoft/fara + our patches Playwright controller executes actions, takes screenshots vLLM server Fara1.5-9B-FP8 · ~17GB slice Sandbox chromium fresh profile, no auth Your own logged-in browser CDP attach :9222 screenshot + history next action drives --cdp_url

02

The observe → think → act loop

You Harness Browser Fara 9B (GPU) task — "find X on this site" screenshot (1440x900 PNG) pixels goal + last 3 screenshots + action history thinks in text, thenemits ONE action chain-of-thought + tool_call playwright executesclick / type / scroll loop repeats ~3s per stepuntil done terminate (final answer) answer + full trajectory+ screenshots critical point? sign-in,payment, submit → PAUSESand asks the human first

03

Anatomy of one step

OBSERVE

Screenshot in

The harness captures the viewport at 1440×900 — the resolution Fara was trained on. Only the last 3 screenshots ride along; older ones drop out of context.

THINK + DECIDE

Model output

Reasoning first, then exactly one grounded action:

The top story is the first
row under the site header...
<tool_call>{"name": "computer_use",
 "arguments": {"action": "left_click",
  "x": 342, "y": 187}}</tool_call>
ACT

Playwright executes

The harness parses the XML, clicks those exact pixels in the real browser, waits for the page to settle, and the loop begins again with a fresh screenshot.

04

Two ways to give it a browser

default · safe

Sandbox launch

The harness spawns a fresh headless chromium — empty profile, no cookies, no auth. This is what the smoke test used. Can't touch your accounts even if it wanted to.

one command

fara-cli --task "..." \
  --endpoint_config endpoint_configs/vllm_config.json
our patch · authed sessions

CDP attach — drive your real session

Attach to your already-logged-in browser over the DevTools protocol. No stored credentials, no automation login to get flagged, and you watch every click live.

two commands

# 1 — relaunch your browser with a debug port
brave --remote-debugging-port=9222

# 2 — point Fara at it
fara-cli --task "..." --cdp_url http://localhost:9222 \
  --endpoint_config endpoint_configs/vllm_config.json

Why this matters: a huge share of real admin work lives in consoles and portals with no API. CDP-attach plus a local model means those flows can be automated with zero cloud spend — and no screenshot of an authed console ever leaves your network.

05

The action space

ActionWhat it doesGroup
left_click right_click double_click triple_clickMouse clicks at predicted (x, y) pixel coordinatesmouse
mouse_move left_click_dragCursor positioning and drag operationsmouse
type keyKeyboard input — text entry and key comboskeyboard
scroll hscrollVertical / horizontal page scrollingnavigate
visit_url history_back web_searchDirect navigation and searchnavigate
pause_and_memorize_factPins a fact so it survives even when old screenshots fall out of the windowused in our smoke run to hold the headline it readmemory
ask_user_questionStops and surfaces a question to the humanhuman gate
wait terminateSleep N seconds / end the task with the final answercontrol

06

Guardrails & the honest capability read

Task typeVerdictWhy
Read / navigate / report"list the users", "check this setting", "find the price" good today WebVoyager-tier work — the 9B scores 86.6% here. Our smoke run was this class: flawless.
Multi-step writes on dense admin UIswizards, multi-page config changes watch it Harder-benchmark tier (63.4% Online-Mind2Web). Misclicks compound. Run headful, human watching. Fara's critical-points training makes it pause before sign-ins, payments, and submits — but a watcher still catches the fumbles.
Anything with an APIofficial SDKs, CLIs, admin APIs… use the API Deterministic beats vision-clicking. Fara's lane is the long tail of UIs with no API.
High-stakes records (legal / medical) not this The model card marks legal/high-stakes domains out of scope — and that matches our own hard rules for records that matter.

Standing caution — prompt injection. Fara reads page content through its eyes. A malicious page can embed instructions aimed at the agent. On an authed session that's a real surface: keep CDP-attach runs to allow-listed destinations (admin.google, known consoles), the same defense-in-depth posture we apply to everything an agent reads.

model microsoft/Fara1.5-9B (MIT) · community FP8 quant · served by vLLM on one workstation-class GPU

harness github.com/microsoft/fara + our small patches: CDP attach to the active tab, SPA-safe navigation waits

trajectories every run saves screenshots + a full action log — a complete record of what it saw and did

ops nothing stays resident — the server wakes on demand and self-expires; every number above measured on our own hardware

ScrappyLabs · Bring your own AI. We keep it wrangled.