get in touch

Why a 4B Specialist Beat Our Tested 9B Browser Agents

Jack Rudenko, CTO of MadAppGang
Jack Rudenko
CTO of MadAppGang

What we learned benchmarking local vision models for screen understanding, UI grounding, OCR and verified browser automation.

The obvious way to pick a model for a browser agent is to take the highest benchmark score or the biggest parameter count. Our local testing gave a more useful answer. The best system depends on what you need right now: understand a screen, find a control, read exact text, or recover safely when the first answer is wrong.

We tested 17 general vision models, 11 browser-agent candidates, three OCR specialists, two dedicated screen parsers and grounders, and 12 composed local-first systems. The hard suite used ten synthetic registration screenshots, 40 understanding prompts, 40 exact UI targets and 40 visible-text questions.

The winner on quality was Microsoft's Fara 1.5 27B. The surprise was how close Qwen 3.6 came, and how well a 4B specialist competed with models many times its size.

TL;DR: For a local browser agent, the biggest model isn't automatically the right pick. Fara 1.5 27B scored highest at 94.0%, Qwen 3.6 came within half a point with about 5.6 times faster grounding, and Clado BrowserOS Action 4B beat both 9B browser agents we tested from a 2.7 GiB package. The bigger win was verification: checking observable browser state took a local cascade from 73.3% to 95.2%.

Comic of two big robots labelled 9B pointing past a Create account button while a tiny 4B robot on a stepladder presses it, under a MadAppGang banner.

The results in one minute

  • Best standalone understanding: Fara 1.5 27B, 90.5%.
  • Best browser-agent composite: Fara 1.5 27B, 94.0%.
  • Best speed and quality challenger: Qwen 3.6 35B-A3B, 93.5% composite and 1.43-second grounding.
  • Fastest model above 80% understanding: Qwen3-VL 30B-A3B, 1.52 seconds.
  • Best small agent: Clado BrowserOS Action 4B, 88.7% composite from a 2.7 GiB package.
  • Best OCR specialist: GLM-OCR 0.9B, 95.0% at 2.31 seconds.
  • Best verified local cascade: LFM2.5 → Qwen3-VL → Fara, 95.2%.
  • Best hybrid cascade: the same local stack with cloud fallback, 100%, with 31 of 40 cases staying local.
Scatter plot comparing browser-agent composite score and grounding latency. Fara 27B has the highest score but the slowest grounding; Qwen 3.6 is close in score and much faster; Qwen3-VL is fastest among high-scoring large models; Clado 4B leads the compact models.

Browser-agent composite against grounding latency for every model with both understanding and grounding results. Higher and farther left is better. Bubble size shows the installed package size.

Understanding and clicking are different skills

Our understanding track asked models to extract visible values, identify validation states, interpret controls and recommend a safe next action. The grounding track asked a simpler question with no room for error: return a point inside a named UI element.

A model can be good at one and bad at the other. LFM2.5-VL 1.6B understood 73.3% of the registration rubric and answered quickly, but it grounded only 5.0% of targets. Holo 4B showed the reverse pattern, with a modest 75.7% understanding score and 92.5% grounding.

It reads the form well and then clicks somewhere near Tasmania. We've all had a colleague like that.

That's why our browser-agent score reports both parts and then takes their equal mean. The composite is useful for ranking. The separate columns tell you what will fail.

Model Understanding Grounding Composite Grounding latency
Fara 1.5 27B 90.5% 97.5% 94.0% 7.96s
Qwen 3.6 35B-A3B 89.5% 97.5% 93.5% 1.43s
Qwen3-VL 30B-A3B 83.3% 95.0% 89.2% 0.56s
Holo 3.1 35B-A3B 88.1% 90.0% 89.0% 2.96s
Clado BrowserOS Action 4B 82.4% 95.0% 88.7% 2.19s
Synthetic Northstar registration welcome screen with work email, password, terms, Google sign-in and Create account controls, one of the test screens.

One of the ten synthetic registration screens from the test set.

Qwen 3.6 vs Fara: the most interesting compromise

Fara still wins, but Qwen 3.6 trails by only half a composite point. Its grounding stage is about 5.6 times faster than Fara's on the Mac we tested on.

The architecture helps explain it. Qwen 3.6 has 35B total parameters but activates roughly 3B per token. So it gets the compute cost of a sparse model without becoming a tiny model on disk. Our installed 4-bit package was about 19 GiB.

Qwen 3.6 also showed why you should look at the raw output. It returned points in several fixed shapes, including x/y ranges and bounding boxes. Once we added model-specific parsing with regression tests, it scored 39 of 40 targets. The parser converted the model's own boxes to their centres. It never looked at the expected answer to pick a point.

Clado 4B is the small-model winner

Clado BrowserOS Action tied the best sub-10B understanding score at 82.4%, reached 95.0% grounding, and takes only 2.7 GiB on disk. On the combined score it beat the 9B versions of Holo and Fara.

So Clado is the most attractive choice when memory, download size or the number of agents per machine matters. Holo 4B is a reasonable alternative when grounding speed matters more than understanding the screen. GLM-4.6V-Flash 9B and MiniCPM-V 8B were also fast, strong screen readers. They weren't part of the grounding subset, so they don't get a browser-agent composite.

Horizontal bar chart of six small browser-agent models. Clado 4B ranks first at 88.7%, followed by Holo 9B, Fara 9B, Holo 4B, Fara 4B and LFM2.5-VL 1.6B.

Browser-agent composite for the tested models at or below 9B. Clado 4B leads at 88.7%.

OCR is better handled as its own route

The OCR track gave a clean result:

A sub-1B OCR specialist tied the best quality and was also the fastest. There's little reason to spend a 20 GiB general model on reading exact text off a screen when you know upfront that this is the task. I made the same argument about narrow models for decisions in my post on Jev and overengineered demos.

LocateAnything and OmniParser also belong on specialist routes. LocateAnything reached 95.0% grounding but was much slower than Qwen3-VL and Qwen 3.6 on this hardware. OmniParser recovered 77.5% of exact visible facts and produced reusable structured screen elements along the way. It's a parser, so you shouldn't compare it directly with chat models.

Synthetic registration completion screen showing a recovery code and workspace actions, the kind of exact text the OCR track had to read.

Valid JSON does not mean success

This was the biggest lesson.

We built a local-first cascade that tried LFM2.5, then Qwen3-VL, then Fara. A deployable response-contract router checked the JSON shape, the required keys and basic types. It was fast. It scored 73.3%.

The responses looked usable, but plausible mistakes in meaning passed straight through.

Valid JSON is a transport guarantee, not evidence that a browser action is correct.

Then we replaced the contract with deterministic postconditions. These stand in for DOM state, accessibility state, hit-testing or a verified screenshot after the action. With them, the local cascade reached 95.2%. Adding cloud fallback for unresolved fields reached 100%, and 31 of 40 cases stayed local. Cloud was called on the other nine.

The grounding cascade told the same story. Qwen3-VL followed by Clado reached 97.5% locally, at 0.68 seconds of replayed warm latency. Cloud fallback was needed for only one of 40 targets to reach 100%.

Read those 100% results carefully. They're a postcondition-verified upper bound. The verifier uses benchmark checks as a stand-in for reliable, observable browser state. That isn't evidence the models can judge their own answers perfectly, and we didn't create a real external account.

Two-panel comic. A robot holds up a scroll reading { valid: true } and says "Perfect JSON!", then a browser behind it shows a "Wrong field" error and an inspector robot asks "Did the page change?".

The local browser agent we would build

For a practical local-first browser agent:

  1. Use a fast specialist for any task you can identify in advance:
    • GLM-OCR for exact text.
    • Qwen3-VL for very fast grounding.
    • Clado when you need one small model for both understanding and grounding.
  2. Verify against observable browser state.
  3. Send only unresolved fields or targets to a stronger local model.
  4. Escalate to a cloud model only when local verification still fails.
  5. Ask for confirmation before credentials, payments, submissions, messages or anything you can't undo.

Qwen 3.6 is now the strongest candidate for the middle of that cascade. We added it to the standalone grounding track after the composed run, though. So a Qwen3-VL → Qwen 3.6 → Fara cascade is the next thing to measure, not a result we can claim today.

If you run agents with a real browser attached, the harness around the model matters as much as the model. And with new vision models landing every week, it helps when your tools hear about them before your users do.

Flow diagram of a local-first browser agent: a fast specialist (GLM-OCR, Qwen3-VL or Clado), verification against browser state, then a stronger local model and a cloud model only for unresolved cases, with confirmation required before credentials, payments, submissions and messages.

What this benchmark proves, and what it doesn't

This is a deterministic engineering benchmark, not a universal model ranking. It uses one Apple Silicon machine, fixed prompts, specific quantizations and one run per case. Latency excludes model loading. Other runtimes and hardware will give different numbers. It isn't sponsored by any of the model developers.

The registration flow is synthetic by design. No real credentials, account, payment, message or external submission was used. The test measures what browser automation needs without turning the evaluation into an uncontrolled live action.

That boundary is worth keeping. Local vision agents are now good enough to be useful.

The production problem is no longer just model accuracy. It's routing, observable verification and safe escalation.

Synthetic plan-selection screen with monthly and annual options and a selected Pro plan, from the synthetic registration flow used in the benchmark.

FAQ

Which local vision model is best for a browser agent?

In our benchmark Fara 1.5 27B had the highest browser-agent composite at 94.0%. Qwen 3.6 35B-A3B scored 93.5% and grounded about 5.6 times faster, so it's the better speed and quality trade on a single Mac.

Can a 4B model compete with bigger browser agents?

Yes, on this test. Clado BrowserOS Action 4B scored 88.7% on the composite, with 82.4% understanding and 95.0% grounding, from a 2.7 GiB package. That beat the 9B versions of Holo and Fara.

Should a browser agent use a separate OCR model?

If you know upfront that the task is reading exact text, yes. GLM-OCR 0.9B tied the best OCR score at 95.0% and was the fastest at 2.31 seconds.

Why isn't valid JSON enough to trust an agent's answer?

A JSON check only proves the shape of the answer. Our schema-only cascade scored 73.3%. Checking observable state, like DOM state or a screenshot after the action, raised the same local cascade to 95.2%.

Does the benchmark prove these models work on real websites?

No. It's one deterministic run per case on one Apple Silicon machine, with synthetic registration screens and no real accounts or submissions. Treat the numbers as measurements for this setup, not a universal ranking.

Sources

Model cards and repositories: Qwen 3.6 35B-A3B, Qwen3-VL 30B-A3B, Fara 1.5 27B, Holo 3.1, Clado BrowserOS Action, GLM-4.6V-Flash, LFM2.5-VL, GLM-OCR, Unlimited-OCR, DeepSeek-OCR, LocateAnything and OmniParser.

The joke in this post was created by AI. Please double-check before laughing.

X icon