Why a 4B Specialist Beat Our Tested 9B Browser Agents
What we learned benchmarking local vision models for screen understanding, UI grounding, OCR and verified browser automation.
The obvious way to pick a model for a browser agent is to take the highest benchmark score or the biggest parameter count. Our local testing gave a more useful answer. The best system depends on what you need right now: understand a screen, find a control, read exact text, or recover safely when the first answer is wrong.
We tested 17 general vision models, 11 browser-agent candidates, three OCR specialists, two dedicated screen parsers and grounders, and 12 composed local-first systems. The hard suite used ten synthetic registration screenshots, 40 understanding prompts, 40 exact UI targets and 40 visible-text questions.
The winner on quality was Microsoft's Fara 1.5 27B. The surprise was how close Qwen 3.6 came, and how well a 4B specialist competed with models many times its size.
TL;DR: For a local browser agent, the biggest model isn't automatically the right pick. Fara 1.5 27B scored highest at 94.0%, Qwen 3.6 came within half a point with about 5.6 times faster grounding, and Clado BrowserOS Action 4B beat both 9B browser agents we tested from a 2.7 GiB package. The bigger win was verification: checking observable browser state took a local cascade from 73.3% to 95.2%.

The results in one minute
- Best standalone understanding: Fara 1.5 27B, 90.5%.
- Best browser-agent composite: Fara 1.5 27B, 94.0%.
- Best speed and quality challenger: Qwen 3.6 35B-A3B, 93.5% composite and 1.43-second grounding.
- Fastest model above 80% understanding: Qwen3-VL 30B-A3B, 1.52 seconds.
- Best small agent: Clado BrowserOS Action 4B, 88.7% composite from a 2.7 GiB package.
- Best OCR specialist: GLM-OCR 0.9B, 95.0% at 2.31 seconds.
- Best verified local cascade: LFM2.5 → Qwen3-VL → Fara, 95.2%.
- Best hybrid cascade: the same local stack with cloud fallback, 100%, with 31 of 40 cases staying local.

Browser-agent composite against grounding latency for every model with both understanding and grounding results. Higher and farther left is better. Bubble size shows the installed package size.
Understanding and clicking are different skills
Our understanding track asked models to extract visible values, identify validation states, interpret controls and recommend a safe next action. The grounding track asked a simpler question with no room for error: return a point inside a named UI element.
A model can be good at one and bad at the other. LFM2.5-VL 1.6B understood 73.3% of the registration rubric and answered quickly, but it grounded only 5.0% of targets. Holo 4B showed the reverse pattern, with a modest 75.7% understanding score and 92.5% grounding.
It reads the form well and then clicks somewhere near Tasmania. We've all had a colleague like that.
That's why our browser-agent score reports both parts and then takes their equal mean. The composite is useful for ranking. The separate columns tell you what will fail.
| Model | Understanding | Grounding | Composite | Grounding latency |
|---|---|---|---|---|
| Fara 1.5 27B | 90.5% | 97.5% | 94.0% | 7.96s |
| Qwen 3.6 35B-A3B | 89.5% | 97.5% | 93.5% | 1.43s |
| Qwen3-VL 30B-A3B | 83.3% | 95.0% | 89.2% | 0.56s |
| Holo 3.1 35B-A3B | 88.1% | 90.0% | 89.0% | 2.96s |
| Clado BrowserOS Action 4B | 82.4% | 95.0% | 88.7% | 2.19s |

One of the ten synthetic registration screens from the test set.
Qwen 3.6 vs Fara: the most interesting compromise
Fara still wins, but Qwen 3.6 trails by only half a composite point. Its grounding stage is about 5.6 times faster than Fara's on the Mac we tested on.
The architecture helps explain it. Qwen 3.6 has 35B total parameters but activates roughly 3B per token. So it gets the compute cost of a sparse model without becoming a tiny model on disk. Our installed 4-bit package was about 19 GiB.
Qwen 3.6 also showed why you should look at the raw output. It returned points in several fixed shapes, including x/y ranges and bounding boxes. Once we added model-specific parsing with regression tests, it scored 39 of 40 targets. The parser converted the model's own boxes to their centres. It never looked at the expected answer to pick a point.
Clado 4B is the small-model winner
Clado BrowserOS Action tied the best sub-10B understanding score at 82.4%, reached 95.0% grounding, and takes only 2.7 GiB on disk. On the combined score it beat the 9B versions of Holo and Fara.
So Clado is the most attractive choice when memory, download size or the number of agents per machine matters. Holo 4B is a reasonable alternative when grounding speed matters more than understanding the screen. GLM-4.6V-Flash 9B and MiniCPM-V 8B were also fast, strong screen readers. They weren't part of the grounding subset, so they don't get a browser-agent composite.

Browser-agent composite for the tested models at or below 9B. Clado 4B leads at 88.7%.
OCR is better handled as its own route
The OCR track gave a clean result:
- GLM-OCR 0.9B: 95.0% at 2.31 seconds.
- Unlimited-OCR 3B: 95.0% at 5.52 seconds.
- DeepSeek-OCR 3B: 90.0% at 3.78 seconds.
A sub-1B OCR specialist tied the best quality and was also the fastest. There's little reason to spend a 20 GiB general model on reading exact text off a screen when you know upfront that this is the task. I made the same argument about narrow models for decisions in my post on Jev and overengineered demos.
LocateAnything and OmniParser also belong on specialist routes. LocateAnything reached 95.0% grounding but was much slower than Qwen3-VL and Qwen 3.6 on this hardware. OmniParser recovered 77.5% of exact visible facts and produced reusable structured screen elements along the way. It's a parser, so you shouldn't compare it directly with chat models.

Valid JSON does not mean success
This was the biggest lesson.
We built a local-first cascade that tried LFM2.5, then Qwen3-VL, then Fara. A deployable response-contract router checked the JSON shape, the required keys and basic types. It was fast. It scored 73.3%.
The responses looked usable, but plausible mistakes in meaning passed straight through.
Valid JSON is a transport guarantee, not evidence that a browser action is correct.
Then we replaced the contract with deterministic postconditions. These stand in for DOM state, accessibility state, hit-testing or a verified screenshot after the action. With them, the local cascade reached 95.2%. Adding cloud fallback for unresolved fields reached 100%, and 31 of 40 cases stayed local. Cloud was called on the other nine.
The grounding cascade told the same story. Qwen3-VL followed by Clado reached 97.5% locally, at 0.68 seconds of replayed warm latency. Cloud fallback was needed for only one of 40 targets to reach 100%.
Read those 100% results carefully. They're a postcondition-verified upper bound. The verifier uses benchmark checks as a stand-in for reliable, observable browser state. That isn't evidence the models can judge their own answers perfectly, and we didn't create a real external account.

The local browser agent we would build
For a practical local-first browser agent:
- Use a fast specialist for any task you can identify in advance:
- GLM-OCR for exact text.
- Qwen3-VL for very fast grounding.
- Clado when you need one small model for both understanding and grounding.
- Verify against observable browser state.
- Send only unresolved fields or targets to a stronger local model.
- Escalate to a cloud model only when local verification still fails.
- Ask for confirmation before credentials, payments, submissions, messages or anything you can't undo.
Qwen 3.6 is now the strongest candidate for the middle of that cascade. We added it to the standalone grounding track after the composed run, though. So a Qwen3-VL → Qwen 3.6 → Fara cascade is the next thing to measure, not a result we can claim today.
If you run agents with a real browser attached, the harness around the model matters as much as the model. And with new vision models landing every week, it helps when your tools hear about them before your users do.

What this benchmark proves, and what it doesn't
This is a deterministic engineering benchmark, not a universal model ranking. It uses one Apple Silicon machine, fixed prompts, specific quantizations and one run per case. Latency excludes model loading. Other runtimes and hardware will give different numbers. It isn't sponsored by any of the model developers.
The registration flow is synthetic by design. No real credentials, account, payment, message or external submission was used. The test measures what browser automation needs without turning the evaluation into an uncontrolled live action.
That boundary is worth keeping. Local vision agents are now good enough to be useful.
The production problem is no longer just model accuracy. It's routing, observable verification and safe escalation.

FAQ
Which local vision model is best for a browser agent?
In our benchmark Fara 1.5 27B had the highest browser-agent composite at 94.0%. Qwen 3.6 35B-A3B scored 93.5% and grounded about 5.6 times faster, so it's the better speed and quality trade on a single Mac.
Can a 4B model compete with bigger browser agents?
Yes, on this test. Clado BrowserOS Action 4B scored 88.7% on the composite, with 82.4% understanding and 95.0% grounding, from a 2.7 GiB package. That beat the 9B versions of Holo and Fara.
Should a browser agent use a separate OCR model?
If you know upfront that the task is reading exact text, yes. GLM-OCR 0.9B tied the best OCR score at 95.0% and was the fastest at 2.31 seconds.
Why isn't valid JSON enough to trust an agent's answer?
A JSON check only proves the shape of the answer. Our schema-only cascade scored 73.3%. Checking observable state, like DOM state or a screenshot after the action, raised the same local cascade to 95.2%.
Does the benchmark prove these models work on real websites?
No. It's one deterministic run per case on one Apple Silicon machine, with synthetic registration screens and no real accounts or submissions. Treat the numbers as measurements for this setup, not a universal ranking.
Sources
Model cards and repositories: Qwen 3.6 35B-A3B, Qwen3-VL 30B-A3B, Fara 1.5 27B, Holo 3.1, Clado BrowserOS Action, GLM-4.6V-Flash, LFM2.5-VL, GLM-OCR, Unlimited-OCR, DeepSeek-OCR, LocateAnything and OmniParser.
The joke in this post was created by AI. Please double-check before laughing.
