A browsing task and the agent-browser CLI.
A local model like Qwen3.8-27B stops guessing addresses and selectors, and gets the task done in fewer tries.
Good to know
Trained and gated How we trained it
Measured result, selection split
Taken from the research repo's README. Held-out test pending.
Qwen3.8-27B, driven by the pi coding agent, 8 selection tasks with 6 runs each (48 rollouts per skill). The held-out test is pending: those runs were still in progress when this was written, so there is no unseen-task result yet.
| Skill | Edits | Correct | Score | Tokens per task | Turns | Decision |
|---|---|---|---|---|---|---|
| No skill (24 rollouts) | 21/24 | 0.7887 | 68.7k | 13.9 | ||
s1 (Opus-trained) | 47/48 | 0.9177 | 42.3k | 9.2 | start | |
q1 | 4 | 43/48 | 0.8367 | 42.0k | 8.3 | rejected |
q2 | q1 with one edit swapped | 48/48 | 0.9454 | 36.4k | 8.5 | accepted |
q3 | 4 more | 47/48 | 0.9215 | 39.4k | 8.1 | rejected |
q4 | 2 of the q3 edits | 44/48 | 0.8542 | 47.0k | 9.1 | see note |
Swipe the table to see more.
q4 note: three of its four failures were 15-minute timeouts on the 50-page catalogue crawl, after the practice site began answering 403 under four concurrent crawls. That is the benchmark throttling, not the skill.
The skill file is q2 plus the same hand-added Setup section as Agent Browse.
What it does
A variant of Agent Browse for Qwen3.8-27B and similar local models that tend to guess selectors and URLs. It teaches an assistant to drive a real browser with the agent-browser command line tool: clicking, filling forms, logging in, hover menus, uploads, JavaScript dialogs, paginated data.
It was trained the same way as Agent Browse: scored tasks, and only edits that measurably helped were kept. It starts from the skill trained for Claude Opus and adds what the Qwen runs showed: never invent a URL or selector you have not seen, use the CLI's own selector syntax, pass longer JavaScript on stdin, check a list against the page's own total before calling it complete, and do not close the browser early.
The test harness, every candidate version (accepted and rejected) and the run data are in the research repo: https://github.com/proskillpacks/skillopt-agent-browse (Part 2 of the write-up).
How it was tested
What you can check yourself. The example on this page, and the files themselves: it is free, so you can open and read all of it first. There are no ratings yet, because nobody has reviewed it yet.
How we tested it.
Trained and scored with Qwen3.8-27B on a local server, driven by the pi coding agent. On the 8 selection tasks (6 runs each) it got 48 of 48 correct at 36.4k tokens per task, against 21 of 24 at 68.7k with no skill and 47 of 48 at 42.3k with the Claude-trained Agent Browse. These are selection-split figures, the split the gate used to pick it. The held-out test is pending. Three of four tuning attempts were rejected by the gate.
Scored on the 8 selection tasks with 6 runs each, with every candidate version and run published in the research repo. The held-out test for this version is not finished. We did not score it with any other model.
How we test, what we do not do, and which assistants we ran it on.
Limits
- Held-out test pending: the figures are on the selection split, which the gate used to choose this skill.
- The cross-check on Claude Opus was not run.
- One local model, one quantisation, one harness (pi). Other models are untested.
- The selection runs were measured under four-way concurrency and site throttling.
- The Setup section of SKILL.md was added by hand and did not go through the validation gate.
Works with
Tested with Qwen3.8-27B on a local OpenAI-compatible server, driven by the pi coding agent, and the agent-browser CLI (https://github.com/vercel-labs/agent-browser). The format is the open SKILL.md format; other assistants and models are untested. You need an assistant that can run shell commands.
For people who use AI assistants
- Install the agent-browser CLI from https://github.com/vercel-labs/agent-browser and check that
agent-browser --helpruns. - Copy the
agent-browse-qwenfolder from GitHub (https://github.com/proskillpacks/skills/tree/main/skills/developers/agent-browse-qwen) into.agents/skills/in your project, or into your tool's own skills folder. - Ask for a browser task in plain words, for example: "Open the page, find the price table and give me the rows." The skill loads by itself. There is no paste-in prompt for this one, because it needs a shell.
Get it
MIT licence. Use it, change it, share it.
Related
- PR Description and Review Prep (Free): Run it on your branch. Get a PR title, summary, risk notes and test plan from the diff, plus the questions a reviewer will ask. Free, MIT-licensed.
- README First-Run Check (Free): Point it at a repo. It follows the README like a new user and checks every step against the repo: OK, BROKEN, MISSING or UNVERIFIED, with evidence and the README edits that fix it.
- Skill Portability Check (Free): A free checker for SKILL.md folders: strict-YAML trap, portable keys, vendor tool names, missing fallbacks, broken links. One Python file, exit code for CI.
- More for developers: Developers
Report a problem with this skill (GitHub, or write to kaiventura.founder@gmail.com).