A browsing task and the agent-browser CLI.
The task done in fewer tool calls: forms, logins, dialogs, uploads and paginated data.
Measured result
Taken from the skill's README.
On 11 unseen tasks the trained skill was about 14% cheaper per task than no skill (29.4k versus 34.2k weighted tokens), with no change in accuracy: 22 of 22 correct in every condition. Two of the three edit sets we tried made things worse and were rejected by the validation gate.
What it does
A skill that teaches an AI assistant to drive a real browser with the agent-browser command line tool in as few tool calls as possible. It covers clicking, filling forms, logging in, hover menus, uploads, drag and drop, JavaScript dialogs, iframes, infinite scroll and pulling data out of rendered or paginated pages.
We did not write it by hand. It was trained against a set of scored tasks, and only the edits that measurably helped were kept. The method is a scaled-down version of Microsoft's SkillOpt paper (arXiv 2605.23904). The test harness, every candidate version and all scores are in the research repo: https://github.com/proskillpacks/skillopt-agent-browse
How it was tested
What you can check yourself. The example on this page, and the files themselves: it is free, so you can open and read all of it first. There are no ratings yet, because nobody has reviewed it yet.
How we tested it.
Trained and scored with Claude Opus 5.5 through the claude CLI and the agent-browser CLI. On 11 unseen tasks (2 runs each) it used 29.4k weighted tokens per task against 34.2k with no skill, with 22 of 22 correct in every condition. Two of three proposed edit sets made things worse and were rejected by the validation gate. One student model and a small task set; the Setup section was added by hand and was not gated.
Scored on 32 auto-checked tasks, with every candidate version and every run published in the research repo. We did not score it with any other model or assistant.
How we test, what we do not do, and which assistants we ran it on.
Limits
- One student model (Claude Opus 5.5). Other models are untested.
- A small task set (32 tasks), mostly public practice sites. No real logins, no bot detection.
- The student already solved every task without a skill, so the gain is cost, not accuracy.
- The Setup section of SKILL.md was added by hand and did not go through the validation gate.
Works with
Tested with Claude Opus 5.5 through the claude CLI and the agent-browser CLI. The format is the open SKILL.md format, but we have not tested other assistants or models. You need the agent-browser CLI (https://github.com/vercel-labs/agent-browser) and an assistant that can run shell commands.
For people who use AI assistants
- Install the agent-browser CLI from https://github.com/vercel-labs/agent-browser and check that
agent-browser --helpruns. - Copy the
agent-browsefolder from GitHub (https://github.com/proskillpacks/skills/tree/main/skills/developers/agent-browse) into.agents/skills/in your project, or into your tool's own skills folder (for Claude Code:~/.claude/skills/).npx skills add proskillpacks/skillsalso works. - Ask for a browser task in plain words, for example: "Open the page, find the price table and give me the rows." The skill loads by itself. There is no paste-in prompt for this one, because it needs a shell.
Get it
MIT licence. Use it, change it, share it.
Related
- Accessibility Quick Audit (Free): Give it a page's HTML or components. Get accessibility issues ranked by impact, with the WCAG criterion, the code and a fix. Free, MIT-licensed.
- Changelog from Commits (Free): Give it a commit range or a pasted git log. Get a user-facing changelog with the commit behind every line, and the vague commits flagged.
- PR Description and Review Prep (Free): Run it on your branch. Get a PR title, summary, risk notes and test plan from the diff, plus the questions a reviewer will ask. Free, MIT-licensed.
- More for developers: Developers
Report a problem with this skill (GitHub, or write to kaiventura.founder@gmail.com).