Home / Developers and data / Agent Browse

Agent Browse

A browsing skill for the agent-browser CLI, trained against scored tasks instead of hand-written. About 14% cheaper per task than no skill on 11 unseen tasks, same accuracy.

Trained and gated How we trained it

Free, MIT licence, hosted on GitHub. Install the whole free set with npx skills add proskillpacks/skills, or copy this skill's folder.

What you give it

A browsing task and the agent-browser CLI.

What you get

The task done in fewer tool calls: forms, logins, dialogs, uploads and paginated data.

Measured result

Taken from the skill's README.

On 11 unseen tasks the trained skill was about 14% cheaper per task than no skill (29.4k versus 34.2k weighted tokens), with no change in accuracy: 22 of 22 correct in every condition. Two of the three edit sets we tried made things worse and were rejected by the validation gate.

What it does

A skill that teaches an AI assistant to drive a real browser with the agent-browser command line tool in as few tool calls as possible. It covers clicking, filling forms, logging in, hover menus, uploads, drag and drop, JavaScript dialogs, iframes, infinite scroll and pulling data out of rendered or paginated pages.

We did not write it by hand. It was trained against a set of scored tasks, and only the edits that measurably helped were kept. The method is a scaled-down version of Microsoft's SkillOpt paper (arXiv 2605.23904). The test harness, every candidate version and all scores are in the research repo: https://github.com/proskillpacks/skillopt-agent-browse

How it was tested

What you can check yourself. The example on this page, and the files themselves: it is free, so you can open and read all of it first. There are no ratings yet, because nobody has reviewed it yet.

How we tested it.

Trained and scored with Claude Opus 5.5 through the claude CLI and the agent-browser CLI. On 11 unseen tasks (2 runs each) it used 29.4k weighted tokens per task against 34.2k with no skill, with 22 of 22 correct in every condition. Two of three proposed edit sets made things worse and were rejected by the validation gate. One student model and a small task set; the Setup section was added by hand and was not gated.

Scored on 32 auto-checked tasks, with every candidate version and every run published in the research repo. We did not score it with any other model or assistant.

How we test, what we do not do, and which assistants we ran it on.

Limits

  • One student model (Claude Opus 5.5). Other models are untested.
  • A small task set (32 tasks), mostly public practice sites. No real logins, no bot detection.
  • The student already solved every task without a skill, so the gain is cost, not accuracy.
  • The Setup section of SKILL.md was added by hand and did not go through the validation gate.

Works with

Tested with Claude Opus 5.5 through the claude CLI and the agent-browser CLI. The format is the open SKILL.md format, but we have not tested other assistants or models. You need the agent-browser CLI (https://github.com/vercel-labs/agent-browser) and an assistant that can run shell commands.

For people who use AI assistants

  1. Install the agent-browser CLI from https://github.com/vercel-labs/agent-browser and check that agent-browser --help runs.
  2. Copy the agent-browse folder from GitHub (https://github.com/proskillpacks/skills/tree/main/skills/developers/agent-browse) into .agents/skills/ in your project, or into your tool's own skills folder (for Claude Code: ~/.claude/skills/). npx skills add proskillpacks/skills also works.
  3. Ask for a browser task in plain words, for example: "Open the page, find the price table and give me the rows." The skill loads by itself. There is no paste-in prompt for this one, because it needs a shell.

Get it

MIT licence. Use it, change it, share it.

Related

Report a problem with this skill (GitHub, or write to kaiventura.founder@gmail.com).