Case study. Published 5 October 2026.
The problem with hand-written skills
Most skills are written by hand. Someone writes advice that sounds sensible, tries it on a few examples, and ships it.
The trouble is that sensible advice can make results worse. Nothing tells you. The skill reads well, the three examples worked, and the extra cost or the new mistake never shows up. We wanted a way to find out.
Our method, in six plain steps
We treat a skill like something you tune. The model stays the same. Only the skill file changes, and only when a score says it should. The method is a small version of Microsoft's SkillOpt paper (arXiv 2605.23904).
- Tasks. A set of jobs where a program can check the answer. Split into train, selection and test.
- Rollout. The model does the train tasks with the current skill. We record every step and the score.
- Reflect. A second model reads the recordings and looks for patterns that repeat across tasks, not one-off fixes.
- Bounded edit. Only a few changes per round: 4, then 3, then 2. If a round fails, we know where to look.
- Gate. The new skill runs on the selection tasks, twice. Those tasks did not shape the edit.
- Keep or reject. The new skill stays only if its score is strictly higher. A rejected idea is written down so it is not proposed again. The test tasks are used once, at the end.
The case: the Agent Browse skill
Agent Browse teaches an AI assistant to browse the web with the agent-browser command line tool. We trained it on 32 browsing tasks (13 train, 8 selection, 11 test), using Claude Opus 5.5 as the model doing the work.
The model got every task right even with no skill. So accuracy gave the gate nothing to measure. We scored cost as well: a wrong answer is 0, a right answer is 0.7 plus up to 0.3 for using fewer tokens.
Result 1: training, on the selection tasks (8 tasks, 2 runs each)
| Step | Edits | Score | Tokens per task | Turns | Decision |
|---|---|---|---|---|---|
| 0, starting skill | 150 words | 0.9435 | 37.7k | 5.2 | start |
| 1 | 4 | 0.9687 | 20.8k | 3.0 | kept |
| 2 | 3 more | 0.9644 | 23.8k | 3.2 | rejected |
| 3 | 1 more | 0.9632 | 24.5k | 3.0 | rejected |
Every run was correct, so the score differences are all about cost.
Result 2: unseen test tasks (11 tasks, 2 runs each)
| Condition | Correct | Tokens per task | Turns per task |
|---|---|---|---|
| No skill | 22/22 | 34.2k | 5.0 |
| The tool vendor's own skill | 22/22 | 109.4k | 5.5 |
| Starting skill | 22/22 | 33.3k | 5.1 |
| Trained skill | 22/22 | 29.4k | 4.1 |
The vendor's skill is about 9k tokens and we put all of it in the prompt, which is harsher than loading it only when needed. The trained skill is about 700 tokens. With two runs per task, the test figures are a guide, not a precise number.
The four edits that helped
All four removed waste the model could not have known about.
- Put several commands in one call, and close the browser in the last one.
- Do not chain work to the page-open command with
&&. It can time out on a slow page that is still usable. - A short recipe for JavaScript dialogs and right-click. The model had been searching the help text for these.
- A short command list, so the model stops asking for help text at all.
The one edit that hurt
The rule was: "Never guess selectors on a page you have not seen." It sounds careful. It made the model take a page snapshot on tasks where guessing already worked, and that cost more than the failures it prevented. It came in the second set of edits. The gate rejected that set. Without the gate, we would have shipped it.
A second case: the starter example
To let anyone try this, we built a small loop that needs no browser: the starter. Its 18 toy tasks ask for a "house style" (date formats, invoice ids, file names) that the task never states. Here the model did not start at 100%, so accuracy itself could be the score. The model was Claude Haiku. One reflection step, one candidate.
| Skill given to the model | Selection (6 tasks, 2 runs) | Test (6 tasks, 2 runs) |
|---|---|---|
| No skill | 0 of 6 (one run) | 0 of 12 |
| Seed: two lines, no rules | 0 of 12 | 0 of 12 |
| Trained candidate (215 words, one step) | 10 of 12, kept | 10 of 12 |
The whole run was 84 model runs and one reflection call, about $1.50. A local 27B model (Qwen 3.8 via Ollama) scored 0 of 18 with no skill. The skill trained with Haiku then got 5 of 6 unseen tasks right on that local model, in one run.
Read it with care. The seed was too thin to show a climb: it went from nothing to a rule set in one step. The 10 of 12 comes from small numbers, and one task is worth 8 points. Four of the 24 unseen runs still slipped on small details. The tasks are toys. They show that the loop works, not that it works on your job.
What we learned
- Sensible is not the same as better.Two of three edit sets read well and made things worse. Only a score on unseen tasks told us.
- Small edits are easier to judge.With few changes per round, a rejection points at one idea.
- Good skills are short.The trained skill is about 700 tokens. The vendor's is about 9k and cost three times as much to use.
- Fix the test setup first.Our first score looked awful, and the cause was the setup, not the skill: browsers that crashed each other and a test site that failed under load.
Where each skill stands today
Scored on tasks, changed in small steps, kept only on a higher score. This is the only skill that has been through the full loop.
Run on real examples before release and checked against a written rubric. Not trained or gated. How we test.
Next to be trained and gated. Their output can be scored automatically. No dates.
Limits
- One skill, one model for the full case. We do not know yet whether the Agent Browse result holds for other models.
- The model started at 100% there, so the gains are about cost. The paper reports larger accuracy gains where models start lower.
- Our checks are automatic, so this only works for jobs a program can score. Many skills cannot be scored that way.
- The Setup section of the Agent Browse skill was added by hand, after problems in our test setup. It did not go through the gate.
Try it yourself
The starter is in the open repo. It needs only Python 3 and either the claude command or any local or hosted model with an OpenAI-style endpoint. It has the tasks, the scoring, the reflection step, the gate and the example run above, with a step-by-step tutorial.
git clone https://github.com/proskillpacks/skillopt-agent-browse
cd skillopt-agent-browse/starter
python3 run.py --skill none --split sel --tag baseline --model haiku