Guides / How we train skills: a case study

How we train skills: a case study

One skill of ours was trained like a model: scored on tasks, changed in small steps, kept only when the score went up. The method, the numbers, and what went wrong. Most of our other skills are not trained this way yet.

Case study. Published 5 October 2026.

The problem with hand-written skills

Most skills are written by hand. Someone writes advice that sounds sensible, tries it on a few examples, and ships it.

The trouble is that sensible advice can make results worse. Nothing tells you. The skill reads well, the three examples worked, and the extra cost or the new mistake never shows up. We wanted a way to find out.

Our method, in six plain steps

We treat a skill like something you tune. The model stays the same. Only the skill file changes, and only when a score says it should. The method is a small version of Microsoft's SkillOpt paper (arXiv 2605.23904).

The training loopSix steps in order: tasks, rollout, reflect, bounded edit, gate, keep or reject.1Taskswith a checker2Rolloutmodel does them3Reflectread the failures4Bounded edita few changes only5Gatescore unseen tasks6Keep or rejectkeep only if higherKept? The next round starts from the new skill.
  1. Tasks. A set of jobs where a program can check the answer. Split into train, selection and test.
  2. Rollout. The model does the train tasks with the current skill. We record every step and the score.
  3. Reflect. A second model reads the recordings and looks for patterns that repeat across tasks, not one-off fixes.
  4. Bounded edit. Only a few changes per round: 4, then 3, then 2. If a round fails, we know where to look.
  5. Gate. The new skill runs on the selection tasks, twice. Those tasks did not shape the edit.
  6. Keep or reject. The new skill stays only if its score is strictly higher. A rejected idea is written down so it is not proposed again. The test tasks are used once, at the end.

The case: the Agent Browse skill

Agent Browse teaches an AI assistant to browse the web with the agent-browser command line tool. We trained it on 32 browsing tasks (13 train, 8 selection, 11 test), using Claude Opus 5.5 as the model doing the work.

The model got every task right even with no skill. So accuracy gave the gate nothing to measure. We scored cost as well: a wrong answer is 0, a right answer is 0.7 plus up to 0.3 for using fewer tokens.

Result 1: training, on the selection tasks (8 tasks, 2 runs each)

StepEditsScoreTokens per taskTurnsDecision
0, starting skill150 words0.943537.7k5.2start
140.968720.8k3.0kept
23 more0.964423.8k3.2rejected
31 more0.963224.5k3.0rejected

Every run was correct, so the score differences are all about cost.

Result 2: unseen test tasks (11 tasks, 2 runs each)

ConditionCorrectTokens per taskTurns per task
No skill22/2234.2k5.0
The tool vendor's own skill22/22109.4k5.5
Starting skill22/2233.3k5.1
Trained skill22/2229.4k4.1

The vendor's skill is about 9k tokens and we put all of it in the prompt, which is harsher than loading it only when needed. The trained skill is about 700 tokens. With two runs per task, the test figures are a guide, not a precise number.

The four edits that helped

All four removed waste the model could not have known about.

  • Put several commands in one call, and close the browser in the last one.
  • Do not chain work to the page-open command with &&. It can time out on a slow page that is still usable.
  • A short recipe for JavaScript dialogs and right-click. The model had been searching the help text for these.
  • A short command list, so the model stops asking for help text at all.

The one edit that hurt

The rule was: "Never guess selectors on a page you have not seen." It sounds careful. It made the model take a page snapshot on tasks where guessing already worked, and that cost more than the failures it prevented. It came in the second set of edits. The gate rejected that set. Without the gate, we would have shipped it.

A second case: the starter example

To let anyone try this, we built a small loop that needs no browser: the starter. Its 18 toy tasks ask for a "house style" (date formats, invoice ids, file names) that the task never states. Here the model did not start at 100%, so accuracy itself could be the score. The model was Claude Haiku. One reflection step, one candidate.

Skill given to the modelSelection (6 tasks, 2 runs)Test (6 tasks, 2 runs)
No skill0 of 6 (one run)0 of 12
Seed: two lines, no rules0 of 120 of 12
Trained candidate (215 words, one step)10 of 12, kept10 of 12

The whole run was 84 model runs and one reflection call, about $1.50. A local 27B model (Qwen 3.8 via Ollama) scored 0 of 18 with no skill. The skill trained with Haiku then got 5 of 6 unseen tasks right on that local model, in one run.

Read it with care. The seed was too thin to show a climb: it went from nothing to a rule set in one step. The 10 of 12 comes from small numbers, and one task is worth 8 points. Four of the 24 unseen runs still slipped on small details. The tasks are toys. They show that the loop works, not that it works on your job.

What we learned

  • Sensible is not the same as better.Two of three edit sets read well and made things worse. Only a score on unseen tasks told us.
  • Small edits are easier to judge.With few changes per round, a rejection points at one idea.
  • Good skills are short.The trained skill is about 700 tokens. The vendor's is about 9k and cost three times as much to use.
  • Fix the test setup first.Our first score looked awful, and the cause was the setup, not the skill: browsers that crashed each other and a test site that failed under load.

Where each skill stands today

Trained and gatedAgent Browse

Scored on tasks, changed in small steps, kept only on a higher score. This is the only skill that has been through the full loop.

Tested by handEvery other skill

Run on real examples before release and checked against a written rubric. Not trained or gated. How we test.

PlannedInvoice Checker, and Spreadsheet Formula and Cleanup

Next to be trained and gated. Their output can be scored automatically. No dates.

Limits

  • One skill, one model for the full case. We do not know yet whether the Agent Browse result holds for other models.
  • The model started at 100% there, so the gains are about cost. The paper reports larger accuracy gains where models start lower.
  • Our checks are automatic, so this only works for jobs a program can score. Many skills cannot be scored that way.
  • The Setup section of the Agent Browse skill was added by hand, after problems in our test setup. It did not go through the gate.

Try it yourself

The starter is in the open repo. It needs only Python 3 and either the claude command or any local or hosted model with an OpenAI-style endpoint. It has the tasks, the scoring, the reflection step, the gate and the example run above, with a step-by-step tutorial.

git clone https://github.com/proskillpacks/skillopt-agent-browse
cd skillopt-agent-browse/starter
python3 run.py --skill none --split sel --tag baseline --model haiku

Open the starter The full research repo