How we test

Exactly what we do before a skill is released, what we do not do, which assistants we ran it on, and how to report a problem.

Written 4 October 2026.

What we do for each skill

  • Two real runs. Before release, each skill is run twice on public input. The first is a normal case: a public web page, a job post, a CSV, code, commits, a status page. The second is deliberately hard: thin, messy, contradictory, blocked, or full of "I think" and "maybe". We read both outputs against the source ourselves, fix what is wrong, re-run the fixed skill, and ship. The first example appears on the skill's page under Before and after.
  • Why two. On 4 October 2026 we gave 69 skills a second, harder run: 16 had a defect that the first run had missed, most often something marked unsure in the input being stated as fact. All 16 were fixed and re-run. Read it: We re-tested 69 skills on harder inputs. 16 failed.
  • Assistants. The runs used Claude Sonnet. Some of the older packs were also checked on Claude Opus. The paste-in prompts were not run in ChatGPT or Gemini.
  • Script checks, no model. The Skill Portability Check script (the free tool at /tools/skill-check/) parses each SKILL.md like a strict YAML parser, flags vendor tool names, vendor paths and missing fallbacks, and warns when a skill that writes text for other people never says what to do with unsure input. A release check on every zip looks for a missing README or prompt, leftover placeholders, em dashes, private addresses and vendor-only wording. It records a sha256 of each shipped prompt in a file, so a changed prompt is visible.

Older packs

The Shopify and rental-host packs went through a stricter process: several tests per skill on real public stores or listings, scored by hand against a written rubric, and released only when every test scored 8 out of 10 or more. For the host pack an independent reviewer re-scored some tests, scored the earlier version lower than we did (as low as 6), and every defect it listed as required was fixed.

Invented test input

Some jobs have no public input, such as a customer complaint and your policy, an invoice, or your own resume bullets. For those we wrote the input ourselves, and the skill's README says so. Examples: the cold email checker, the angry customer reply check, the invoice checker (planted errors) and the resume rewriter's sample bullets. Web pages, job posts, CSVs and code in other tests are real.

What we do not do

  • We have not run every skill many times. Two runs show the skill can do the job on a normal and a hard input. They do not show it always will.
  • Newer skills have had no independent review, and a second run is still our own reading.
  • We have not tested every assistant. Different models give different output.
  • We cannot check your data. Always review the output before you use it.
  • As of 4 October 2026 we have no customer reports on the paid skills, so all of this comes from our own tests.

When something is wrong

Write to the address on the About page with what you gave the skill and what it returned. We reproduce it and change the skill, and the fix goes to everyone who has it. Changes we made after our own test runs, as examples: the spreadsheet skill now has to show a test table, the draft editor was told never to modify your files, and the resume skill no longer writes placeholders that assume a result.

Changes we make to the free skills after feedback are listed in the changelog.