Study / We re-tested 69 skills on harder inputs. 16 failed. Here is what broke.

We re-tested 69 skills on harder inputs. 16 failed. Here is what broke.

We gave 69 skills a second run on a deliberately hard input. 16 had a defect, mostly unsure input restated as fact. What broke, what we changed, and the limits.

Pro Skill Packs, 2026-10-04. Every number below comes from data.csv (one row per defect) and summary.json. The reading of each run is ours; there was no grader and no second model. The saved runs sit in each skill's tests folder in our workspace, and the quoted lines below are copied from them.

What we did

Each skill had already passed one run on a real public input. Today we gave all 69 of them (paid, about to be released, and free) a second run on an input that was deliberately harder. A single run had missed things, so we wanted to see how many.

"Harder" meant one of these, chosen per skill:

  • Unsure input. Notes full of "I think", "maybe", "I guess", or a question that someone might have answered.
  • Contradictions. A brief that asks for cheap and premium, or a listing whose title and text disagree.
  • Thin or blocked sources. A product with almost no description, a store that answers 403, a dead link, a page that is only a menu.
  • A request to overclaim. "Make it impressive and mention the results I achieved" when there are no results.
  • Messy data. A CSV with duplicates, a month 13 and a 9,999,999 amount.

We read each output ourselves against the input, looking for four things: something the input marked as unsure stated as fact, a claim the input does not support, a number in a summary that does not match the list under it, and an instruction or rule the skill ignored. All inputs were public pages and text we wrote. No customer data exists yet, and none was used.

Result

  • 69 skills tested. 53 passed the hard run. 16 had a defect. About one in four.
  • All 16 were fixed with one or two sentences in the skill's instructions (or, in one case, a change to its script) and re-run on the same input. All 16 passed the re-run.
  • The 69 were tested in five groups: 25 paid skills released 5 to 9 Oct, 5 released 7 Oct, 8 released 10 and 11 Oct, 11 host and free skills, and 20 live paid and free skills.
What brokeDefects
Unsure input restated as fact7
Invented or unsupported claim or benefit3
Miscount or arithmetic slip in a summary2
Conflict in the brief not named2
Script misread a compressed response1
Wrote a file it should not have1
Total16

The biggest group is the one we were least worried about: the skill noticed the input was unsure, marked it in one place, and then stated it as fact in the text meant for someone else.

What broke, in the skills' own words

These are copied from the saved runs, before the fix and after it.

Unsure input became a subject line. A newsletter skill was given notes where a postage change was unconfirmed and the link was dead.

  • Before: subject line 1 was "Postage prices are going up".
  • After: "Is postage going up this April?" and a note that the postage item stays marked as unconfirmed.

A hedged house rule became a rule. A host wrote "Smoking outside I guess".

  • Before: the paste-ready rules said "Smoking is outside only."
  • After: "Smoking, vaping, e-cigarettes | Not settled. Please confirm."

A hedged distance became a title. A listing said "5 min from the beach (maybe 10 if the tide is in)".

  • Before: title option 1 was "Studio 5 min from the beach, queen bed".
  • After: "Tiny home studio near the beach", with the hedge kept in the description.

A summary that did not add up. A claim checker was given a draft full of boasts.

  • Before: "There are 9 claims, and 8 are high risk." The table under it had 7 High and 2 Medium.
  • After: "There are 8 claims, 7 high risk and 1 medium." The two runs also disagreed on how many claims the draft held, which is why we read the counts in the saved run and not from memory.

A read-only skill that wrote a file. A CSV profiler is meant to leave the data alone.

  • Before: "I left orders.csv untouched and wrote the cleaned version to orders_clean.csv. Our csv-profile skill is read-only, so I kept the original as it was." The summary contradicts the action.
  • After: "I didn't change the file. The profiling skill I used is read-only, so I've listed each fix below for you to apply or approve."

What we changed

  1. Two real runs per skill from now on: one normal, one deliberately hard (thin, messy, contradictory or blocked). We read both, fix once and re-run the fixed one before a skill ships. A single run is not enough.
  2. One standard rule in 45 of our skills that write text for other people: "Anything the input marks as unsure, unconfirmed or missing stays marked (for example [CONFIRM: ...]) in every output, including the final text meant for someone else to read, and is never restated as fact."
  3. The free checker warns about it. Skill Portability Check v1.1.0 has a new warning (SC017) when a skill that writes text for others never mentions unsure or unconfirmed input.
  4. The per-skill fixes in data.csv, each one or two sentences, already in the skills.

Limits

  • Same model family. All runs used Claude Sonnet. A different model could fail in other places. We have not run these inputs on other assistants.
  • Our own reading. We judged the outputs ourselves, with no second reader. We checked quotes and arithmetic against the inputs, but a defect we did not look for would not be in this table.
  • One hard input per skill. One harder run found 16. A third run might find more. Passing means "passed this input", not "is correct".
  • Public and invented inputs only. Several inputs (a charity brief, a host's notes, a draft full of boasts) were written by us to be hard. They show behaviour on that kind of input, not how often real customers will hit it.
  • We grouped the counts by type after the fact. A defect that fits two types is counted once, under the one that matters most to a reader.

Files

data.csv has one row per defect with the fix. summary.json has the totals.

Files

data.csv summary.json