SkillOptimizer
Methodology

Every study on this site is a careful measurement of a narrow thing. This page lists what the studies do not establish, so that no result is read as more than it is.

Where the numbers come from

  • One machine, one harness, one catalog. Every figure is measured-local: run on one

machine, in the confined arena, on skills from the AirMarket catalog. Nothing here is measured on real users' runs, and nothing says how another marketplace's skills would behave.

  • One assistant model at a time. The assistant under test is Sonnet 5, and the smaller-model

study adds Haiku 4.5. A new model version can change any of these results, and none has been re-measured on one.

  • The arena is not an ordinary install. Its confined shell replaces the host's own tools.

Differences between arms are fair, because every arm runs the same way, but the absolute AI usage a user sees at home has not been compared with the arena's.

What the tasks can and cannot show

  • Tasks come from the skill's own text. A model writes each task and its required checks

from the skill, and another call grades the runs. The tests cover what a skill says it does, not what its users might ask of it. No agreement rate between the grader and a human reader is published with these studies.

  • Untestable skills are unknown, not negative. A skill that needs a network, a credential,

a live service or a person cannot run in the arena. Nothing is known about whether it can be made lighter, and it is never counted as a failure.

How far the results reach

  • Few skills, few tasks, three runs. Rule checks were compared on

8 skills and prepared facts on 31, each on one or two tasks per skill. Three runs is a floor, not a confidence bound. None of these results has passed the project's separate confirmation track, which SPEC-029 describes and which requires independent tasks left untouched during development. They are development observations, and a tie in quality on a few tasks is not proof that quality never drops.

  • Ranked is not random. The rule-check study started from

15 skills ranked as the best fit for that kind of check, so its acceptance rate describes skills where the check applies, not how often it applies. On a random draw of the catalog, 3 of 64 skills reached a test at all.

  • Seen is not unseen. Rule checks were accepted on

4 of 8 skills in the study that found them and on 0 of 18 skills measured out of sample. The two are never pooled, and the honest headline for a new skill is the second.

  • One rule was learned in sample. The rule that a rule check may not forbid what the skill

exists to do was read off the same comparisons it explains. Out of sample it refused 2 rule checks before measurement, and it let through 5 whose task grade then regressed, of which 0 of 5 were accepted because other conditions caught them. How often it refuses a check that would have worked is unobservable, because a refused check is never run.

Study by study

  • Cheaper on a smaller model. Quality held means within 0.05 on the tasks tested. Across

the study, quality held on 27 of 41 comparable tasks. The result is a cost ratio at the prices of the day; it does not reduce AI usage and is never added to a usage figure.

  • Delegation. One first-party pilot skill, two tasks and

3 runs per arm. It shows why the mechanism lost, not how often it would.

  • The optimization leaderboard. It predates the three-arm protocol: it compares two

versions, its runs were scored by checks written for each skill in its benchmark, and it ranks on skill tokens, which are a lower bound on what a skill costs. 1 skill is excluded with the reason on the board.