SkillOptimizer
Methodology

Every measured test on this site runs the same tasks with two kinds of version of your skill.

VersionWhat runsWhat it establishes
Originalthe tasks with your skill as you published itwhat the skill costs today, and the results to match
Lighterthe tasks with a lighter version AIR builtwhether the same job gets done with fewer tokens

AIR may build up to two lighter versions of one skill: a rule check, which blocks a command the skill already forbids, and prepared facts, which gather what the skill reads on every run before the assistant starts. Each one runs the same tasks as the original and is judged on its own.

The order of the versions rotates from task to task. An earlier cycle ran the same version first every time, and the provider's prompt cache turned the same data into a large saving or a large loss depending only on that order. SPEC-030 section 5 records it.

What decides the answer

A lighter version is kept only when it uses fewer complete execution tokens than the original on the median, every run agrees on the direction, and its grade is not more than 0.05 below the original's on the median and on every task. "What the answers mean" states the rule in full.

Measured example

On ce-work, one of the accepted rule checks, the median run used 1,065,055 tokens with the original skill and 539,226 with the lighter one, over 6 runs per version, with the skill file byte for byte unchanged. Prepared facts were kept on 7 of 31 compared skills.

A note on the research studies

Some research studies listed on this site used other designs, and each names its own versions on its page. The optimization leaderboard is older and compares two versions, original and improved. "Known limits" repeats that.