Every measured test on this site runs the same tasks with two kinds of version of your skill.
| Version | What runs | What it establishes |
|---|---|---|
| Original | the tasks with your skill as you published it | what the skill costs today, and the results to match |
| Lighter | the tasks with a lighter version AIR built | whether the same job gets done with fewer tokens |
AIR may build up to two lighter versions of one skill: a rule check, which blocks a command the skill already forbids, and prepared facts, which gather what the skill reads on every run before the assistant starts. Each one runs the same tasks as the original and is judged on its own.
The order of the versions rotates from task to task. An earlier cycle ran the same version first every time, and the provider's prompt cache turned the same data into a large saving or a large loss depending only on that order. SPEC-030 section 5 records it.
What decides the answer
A lighter version is kept only when it uses fewer complete execution tokens than the original on the median, every run agrees on the direction, and its grade is not more than 0.05 below the original's on the median and on every task. "What the answers mean" states the rule in full.
Measured example
On ce-work, one of the accepted rule checks, the median run used 1,065,055 tokens with the original skill and 539,226 with the lighter one, over 6 runs per version, with the skill file byte for byte unchanged. Prepared facts were kept on 7 of 31 compared skills.
A note on the research studies
Some research studies listed on this site used other designs, and each names its own versions on its page. The optimization leaderboard is older and compares two versions, original and improved. "Known limits" repeats that.