The obvious way to make a skill cheaper is to make its file shorter. It does not work, and the site never presents a shorter file as an improvement. The quantity that matters is the complete AI usage of the task: every token the assistant reads and writes from the prompt to the final answer. The skill file is one input to that, and usually a small one.
Where the usage goes
An assistant works in steps. On each step it sends the whole conversation so far back to the model, including its instructions, its tool definitions, the skill and every result it has seen. One extra step therefore bills the entire context again, and the context is many times larger than any skill file. SPEC-030 section 0 sets out the arithmetic.
The leaderboard shows the scale. Its skill tokens count everything a run adds beyond the context the assistant starts with, and averaged over its rows they come to 2.5 percent of the session's tokens. Most of a session is that starting context, the assistant's instructions and tool definitions, sent again on every step. The number of steps multiplies that cost; the size of the skill file barely moves it.
What the measured wins have in common
None of the measured improvements works by shortening the file.
- Rule checks leave the file byte for byte unchanged. Every accepted rule check,
4 of 4, shipped the original skill file untouched. The saving comes from stopping work the skill already forbids: a run that breaks the rule goes on to do the forbidden work, and that work costs steps.
- Prepared facts make the file larger. The prepared-facts version appends a block of fixed
read-only commands to the skill, so every compared skill file grew. It was still accepted on 7 of 31 skills, because those commands hand the assistant facts it would otherwise spend steps collecting.
What happened when text was moved out
The delegation study moved a long reference out of the assistant's prompt and gave it to either a cheaper model or a lookup script. The prompt got smaller. The runs got more expensive, because the assistant made the lookup several times and every call was another step that billed the whole context again.
| Where the reference went | Main-session tokens against the original, mean of two tasks |
|---|---|
| a cheaper model | 88.1 percent more |
| a lookup script | 300.5 percent more |
Counting the cheaper model's own tokens as well, the first arm used 155.8 percent more in total. That study is one first-party pilot skill and two tasks, so it shows the mechanism rather than a rate.
The project's earlier attempts at shortening skills point the same way. In the comparisons that were run, compressing skill files made every run more expensive, and a rewriter that removed instructions came back with runs that explored, retried and reread more to make up for what was gone. SPEC-030 section 0 and SPEC-031 section 0 record both. Those comparisons predate the exported tables, so their figures are not repeated here.
What this means for a result
- A before-and-after that shows fewer bytes shows fewer bytes. It says nothing about AI usage
until the task has been run.
- An improved version may be larger than its original, and that is not a warning sign.
- The only accepted evidence of lower usage is the three-arm test described in "The three
arms", at three runs or more, with the mission kept.