SkillOptimizer
Methodology

Every measured test ends in one of three answers: lighter, already light, or can't be tested here. Two further states sit beside them, and one result lives on a separate axis. This page defines each one and the rule behind it.

Lighter

"Lighter" is shown only after a measured test passes. A static analysis that finds something to change is not enough. The lighter version passes when every one of these holds:

  1. At least three runs per task, per version. SPEC-031 section 4 sets the minimum.
  2. It uses fewer tokens. The median complete execution tokens of the lighter version are

below the original's, and every run agrees: if some runs are cheaper and some are dearer, the result is inconclusive. SPEC-031 sections 4 and 5 set both halves.

  1. The mission still succeeds. The lighter version's grade is not more than 0.05 below the

original's, on the median and on every task separately. SPEC-031 section 5 sets the tolerance.

  1. It does not add steps. A version that needs more tool calls than the original fails,

even when its token figure improved, because a mechanism that adds steps is not understood well enough to trust. The project's stop rules, in its compiler task list, set this.

The measured improvement families pass this rule rarely, and the pages say how rarely. A rule check, the hook in the studies' terms, was accepted on 4 of 8 compared skills, each with its skill file byte for byte unchanged. Prepared facts, the precompute in the studies' terms, were accepted on 7 of 31 compared skills. Both are statements about the tasks those skills were tested on, at three runs.

Not tested yet

When a check finds a change it can make, such as a rule the assistant could be held to, the answer is "one thing to try, not tested yet" until a measured test passes. The distinction matters because an analysis that finds a change is a weak predictor of a test that passes. Rule checks measured on skills outside the study that found them were accepted on 0 of 18. On a random draw of the catalog, 3 of 64 skills reached a test at all, and 0 of the 3 were accepted.

Nothing to lighten

"Nothing to lighten" is the answer when a test found nothing in the skill that a known way of lightening applies to, so no lighter version was built and nothing was compared. It says the methods do not fit this skill. It does not say the original is already light, because that was never measured. On the random draw, 30 of 64 skills had no deterministic work that any of the detectors could find.

Already light

"Already light" is the answer when a test built lighter versions, ran them and none reached "lighter", and it is a result, not an error. It covers several different findings, and the page for each skill says which one applies:

  • The result was inconclusive. Of the skills compared with prepared facts,

18 were inconclusive because their runs disagreed in sign. A median saving is not a win when single runs point both ways.

  • The lighter version broke the task or cost more. dev-commit is the clearest case. Its

rule check forbade committing, which is what the skill exists to do. The grade fell from 0.8 to 0, and the run cost 51.9 percent more with 5 more tool calls as the assistant worked against the check.

Can't be tested here

The confined arena has no network, no credentials and no person to answer questions. A skill that needs any of those gets this answer, with the reason. It is never counted as a failure, and the security review still runs.

Licence keeps it as is

A skill whose licence does not allow changes is never made lighter, whatever a check finds. Pasted text and uploads count as the user's own. A public link with no licence, or with licence terms we don't recognise that say nothing against changes, can be made lighter: we change a skill unless its licence says we may not. The security review still runs. The site's product definition records this rule among its decisions.

Cheaper to run, on its own axis

Some skills let a smaller model match the default model's results. That is a money result, not a token result: it lowers the bill per run without lowering token usage, and it is never added to a token figure. In the smaller-model study, 16 of 34 skills held quality on every comparable task and cost less, at a median of 65.7 percent of the default model's cost. Held quality means within 0.05 on the tasks tested, which says nothing about tasks that were not.