SkillOptimizer

Research hub

Most answers are "no profitable optimization," and the changes that do pay are rare. Every number here shows its n, its evidence class and the commit it came from, and a null is a result, never a failure.

Leaderboard

Result

Does the lighter version do the same job with fewer tokens? Per SPEC-027, with exclusions shown rather than dropped.

Open the leaderboard →

Results

Rule checks

Does a rule check built from a skill's own prohibition use less AI usage with the same results?

AI usage (tokens) · measured-local · 2026-09-18

Prepared facts

Does gathering a skill's routine facts once, before the assistant starts, use less AI usage with the same results?

AI usage (tokens) · measured-local · 2026-09-26

Cheaper on a smaller model

Does a skill let a smaller model match the default model's results for less money?

Money · measured-local · 2026-09-20

Nulls

A null is titled and shown like any other result. These are the things we tried that did not pay, with the same denominators as everything else.

Rule checks from absolute prohibitions

Does a rule check pay when the prohibition has no condition?

AI usage (tokens) · measured-local · 2026-09-19

Rule checks on file paths

Does a rule check that blocks writes to named files pay?

AI usage (tokens) · measured-local · 2026-09-19

Rule checks from conditional prohibitions, the rest of the band

Does the one paying rule-check shape repeat on the candidates not yet measured?

AI usage (tokens) · measured-local · 2026-09-19

Does refusing a rule check that forbids the skill's own purpose hold on skills it never saw?

35 Unseen skills the rule could see

AI usage (tokens) · measured-local · 2026-09-19

Helpers built from documented steps

Does moving a skill's documented creation steps into a helper reach a test at all?

AI usage (tokens) · measured-local · 2026-09-19

Improvement on a random draw

On a random draw of skills, how often does any improvement reach a test?

AI usage (tokens) · measured-local · 2026-09-19

Does moving a lookup step off the main assistant use less AI usage?

-88.1% Arm B, mean main-session token reduction; negative is dearer

AI usage (tokens) · measured-local · 2026-09-22

Rule checks out of sample

Across every band measured on unseen skills, how many rule checks paid?

AI usage (tokens) · measured-local · 2026-09-19