SkillOptimizer
Research hub

Optimization leaderboard

Does the improved version do the same job with less AI usage?

measured-local2026-09-112ebfd81
  • docs/specs/SPEC-027-optimization-leaderboard.md · 5e9cb0e · 2026-09-15
  • docs/experiments/MEASUREMENT-INTEGRITY.md · 8fce0d5 · 2026-09-15

Headline figures

12

Skills on the board

measured-local · 2026-09-11

7of 12

Won: leaner on skill tokens and correctness held

measured-local · n=12 · 2026-09-11

3of 7

Won and more correct

measured-local · n=7 · 2026-09-11

1

Excluded, with the reason on the board

measured-local · 2026-09-11

-20.6%

Winners, skill tokens change, as a sum of per-skill means

measured-local · 2026-09-11

-10.8%

Winners, billed money change

measured-local · 2026-09-11

-13.4%

Winners, wall-clock change

measured-local · 2026-09-11

-10.0%

All included skills, skill tokens change

measured-local · 2026-09-11

2.5%

Share of a session the skill accounts for, mean over rows

measured-local · 2026-09-11

Notes

  • The board ranks on skill tokens, a lower bound: largest request prefix minus first request prefix, plus output, per SPEC-027 section 1. Negative is better in every change column.
  • Money is the billed cost of the whole session, a different population from skill tokens, per SPEC-027 section 7.
  • Wins and losses are both large, so the per-skill row is the result and the all-skills mean is not.

Every skill, with its outcome

Exclusions are shown, not dropped: a skill with no successful run still appears, with the reason.

Showing 13 of 13 rows. n: Successful runs behind the row, both versions together; for an excluded skill, every recorded run.

skilloutcomeexcluded_reasonskill_tokens_changecost_changeseconds_changecorrectness_originalcorrectness_optimizedinstallsn
stitch-loopwon—-0.4713-0.2349-0.467812/1212/1253,43924
json-canvaswon—-0.3821-0.18640.988812/1212/1267,69324
scaffold-exerciseswon—-0.2612-0.2572-0.570612/1212/12312,77924
skill-creatorwon—-0.1412-0.04080.13227/129/12374,47024
hyperframeswon—-0.13740.0837-0.268911/1211/12476,61624
obsidian-cliwon—-0.0529-0.0497-0.10927/1211/1273,60324
vercel-cli-with-tokenswon—-0.0227-0.0759-0.068310/1212/1294,03824
deploy-to-verceldid_not_win—-0.1432-0.0861-0.05828/126/12121,85024
vercel-optimizedid_not_win—0.02460.096-0.21179/1211/1266,26224
domain-modelingdid_not_win—0.09240.0313-0.088612/1212/12583,52824
gsapdid_not_win—0.11660.18790.75338/128/1295,13724
impeccabledid_not_win—0.89530.49480.38712/126/12265,36124
mcp-builderexcludedno_successful_cells—————111,81524

Superseded figures

Winners, skill tokens / money / time

-21.2% / -12.0% / -22.8% → -20.6% / -10.8% / -13.4%

SPEC-027 reproduced the 2026-09-10 meeting evidence at commit 8d6b27b. Commit 9325ef6 re-ran five skills on 2026-09-11 after their improved version was found installed with dangling pointers. The committed results tree has changed since the meeting evidence and now gives the figures here. SPEC-027 section 9 test 1 still pins the meeting evidence.

docs/specs/SPEC-027-optimization-leaderboard.md

Winners, as first printed on the meeting board

-20.5% / -11.5% / -22.8% → -21.2% / -12.0% / -22.8%, then the current figures

The meeting board took an unpaired mean for skill-creator; SPEC-027 section 2 pairs per fixture.

docs/specs/SPEC-027-optimization-leaderboard.md

All skills footer, skill tokens / money / time

-9.6% / -3.8% / -11.1%, then -9.8% / -4.1% / -11.2% → -10.0% / -3.0% / -6.9%

The meeting footer did not reproduce from its own evidence; SPEC-027 recomputed it, and the results tree has changed since.

docs/specs/SPEC-027-optimization-leaderboard.md

results-all summary: cheaper / correctness not worse / both, of the skills measured

7 / 11 / 6 of 13 → 6 / 9 / 5 of 11 skills with complete usage

The 2026-09-15 re-score on complete execution tokens excludes mcp-builder, which has no successful run, and gsap, which has two runs without usage. The board keeps the SPEC-027 ranking and marks gsap in integrity_2026_09_15.

docs/experiments/MEASUREMENT-INTEGRITY.md