Optimization leaderboard
Does the improved version do the same job with less AI usage?
- docs/specs/SPEC-027-optimization-leaderboard.md · 5e9cb0e · 2026-09-15
- docs/experiments/MEASUREMENT-INTEGRITY.md · 8fce0d5 · 2026-09-15
Headline figures
Skills on the board
measured-local · 2026-09-11
Won: leaner on skill tokens and correctness held
measured-local · n=12 · 2026-09-11
Won and more correct
measured-local · n=7 · 2026-09-11
Excluded, with the reason on the board
measured-local · 2026-09-11
Winners, skill tokens change, as a sum of per-skill means
measured-local · 2026-09-11
Winners, billed money change
measured-local · 2026-09-11
Winners, wall-clock change
measured-local · 2026-09-11
All included skills, skill tokens change
measured-local · 2026-09-11
Share of a session the skill accounts for, mean over rows
measured-local · 2026-09-11
Notes
- The board ranks on skill tokens, a lower bound: largest request prefix minus first request prefix, plus output, per SPEC-027 section 1. Negative is better in every change column.
- Money is the billed cost of the whole session, a different population from skill tokens, per SPEC-027 section 7.
- Wins and losses are both large, so the per-skill row is the result and the all-skills mean is not.
Every skill, with its outcome
Exclusions are shown, not dropped: a skill with no successful run still appears, with the reason.
Showing 13 of 13 rows. n: Successful runs behind the row, both versions together; for an excluded skill, every recorded run.
| skill | outcome | excluded_reason | skill_tokens_change | cost_change | seconds_change | correctness_original | correctness_optimized | installs | n |
|---|---|---|---|---|---|---|---|---|---|
| stitch-loop | won | — | -0.4713 | -0.2349 | -0.4678 | 12/12 | 12/12 | 53,439 | 24 |
| json-canvas | won | — | -0.3821 | -0.1864 | 0.9888 | 12/12 | 12/12 | 67,693 | 24 |
| scaffold-exercises | won | — | -0.2612 | -0.2572 | -0.5706 | 12/12 | 12/12 | 312,779 | 24 |
| skill-creator | won | — | -0.1412 | -0.0408 | 0.1322 | 7/12 | 9/12 | 374,470 | 24 |
| hyperframes | won | — | -0.1374 | 0.0837 | -0.2689 | 11/12 | 11/12 | 476,616 | 24 |
| obsidian-cli | won | — | -0.0529 | -0.0497 | -0.1092 | 7/12 | 11/12 | 73,603 | 24 |
| vercel-cli-with-tokens | won | — | -0.0227 | -0.0759 | -0.0683 | 10/12 | 12/12 | 94,038 | 24 |
| deploy-to-vercel | did_not_win | — | -0.1432 | -0.0861 | -0.0582 | 8/12 | 6/12 | 121,850 | 24 |
| vercel-optimize | did_not_win | — | 0.0246 | 0.096 | -0.2117 | 9/12 | 11/12 | 66,262 | 24 |
| domain-modeling | did_not_win | — | 0.0924 | 0.0313 | -0.0886 | 12/12 | 12/12 | 583,528 | 24 |
| gsap | did_not_win | — | 0.1166 | 0.1879 | 0.7533 | 8/12 | 8/12 | 95,137 | 24 |
| impeccable | did_not_win | — | 0.8953 | 0.4948 | 0.387 | 12/12 | 6/12 | 265,361 | 24 |
| mcp-builder | excluded | no_successful_cells | — | — | — | — | — | 111,815 | 24 |
Superseded figures
Winners, skill tokens / money / time
-21.2% / -12.0% / -22.8% → -20.6% / -10.8% / -13.4%
SPEC-027 reproduced the 2026-09-10 meeting evidence at commit 8d6b27b. Commit 9325ef6 re-ran five skills on 2026-09-11 after their improved version was found installed with dangling pointers. The committed results tree has changed since the meeting evidence and now gives the figures here. SPEC-027 section 9 test 1 still pins the meeting evidence.
docs/specs/SPEC-027-optimization-leaderboard.md
Winners, as first printed on the meeting board
-20.5% / -11.5% / -22.8% → -21.2% / -12.0% / -22.8%, then the current figures
The meeting board took an unpaired mean for skill-creator; SPEC-027 section 2 pairs per fixture.
docs/specs/SPEC-027-optimization-leaderboard.md
All skills footer, skill tokens / money / time
-9.6% / -3.8% / -11.1%, then -9.8% / -4.1% / -11.2% → -10.0% / -3.0% / -6.9%
The meeting footer did not reproduce from its own evidence; SPEC-027 recomputed it, and the results tree has changed since.
docs/specs/SPEC-027-optimization-leaderboard.md
results-all summary: cheaper / correctness not worse / both, of the skills measured
7 / 11 / 6 of 13 → 6 / 9 / 5 of 11 skills with complete usage
The 2026-09-15 re-score on complete execution tokens excludes mcp-builder, which has no successful run, and gsap, which has two runs without usage. The board keeps the SPEC-027 ranking and marks gsap in integrity_2026_09_15.
docs/experiments/MEASUREMENT-INTEGRITY.md