ResultAI usage (tokens)
Rule checks
Does a rule check built from a skill's own prohibition use less AI usage with the same results?
measured-local2026-09-185658cf4
- docs/experiments/COMPILER-2026-09-19.md · 5658cf4 · 2026-09-19
Headline figures
15
Skills attempted
measured-local · 2026-09-18
8of 15
Reached a three-arm comparison
measured-local · n=15 · 2026-09-18
4of 8
Accepted
95% interval 21.5% to 78.5%
measured-local · n=8 · 2026-09-18
26.9%
Median token reduction over the accepted skills
measured-local · 2026-09-18
4of 4
Accepted with SKILL.md byte-identical
measured-local · n=4 · 2026-09-18
5
No baseline advantage
measured-local · 2026-09-18
2
Cannot run offline
measured-local · 2026-09-18
Notes
- Ranked candidates: these 15 were the top of a ranking of 5,914 skills for the conditional-prohibition shape, so the acceptance rate is where the rule check applies, not how often it applies. The random-draw counterpart is the compiler-loop study.
- The mission rule, that a rule check may not forbid what the skill is for, was read off these eight comparisons; its fit here is in-sample by construction. See the mission-rule study.
Tables
n: Runs per arm behind the row: repetitions times tasks compared, or the graded runs of the eligibility check when the skill was not compared. 15 rows.
| skill | name | shape | installs | status | status_superseded | decision | accepted | artifacts | skill_md_unchanged | token_reduction | spread_min | spread_max | straddles_zero | net_call_change | tokens_none | tokens_original | tokens_optimized | score_none | score_original | score_optimized | mission_preserved | retention_band | R | tasks | n | evidence_class | date | commit |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ce-work | Work Execution Command | workflow | — | compared | — | accepted | Yes | hook | Yes | 0.4937 | 0.1249 | 0.6397 | No | -1.5 | 236,619 | 1,065,055 | 539,226 | 0.8 | 1 | 1 | Yes | full | 3 | 2 | 6 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| code-review-expert | Code Review Expert | reference-lookup | 13,180 | compared | — | accepted | Yes | hook | Yes | 0.3028 | 0.1089 | 0.3028 | No | -1 | 23,913 | 44,482 | 31,013 | 0.8 | 1 | 1 | Yes | full | 3 | 1 | 3 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| rds-sqlserver | Amazon RDS for SQL Server | workflow | 3,348 | compared | — | accepted | Yes | hook | Yes | 0.2341 | 0.0043 | 0.3872 | No | -1 | 16,663 | 76,880 | 58,879 | 0.2 | 1 | 1 | Yes | full | 3 | 1 | 3 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| github-project-triage | GitHub Project Triage | reference-lookup | 89 | compared | — | accepted | Yes | hook | Yes | 0.0378 | 0.0199 | 0.2618 | No | 0 | 21,476 | 51,049 | 49,117 | 0.6 | 0.8 | 1 | Yes | full | 3 | 1 | 3 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| write-memo | Write Investment Memo Skill | reference-prose | 7 | compared | — | rejected_mission_failed | No | hook | Yes | 0.2381 | 0.2171 | 0.5397 | No | 0 | 102,254 | 214,300 | 163,284 | 0.8 | 1 | 0.6 | No | lost | 3 | 2 | 6 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| cloudflare-one-migrations | Cloudflare One Migrations | reference-prose | 55,775 | compared | — | inconclusive_repetitions_straddle_zero | No | hook | Yes | -0.0493 | -0.572 | 0.2857 | Yes | 0 | 26,654 | 51,907 | 54,465 | 0.2 | 0.8 | 0.8 | Yes | full | 3 | 1 | 3 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| dev-commit | Dev Commit | workflow | 2 | compared | — | rejected_not_cheaper | No | hook | Yes | -0.5188 | -1.3633 | -0.2421 | No | 5 | 149,257 | 169,655 | 257,674 | 0 | 0.8 | 0 | No | lost | 3 | 2 | 6 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| prp-commit | Commit Intended Work | workflow | — | compared | — | rejected_not_cheaper | No | hook | Yes | -0.7969 | -2.117 | -0.6558 | No | 3 | 47,001 | 65,366 | 117,453 | 0 | 0.2 | 0 | No | lost | 3 | 1 | 3 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| architecture-decision | ADR-[NNNN]: [Title] | reference-lookup | — | no_baseline_advantage | — | — | No | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 3 | 2 | 6 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| azure-reliability | Azure Reliability Assessment & Configuration | reference-lookup | 266,100 | unmeasurable | — | — | No | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 3 | 0 | 0 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| ce-work-beta | Work Execution Command | workflow | — | no_baseline_advantage | — | — | No | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 3 | 2 | 6 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| do | Do Plan | workflow | 6,528 | no_baseline_advantage | — | — | No | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 3 | 2 | 6 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| docs-authoring | Docs Authoring | workflow | — | no_baseline_advantage | — | — | No | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 3 | 2 | 6 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| migrate-to-codex | Migrate to Codex | workflow | 2,068 | unmeasurable | — | — | No | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 3 | 0 | 0 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |
| plan-payroll | Plan Payroll | workflow | 1,533 | no_baseline_advantage | — | — | No | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 3 | 2 | 6 | measured-local | 2026-09-18 | 5658cf4666bb9158f16708aa0a48dfb3c59d752e |