SkillOptimizer
Research hub
ResultAI usage (tokens)

Prepared facts

Does gathering a skill's routine facts once, before the assistant starts, use less AI usage with the same results?

measured-local2026-09-2652363d2
  • docs/experiments/PRECOMPUTE-2026-09-26.md · 52363d2 · 2026-09-26
  • data: experiments/airmarket-6k/campaign/round-evidence/r9-precompute · 52363d2

Headline figures

31

Skills compared

measured-local · 2026-09-26

7of 31

Accepted

95% interval 11.4% to 39.8%

measured-local · n=31 · 2026-09-26

16.6%

Median token reduction over every compared skill

measured-local · 2026-09-26

18

Inconclusive: the three repetitions disagree in sign

measured-local · 2026-09-26

1

Inconclusive: too few repetitions

measured-local · 2026-09-26

4

Rejected for a mission loss

measured-local · 2026-09-26

1

Rejected, not cheaper

measured-local · 2026-09-26

7

Compared skills whose mission regressed on some task, whatever the decision

measured-local · 2026-09-26

Notes

  • SKILL.md stays byte-identical except one appended block of fixed read-only commands that Claude Code expands at skill load; every component is justified by the original arm's traces, never by the prose.
  • Picking did not help: the cohorts table compares Jev's pick, the trace pick and a random draw, and their intervals overlap.
  • 118 of the 174 skills that beat the base model in earlier studies were plannable; that count is in the source document and is not derived from these rows.

Superseded figures

A figure this project published and later corrected, kept beside what replaced it (CLAUDE.md, "Correcting a measurement instrument").

Outcome split of the 24 skills not accepted

17 inconclusive, 5 rejected for a mission loss → 18 inconclusive by sign, 4 rejected for a mission loss, as the records decide them

The source document's split differs from the recorded decisions by one skill, and the records do not show which skill the document moved. The accepted count, its interval and the median saving are unchanged. mission_preserved is carried on every row, because two inconclusive skills and the not-cheaper one also scored below the original on some task.

docs/experiments/PRECOMPUTE-2026-09-26.md

Tables

Prepared facts, R9, three repetitions per arm

n: Runs per arm behind the row: repetitions times tasks compared, or the graded runs of the eligibility check when the skill was not compared. 31 rows.

skillnameshapeinstallsstatusstatus_supersededdecisionacceptedartifactsskill_md_unchangedtoken_reductionspread_minspread_maxstraddles_zeronet_call_changetokens_nonetokens_originaltokens_optimizedscore_nonescore_originalscore_optimizedmission_preservedretention_bandRtaskscohortscomponentsjev_scoresubsumable_calls_per_runnevidence_classdatecommit
product-visionProduct Visionreference-prose2,585compared—acceptedYesprecomputeNo0.43750.41140.4455No-141,59653,41930,0500.60.80.8Yesfull32randomlisting;inputs0.271.56measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
approachability-auditApproachability Auditreference-prose—compared—acceptedYesprecomputeNo0.32410.29960.3326No-121,83124,91316,8380.811Yesfull31tracelisting;inputs0.232.3333measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
marquee-loopMarquee Skillworkflow—compared—acceptedYesprecomputeNo0.22590.2010.274No-127,83130,36223,5030.7511Yesfull31randomlisting;inputs0.2123measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
github-project-triageGitHub Project Triagereference-lookup89compared—acceptedYesprecomputeNo0.22570.21360.2569No021,47651,04939,5270.60.81Yesfull31jevlisting;inputs0.551.3333measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
rds-sqlserverAmazon RDS for SQL Serverworkflow3,348compared—acceptedYesprecomputeNo0.22410.21860.3783No-116,66376,88059,6530.211Yesfull31randomlisting0.2313measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
jqjq - Built-in JSON Processorreference-lookup—compared—acceptedYesprecomputeNo0.2140.20890.2143No-133,67540,68331,9770.411Yesfull31randomlisting;inputs0.123measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
writing-coreAcademic Writing Corereference-lookup—compared—acceptedYesprecomputeNo0.19890.17280.2622No-130,90251,32441,1170.811Yesfull31tracelisting;inputs0.232.6673measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
ads-planPaid Media Planworkflow—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.4581-0.18030.7078Yes-337,85459,86432,4400.811Yesfull31tracelisting;inputs0.252.3333measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
git-advanced-workflowsGit Advanced Workflowsreference-lookup17,900compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.393-0.07450.4552Yes-5114,182247,955150,49800.60Nolost31tracelisting;inputs;git0.233measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
programmingProgrammingreference-lookup11compared—inconclusive_insufficient_repetitionsNoprecomputeNo0.38940.36420.6266No-7309,224776,690474,2510.611Yesfull21jevlisting;bundle:references/python/README.md0.4922measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
abp-application-layerABP Application Layer Patternsreference-lookup16compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.3842-0.06290.6203Yes-3141,46598,74960,8080.811Yesfull31tracelisting;inputs0.282.3333measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
aws-well-architected-reviewAWS Well-Architected Reviewreference-lookup108compared—rejected_mission_failedNoprecomputeNo0.3630.34450.5339No-122,05430,23619,2590.60.80.6Nolost31jevlisting;inputs0.4823measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
axiom-audit-iapIn-App Purchase Auditor Agentreference-lookup—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.3377-0.16770.5364Yes-223,70198,61265,3100.250.750.75Yesfull31tracelisting;inputs0.252.3333measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
debug-gradient-flowDebugging Gradient Flow in Trainingreference-lookup3compared—rejected_mission_failedNoprecomputeNo0.30330.14180.3377No-2100,148128,77889,7240.250.750.5Nopartial32tracelisting;inputs0.22.6676measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
xcode-project-analyzerXcode Project Analyzerreference-prose—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.2769-0.0680.4219Yes-120,84638,33627,7210.811Yesfull31jevlisting0.411.6673measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
blog-analyzeBlog Analyzereference-lookup—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.1661-0.31410.5742Yes-119,35589,79874,8860.411Yesfull31jevlisting;inputs0.4213measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
prp-commitCommit Intended Workworkflow—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.145-0.1370.145Yes-147,00165,36655,89000.20.4Yesfull31jevlisting0.413measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
stat-research-orchestratorStatistical Research Orchestratorreference-lookup—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.1326-0.43250.2051Yes-146,46377,53267,253011Yesfull31randominputs0.213measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
dmr-from-django-ninjaDMR from django-ninjaworkflow45compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.1325-0.95350.2918Yes-453,954396,121343,65000.80.8Yesfull31jev;tracelisting;inputs0.552.6673measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
unity-netcodeUnity Netcode for GameObjects Skillsreference-lookup—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.1312-0.66270.3321Yes-1518,656548,296476,3480.811Yesfull32tracelisting0.2636measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
catchupContext Catchupworkflow—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.1187-0.76170.12Yes-146,46574,31865,4970.811Yesfull31jevinputs;git0.781.6673measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
gh-awGitHub Agentic Workflowsreference-lookup36compared—rejected_mission_failedNoprecomputeNo0.11180.03880.3452No-146,230138,152122,7000.210.8Nostrong31tracelisting0.3133measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
animation-principlesAnimation Principlesreference-prose—compared—rejected_mission_failedNoprecomputeNo0.08470.07580.1106No-0.564,63862,16956,902010.75Nostrong32randomlisting;inputs0.2126measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
code-review-expertCode Review Expertreference-lookup13,180compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo0.0765-0.02360.126Yes023,91344,48241,0790.811Yesfull31jevlisting0.4623measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
threejs-interactionThree.js Interactionreference-lookup—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo-0.0158-0.01980.2373Yes049,99168,34069,4190.40.60.6Yesfull31tracelisting;inputs0.182.6673measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
modeling-product-usage-metricsModeling product-usage metricsreference-lookup—compared—rejected_not_cheaperNoprecomputeNo-0.0549-0.3406-0.0269No022,89245,36247,8540.60.80.4Nolost31randomlisting0.221.6673measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
publish-project-to-githubPublish Project to GitHubworkflow—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo-0.0592-0.38860.1502Yes0.592,614336,990356,94200.20Nolost32jevlisting0.411.56measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
retrospectiveSprint Retrospectivereference-lookup—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo-0.1684-0.39880.0061Yes133,58059,94770,0430.60.80.8Yesfull31tracelisting0.342.6673measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
update-prUpdate PRworkflow14compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo-0.2156-0.22170.1959Yes228,57250,38161,2410.811Yesfull31jevlisting0.5613measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
draw-io-diagram-generatorDraw.io Diagram Generatorreference-lookup2,724compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo-0.3902-0.88630.1497Yes163,891153,082212,8070.611Yesfull31randombundle:assets/templates/architecture.drawio0.1213measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
wiki-lintLint the wikiworkflow—compared—inconclusive_repetitions_straddle_zeroNoprecomputeNo-0.475-0.56230.07Yes387,993114,576168,9980.7511Yesfull31jevlisting0.671.6673measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645

Prepared facts by selection cohort

n: Skills compared in the cohort. 3 rows.

cohortskillscomparedacceptedinterval_lowinterval_highmedian_token_reductionnevidence_classdatecommit
jev121210.01490.35390.138712measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
trace121220.0470.4480.251112measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645
random8840.21520.78480.17338measured-local2026-09-2652363d25071579bcd00029e20f4278c965578645