SkillOptimizer
Research hub

Nulls

A null is a result and is stated the same way as a win: what we measured, on how many skills, with what interval. These are the artifact families and bands that did not pay, out of sample or in general.

Rule checks from absolute prohibitions

Does a rule check pay when the prohibition has no condition?

measured-local2026-09-1980e4a13
20

Skills attempted

measured-local · 2026-09-19

12of 20

Reached a three-arm comparison

measured-local · n=20 · 2026-09-19

0of 12

Accepted

95% interval 0.0% to 24.3%

measured-local · n=12 · 2026-09-19

Full denominator and tables →

Tasks the assistant already solves

Does a stricter task prompt reduce tasks the assistant solves perfectly without the skill?

measured-local2026-09-1980e4a13
10

Skills drawn

measured-local · 2026-09-19

2

Skills that cannot run offline, no task graded

measured-local · 2026-09-19

16

Tasks graded without the skill

measured-local · 2026-09-19

Full denominator and tables →

Rule checks on file paths

Does a rule check that blocks writes to named files pay?

measured-local2026-09-1980e4a13
10

Skills attempted

measured-local · 2026-09-19

4of 10

Reached a three-arm comparison

measured-local · n=10 · 2026-09-19

0of 4

Accepted

95% interval 0.0% to 49.0%

measured-local · n=4 · 2026-09-19

Full denominator and tables →

Rule checks from conditional prohibitions, the rest of the band

Does the one paying rule-check shape repeat on the candidates not yet measured?

measured-local2026-09-1980e4a13
5

Skills attempted

measured-local · 2026-09-19

2of 5

Reached a three-arm comparison

measured-local · n=5 · 2026-09-19

0of 2

Accepted

95% interval 0.0% to 65.8%

measured-local · n=2 · 2026-09-19

Full denominator and tables →

The mission rule, out of sample

Does refusing a rule check that forbids the skill's own purpose hold on skills it never saw?

measured-local2026-09-1980e4a13
35

Unseen skills the rule could see

measured-local · 2026-09-19

2

Rule check refused by the rule before measurement

measured-local · 2026-09-19

18

Comparisons the rule let through

measured-local · 2026-09-19

Full denominator and tables →

Helpers built from documented steps

Does moving a skill's documented creation steps into a helper reach a test at all?

measured-local2026-09-195658cf4
15

Skills attempted

measured-local · 2026-09-19

0of 15

Reached a three-arm comparison

measured-local · n=15 · 2026-09-19

0of 0

Accepted

measured-local · n=0 · 2026-09-19

Full denominator and tables →

Improvement on a random draw

On a random draw of skills, how often does any improvement reach a test?

measured-local2026-09-190d3b836
8

Rounds run

measured-local · 2026-09-19

2of 8

Rounds that compiled anything

measured-local · n=8 · 2026-09-19

64

Skills drawn

measured-local · 2026-09-19

Full denominator and tables →

Handing a step to a cheaper model or to code

Does moving a lookup step off the main assistant use less AI usage?

measured-local2026-09-22da8b53a
-88.1%

Arm B, mean main-session token reduction; negative is dearer

measured-local · 2026-09-22

-1.56

Arm B, mean total token reduction including the delegate

measured-local · 2026-09-22

-2.18

Arm B, mean cost reduction

measured-local · 2026-09-22

Full denominator and tables →

Rule checks out of sample

Across every band measured on unseen skills, how many rule checks paid?

measured-local2026-09-1980e4a13
18

Unseen skills compared

measured-local · 2026-09-19

0of 18

Unseen skills accepted

95% interval 0.0% to 17.6%

measured-local · n=18 · 2026-09-19

4of 8

Seen, ranked lane K skills accepted

95% interval 21.5% to 78.5%

measured-local · n=8 · 2026-09-19

Full denominator and tables →