Nulls
A null is a result and is stated the same way as a win: what we measured, on how many skills, with what interval. These are the artifact families and bands that did not pay, out of sample or in general.
Rule checks from absolute prohibitions
Does a rule check pay when the prohibition has no condition?
Skills attempted
measured-local · 2026-09-19
Reached a three-arm comparison
measured-local · n=20 · 2026-09-19
Accepted
95% interval 0.0% to 24.3%
measured-local · n=12 · 2026-09-19
Tasks the assistant already solves
Does a stricter task prompt reduce tasks the assistant solves perfectly without the skill?
Skills drawn
measured-local · 2026-09-19
Skills that cannot run offline, no task graded
measured-local · 2026-09-19
Tasks graded without the skill
measured-local · 2026-09-19
Rule checks on file paths
Does a rule check that blocks writes to named files pay?
Skills attempted
measured-local · 2026-09-19
Reached a three-arm comparison
measured-local · n=10 · 2026-09-19
Accepted
95% interval 0.0% to 49.0%
measured-local · n=4 · 2026-09-19
Rule checks from conditional prohibitions, the rest of the band
Does the one paying rule-check shape repeat on the candidates not yet measured?
Skills attempted
measured-local · 2026-09-19
Reached a three-arm comparison
measured-local · n=5 · 2026-09-19
Accepted
95% interval 0.0% to 65.8%
measured-local · n=2 · 2026-09-19
The mission rule, out of sample
Does refusing a rule check that forbids the skill's own purpose hold on skills it never saw?
Unseen skills the rule could see
measured-local · 2026-09-19
Rule check refused by the rule before measurement
measured-local · 2026-09-19
Comparisons the rule let through
measured-local · 2026-09-19
Helpers built from documented steps
Does moving a skill's documented creation steps into a helper reach a test at all?
Skills attempted
measured-local · 2026-09-19
Reached a three-arm comparison
measured-local · n=15 · 2026-09-19
Accepted
measured-local · n=0 · 2026-09-19
Improvement on a random draw
On a random draw of skills, how often does any improvement reach a test?
Rounds run
measured-local · 2026-09-19
Rounds that compiled anything
measured-local · n=8 · 2026-09-19
Skills drawn
measured-local · 2026-09-19
Handing a step to a cheaper model or to code
Does moving a lookup step off the main assistant use less AI usage?
Arm B, mean main-session token reduction; negative is dearer
measured-local · 2026-09-22
Arm B, mean total token reduction including the delegate
measured-local · 2026-09-22
Arm B, mean cost reduction
measured-local · 2026-09-22
Rule checks out of sample
Across every band measured on unseen skills, how many rule checks paid?
Unseen skills compared
measured-local · 2026-09-19
Unseen skills accepted
95% interval 0.0% to 17.6%
measured-local · n=18 · 2026-09-19
Seen, ranked lane K skills accepted
95% interval 21.5% to 78.5%
measured-local · n=8 · 2026-09-19