Null resultAI usage (tokens)
Improvement on a random draw
On a random draw of skills, how often does any improvement reach a test?
measured-local2026-09-190d3b836
- docs/experiments/COMPILER-LOOP-NULL-2026-09-19.md · 0d3b836 · 2026-09-19
Headline figures
8
Rounds run
measured-local · 2026-09-19
2of 8
Rounds that compiled anything
measured-local · n=8 · 2026-09-19
64
Skills drawn
measured-local · 2026-09-19
3of 64
Reached a comparison
measured-local · n=64 · 2026-09-19
4.7%
Share that reached a comparison
measured-local · 2026-09-19
30of 64
No deterministic work any generator could find
measured-local · n=64 · 2026-09-19
46.9%
Share with no deterministic work found
measured-local · 2026-09-19
0of 3
Accepted
measured-local · n=3 · 2026-09-19
Notes
- The loop stopped at its eight-round limit, neither converged nor gave up: only rounds that compiled something count toward its registered thresholds.
- Two skills lost their artifact to an instrument defect: scaffold files were dropped before reaching the container. They are the artifact_not_installed outcome, kept as that status rather than counted as skills with nothing to compile.
- 46.9 percent is a property of this corpus as seen by three deterministic detectors, not of agent skills in general.
Tables
n: Skills with this outcome. 8 rows.
| outcome | skills | share | n | evidence_class | date | commit |
|---|---|---|---|---|---|---|
| no_artifact | 30 | 0.4688 | 30 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| unmeasurable | 13 | 0.2031 | 13 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| no_baseline_advantage | 8 | 0.125 | 8 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| tasks_not_discriminating | 6 | 0.0938 | 6 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| compared | 3 | 0.0469 | 3 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| artifact_not_installed | 2 | 0.0313 | 2 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| artifact_refused | 1 | 0.0156 | 1 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| insufficient_repetitions | 1 | 0.0156 | 1 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
n: Skills drawn in the round. 8 rows.
| round | drawn | compiled | accepted | rules_added | n | evidence_class | date | commit |
|---|---|---|---|---|---|---|---|---|
| 1 | 8 | 1 | 0 | 0 | 8 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| 2 | 8 | 0 | 0 | 0 | 8 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| 3 | 8 | 0 | 0 | 0 | 8 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| 4 | 8 | 0 | 0 | 0 | 8 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| 5 | 8 | 2 | 0 | 0 | 8 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| 6 | 8 | 0 | 0 | 0 | 8 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| 7 | 8 | 0 | 0 | 0 | 8 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| 8 | 8 | 0 | 0 | 0 | 8 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
n: Runs per arm behind the row: repetitions times tasks compared, or the graded runs of the eligibility check when the skill was not compared. 3 rows.
| skill | name | shape | installs | status | status_superseded | decision | accepted | artifacts | skill_md_unchanged | token_reduction | spread_min | spread_max | straddles_zero | net_call_change | tokens_none | tokens_original | tokens_optimized | score_none | score_original | score_optimized | mission_preserved | retention_band | R | tasks | n | evidence_class | date | commit |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| webiny-api-architect | API Architecture Patterns | reference-lookup | 1 | compared | — | inconclusive_repetitions_straddle_zero | No | hook | Yes | 0.229 | -0.2595 | 0.4176 | Yes | -2 | 74,934 | 129,805 | 100,075 | 0 | 0.6 | 0.8 | Yes | full | 3 | 1 | 3 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| axiom-swift | Swift Language & Platform | reference-lookup | — | compared | — | inconclusive_repetitions_straddle_zero | No | hook | Yes | 0.0057 | -0.0057 | 0.0206 | Yes | -1 | 21,050 | 55,930 | 55,610 | 0.4 | 1 | 1 | Yes | full | 3 | 1 | 3 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |
| aws-cloudformation | CloudFormation | workflow | — | compared | — | rejected_not_cheaper | No | hook | Yes | -0.4342 | -0.7728 | -0.2091 | No | 0 | 92,266 | 215,759 | 309,437 | 0.8 | 1 | 1 | Yes | full | 3 | 2 | 6 | measured-local | 2026-09-19 | 0d3b836079506c06e00a5ae5fa4d0eede75aa628 |