Every arm runs each task at least three times. A run is one complete attempt by the assistant, from the prompt to its final answer, and the studies call each run a repetition. SPEC-031 section 4 sets the minimum, and no saving measured on fewer runs is accepted, whatever its size.
The rule
- The median, not the mean. The saving for a skill compares the median AI usage of the
improved version's runs with the median of the original's. One runaway run should not decide a skill.
- The spread is shown beside the figure. Every row carries the smallest and largest saving
across its runs, so a reader can see how far a single run could have moved it.
- Runs that disagree in sign make the row inconclusive. When some runs of the improved
version are cheaper than the original and others are dearer, the row is reported as inconclusive, never as a win, however good its median looks.
Why one run is not enough
The project learned this the expensive way. One skill, on identical tasks, measured clearly cheaper on one run and clearly dearer on the next. SPEC-031 section 0 records the case. At a single run, a real saving and a coin flip look the same.
The exported studies show the same thing at scale.
- Of the rule checks built from absolute prohibitions,
11 of the 12 compared skills had runs that disagreed in sign. cpp-testing had a median saving of 29.0 percent, and its individual runs ranged from 77.2 percent dearer to 60.6 percent cheaper. On one run it could have been reported as either.
- Of the 31 skills compared with prepared facts, the
precompute family, 18 were inconclusive because their runs disagreed in sign. That is the single largest outcome in the study, larger than the 7 accepted.
- Even an accepted result moves a lot from run to run.
ce-work, the largest accepted rule
check, saved 49.4 percent at the median, with single runs ranging from 12.5 to 64.0 percent. The sign held on every run, which is why it counts; the size of the saving is far less certain than its direction.
What three runs buy, and what they do not
Three runs are a minimum, not a confidence bound. They rule out a saving that a single lucky run produced, and they make the spread visible. They do not establish how large the saving is with any stated confidence, and they say nothing about tasks the skill was not tested on. A result at three runs on one or two tasks is a measured observation, and the pages describe it that way.
The cost is real: three runs roughly triple what a test costs. That is why eligibility comes first, so that skills with nothing to improve are dropped before any run is repeated. SPEC-031 section 4 puts it as three runs on a third of a corpus costing about what one run on all of it does.