Quality on this site means one thing: the share of a task's required checks that a run passed. A separate grader decides each check, and it never learns which arm it is scoring.
Who does what
| Role | Who | What it sees |
|---|---|---|
| The assistant under test | Sonnet 5, running the task in the confined arena | the task, and the skill when its arm installs one |
| The task writer | Haiku 4.5, before any run | the skill's own text, from which it writes each task and its checks |
| The grader | Haiku 4.5, a separate call per run | the task, the checks, the assistant's final answer and the files in the workspace when the run ended |
Two studies differ. The smaller-model study runs the assistant on Haiku 4.5 as well, because that is the question it asks. The optimization leaderboard predates this grader, and its runs were scored by checks written for each skill in its own benchmark, the bench oracles that SPEC-027 names.
The grader is not told which arm produced the run, whether a skill was installed, or that one of the versions is the improved one. Telling a grader that one side is "the optimized version" is how a grader learns to prefer it. SPEC-030 section 5 requires grading to be blind to arm. The grader's input names no arm, and the workspace it sees leaves out the directory where the skill is installed, so the skill's own files cannot give the arm away. The assistant's answer can still mention a skill; the grader is instructed to judge only whether each check is met.
How a check is decided
For every check the grader returns passed or failed and quotes the text or file content that decides it. A check passes only when the answer or the workspace demonstrates it. When the evidence is absent or ambiguous, or the grader would have to assume something to make it true, the check fails. Effort, length and confident wording earn nothing, and a different approach from the one the grader would take is not penalised as long as the check is met.
What makes a good check
A task carries three to five required checks, and each one must be decidable from the final answer and the final workspace. SPEC-030 section 5 bans checks on process. An early pilot showed why: a check such as "formed a hypothesis before fixing" failed an assistant that went straight to the correct fix, so both arms scored badly while both had diagnosed the problem correctly. A check that a competent direct solution fails is measuring ritual, not quality.
From checks to a decision
A run's grade is the share of its checks that passed. The improved version keeps the mission when its grade is not materially below the original's, and materially means a drop of more than 0.05. SPEC-031 section 5 sets that tolerance. With three to five checks a task, one check is worth a fifth to a third of the grade, so the tolerance absorbs rounding and nothing else: losing a single check counts as losing quality.
The comparison holds on the median and on every task separately. An average across tasks can hide one task that broke behind another that improved, and the project has seen exactly that happen to a gate the skill was meant to keep. SPEC-030 section 2.2 records the case.
When the grader is wrong
The grader is an instrument, not the thing under test. Twice it scored correct work as failing, on every arm at once, and both times the cause was the evidence it was handed rather than its judgment.
- The workspace was read before it was saved. The arena's workspace lives inside the
container, and the grader once read it before the run's files were copied out. Every arm was graded on the workspace it started with, and all three scored the same zero while the grader's own quoted evidence showed the assistant had found the bug. SPEC-030 section 0.1 records the defect.
- The grader read log files instead of the files its checks named. The workspace
collector had a size budget and spent it on package-manager logs before it reached the project's own files, so every check read "not captured". A run that reported all its tests passing scored zero, and the same test measured again after the fix scored full marks. The structural compiler write-up records it.
Both were fixed, each with a test that fails if the old behaviour returns, and both scorings were kept. A corrected figure replaces an old one only when the defect can be shown without looking at the score it produced, was found independently of wanting a result, and is pinned by a test. When a correction moves a published figure, the exported study lists the superseded figure and the reason beside the one it reports.
What blind grading cannot promise
The grader is itself a model. These studies publish no agreement rate between it and a human reader. The tasks and checks are written from each skill's own text, so they test what the skill says it does, and a task the skill's author never imagined is not covered. Known limits lists both.