SkillOptimizer
Methodology

The skills in these studies are third-party text from a public catalog, and a test asks an assistant to follow that text unattended, many times over. The arena is the place where that happens without the skill being able to reach anything outside the test.

How a run is confined

The assistant itself runs on the host, but its own file and shell tools are switched off. It has two tools: one that loads the skill, and one shell whose commands run inside a container. Every read, write and command goes through that shell. SPEC-030 section 5 describes the setup.

  • No network. The container has no network access, so a skill cannot download, upload or

call a live service.

  • No credentials. No host keys, tokens or login sessions are visible inside it.
  • Reduced privileges. The container runs with all of its Linux capabilities dropped, and

its system files are read-only; only the workspace can be written.

  • A disposable workspace. The task's files live in a temporary filesystem inside the

container. When the run ends, the workspace is copied out and only then graded. If the copy fails, the failure is recorded; it is never graded as an empty result.

A security review answers a different question, whether a file contains something harmful. The arena answers whether it is safe to let an assistant follow a skill's instructions unattended, and it does that by limiting what any instruction can reach.

What cannot be tested here

Some skills exist to do exactly what the arena forbids: call a live service, use a login, administer a running system, or ask a person a question. Those skills get the answer "can't be tested here", with the reason. In the rule-check study 2 skills could not run offline.

An excluded skill is not a skill that failed. Whether it can be made lighter is unknown, and the studies count it outside the denominator rather than as a negative.

A skill whose tasks need the assistant's own host configuration, for example one that migrates a host's settings, is also out of scope. The arena refuses to start with host policy installed in the workspace, and that check runs before any spend.

The arena is frozen

The studies run on their own copy of the arena, vendored into the campaign rather than shared with other experiments. Each vendored file records the path and hash of the file it was copied from, and a test fails when that source changes, so a change upstream becomes a deliberate decision about this copy and never a silent drift. The shared evaluation sandbox is pinned byte for byte by earlier retained evidence, which is why it is never edited in place.

Every way the vendored arena differs from its source is a reviewed difference with its own tests:

  • Rule checks are enforced inside the arena. When the improved version ships a rule

check, the hook in the studies' terms, the arena refuses the forbidden tool call the way the assistant's host would before running it.

  • Prepared facts are expanded inside the arena. When the improved version appends fixed

read-only commands to its skill, the precompute in the studies' terms, the arena runs only those declared commands at the moment the skill loads, as the host would, and logs each one.

  • A delegated step leaves the container through one narrow door. For the delegation

study, a host process reads one pinned reference file out of the container and makes one bounded model call with tools disabled, returning only the result. That call's usage is counted apart from the main session.

What the arena does not model

The assistant's own file and shell tools are replaced by the confined shell, and the skill is loaded by the arena's own loader rather than the host's. The comparison between arms is fair, because every arm runs the same way, but how closely the absolute AI usage of a run here matches an ordinary install, where the host's own tools are available, has not been measured. The studies report differences between arms, not the usage a particular user will see.