A number without its origin cannot be checked, so every figure in the research hub carries four things beside it: its evidence class, its n, the date of the last run behind it, and the commit the figure comes from. The exported tables carry all four on every row, and every study names the source document it was taken from.
The three classes
| Class | What it means | On this site |
|---|---|---|
| estimate | computed from a model of cost or behaviour, with no run behind it | labelled wherever it appears, never pooled with a measured figure |
| measured-local | measured by running the task, on one machine, in one harness | every published figure |
| measured-production | measured on real users' real runs | none yet |
The classes come from the project's benchmark discipline, recorded in the roadmap and in SPEC-008. Every published figure on this site is measured-local. That means it was produced by running real tasks with a real assistant, on one machine, on skills from the AirMarket catalog, with tasks written for the test. It does not mean the figure will hold on your machine, your tasks or your host. No figure here is measured-production, and none is described as one.
An estimate is never added to a measured figure, and two different populations are never pooled into one headline. Seen and unseen skills are the clearest case. Rule checks were accepted on 4 of 8 ranked skills in the study that found them and on 0 of 18 skills measured out of sample. Those two figures are reported side by side and never as one rate.
What n means
n is always stated where a figure appears, and each table says in its own header what it counts: skills, tasks or runs. A share always comes with its numerator, its denominator and, where the study computed one, its confidence interval. The prepared-facts study's headline is shown as 7 of 31 compared skills, never as a percentage alone.
Exclusions are stated with the figure, not under it. A skill that could not run offline, or whose runs did not survive, is named outside the denominator. A reader who sees only the percentage has not seen the result.
A result at a small n is an observation, not a reliability claim. When a page describes one, it says how many skills and runs stand behind it.
When a figure changes
Sometimes a measurement instrument turns out to have been wrong, or a figure printed in a source document does not reproduce from the records behind it. In both cases the site reports the figure the records give, and the exported study lists the superseded figure, what replaced it and why. Nothing is silently overwritten.
Four were figures that did not reproduce from their records, and in each the exported figure is the one reported:
- Cheaper on a smaller model. The study's own summary printed a median cost ratio of
0.42, because tasks with no comparable run on the default model recorded a ratio of zero. The reported figure is 0.657, the median over the 16 skills that held quality and cost less. The source document's split of the other skills did not reproduce either, and the reported split is 13 that dropped quality and 5 that held quality but were not cheaper.
- Prepared facts. The source document split the skills it did not accept into
17 inconclusive and 5 rejected for a mission loss. The records decide 18 inconclusive and 4 rejected for a mission loss. The 7 accepted are unchanged.
- Rule checks from absolute prohibitions. The source document said every compared skill
had runs that disagreed in sign. The records show 11 of 12. Nothing was accepted either way.
- The optimization leaderboard. Five of its skills were run again after their improved
versions were found installed with broken references. The winners' figures it first printed, a fall of 21.2, 12.0 and 22.8 percent in skill tokens, money and time, therefore no longer reproduce. The reported figures are 20.6 percent fewer skill tokens, 10.8 percent less money and 13.4 percent less time.