Results
No results have been published yet. This page will show, for each application:
- Pass rate by agent and task family (n shown)
- Median dimension scores by agent and task family
- Most frequent failure labels with links to illustrative runs
Tables are built from scores.csv. Each regeneration is logged in the results changelog with the commit hash of the data used.
Table format (preview)
Blender — pass rate
| Family | n | claude-cu | openai-cua |
|---|---|---|---|
| Orientation | — | — | — |
| Navigation | — | — | — |
| Selection and inspection | — | — | — |
| Creation and modification | — | — | — |
| Multi-step workflow | — | — | — |
| Recovery | — | — | — |
Blender — median dimension scores (0–3)
| Agent | Visual perception | Spatial reasoning | Action execution | Error recovery | Execution efficiency |
|---|---|---|---|---|---|
| claude-cu | — | — | — | — | — |
| openai-cua | — | — | — | — | — |