Results

No results have been published yet. This page will show, for each application:

  1. Pass rate by agent and task family (n shown)
  2. Median dimension scores by agent and task family
  3. Most frequent failure labels with links to illustrative runs

Tables are built from scores.csv. Each regeneration is logged in the results changelog with the commit hash of the data used.

Table format (preview)

Blender — pass rate

Familynclaude-cuopenai-cua
Orientation———
Navigation———
Selection and inspection———
Creation and modification———
Multi-step workflow———
Recovery———

Blender — median dimension scores (0–3)

AgentVisual perceptionSpatial reasoningAction executionError recoveryExecution efficiency
claude-cu—————
openai-cua—————