Method

One run

Operator prepares start state ─► Agent receives task prompt (verbatim)
        │
        ▼
Agent acts ─► Operator screenshots ─► Operator logs step ─► repeat
        │
        ▼
Stop (agent done / limit / intervention)
        │
        ▼
Operator verifies result independently ─► Scores five dimensions ─► Labels failures

The agent interacts through the GUI only: screenshots in, mouse and keyboard out. No scripting consoles, APIs or plugins. The operator never helps mid-run.

Task ladder

Every application follows the same six families so results compare across tools.

FamilyTests
OrientationLaunch state, dialogs, workspace, opening a file
NavigationOrbit, pan, zoom, views and display modes
Selection and inspectionPicking a specific element; reading a property
Creation and modificationPrimitives, transforms, exact values, materials or parameters
Multi-step workflowA realistic short job composed of the above
RecoveryAny of the above with a seeded obstacle

Five dimensions, 0–3 each

DimensionQuestion
Visual perceptionDid the agent correctly read the screen?
Spatial reasoningDid it understand the 3D scene from 2D views?
Action executionDid it perform the right action on the right target?
Error recoveryWhen something went wrong, did it notice and fix it?
Execution efficiencyHow many steps versus a competent human reference path?

Full anchors: scoring rubric. Controlled labels: failure taxonomy.

Reporting rules

Limitations