Method
One run
Operator prepares start state ─► Agent receives task prompt (verbatim)
│
▼
Agent acts ─► Operator screenshots ─► Operator logs step ─► repeat
│
▼
Stop (agent done / limit / intervention)
│
▼
Operator verifies result independently ─► Scores five dimensions ─► Labels failures
The agent interacts through the GUI only: screenshots in, mouse and keyboard out. No scripting consoles, APIs or plugins. The operator never helps mid-run.
Task ladder
Every application follows the same six families so results compare across tools.
| Family | Tests |
|---|---|
| Orientation | Launch state, dialogs, workspace, opening a file |
| Navigation | Orbit, pan, zoom, views and display modes |
| Selection and inspection | Picking a specific element; reading a property |
| Creation and modification | Primitives, transforms, exact values, materials or parameters |
| Multi-step workflow | A realistic short job composed of the above |
| Recovery | Any of the above with a seeded obstacle |
Five dimensions, 0–3 each
| Dimension | Question |
|---|---|
| Visual perception | Did the agent correctly read the screen? |
| Spatial reasoning | Did it understand the 3D scene from 2D views? |
| Action execution | Did it perform the right action on the right target? |
| Error recovery | When something went wrong, did it notice and fix it? |
| Execution efficiency | How many steps versus a competent human reference path? |
Full anchors: scoring rubric. Controlled labels: failure taxonomy.
Reporting rules
- Pass rate is the headline, decided by the operator’s verification checks, never by the agent’s own report.
- Dimension scores are the median of at least three runs. n is always shown.
- No single aggregate number is shown without the per-dimension breakdown beside it.
- Every published number traces to rows in
scores.csv, and every row traces to a run folder with screenshots.
Limitations
- Results depend on application version, resolution, theme and locale; they are a snapshot under stated conditions.
- One operator introduces rater bias; the rubric anchors and periodic re-scoring are the mitigation, not a cure.
- GUI-only is a deliberate constraint. Many of these tools are more tractable via APIs; nothing here speaks to that.
- Several applications require commercial licences, which limits independent reproduction.