Agentic-3D
How well can AI computer-use agents operate professional 3D applications through their graphical interfaces alone?
Agentic-3D is a manual, evidence-led benchmark. A human operator gives a computer-use agent a task in a 3D application, records a screenshot after every action, verifies the result independently and scores five dimensions: visual perception, spatial reasoning, action execution, error recovery and execution efficiency.
The method, the scoring rubric, the failure taxonomy and the raw scores are published on this site. Every published number traces to a scored run with its screenshot sequence.
Status
Framing and task-authoring stage. No results have been published yet. The method, rubric and first task cards are in place. Results will appear on this page, application by application, starting with Blender.
Navigate
- Method — how a run is conducted and scored
- Applications — the 22 tools under study and their status
- Results — pass rates and dimension scores (empty until the first application completes)
- Findings — what the evidence shows
- Raw scores — the CSV every table is built from
Why this matters
Computer-use agents are evaluated almost entirely on browsers, office software and terminals. 3D applications are a harder regime: dense viewports, a 2D projection of 3D state, long stateful workflows and silent failure modes. The architecture, engineering, construction, design, games, simulation and geospatial industries depend on these tools. Whether agents can operate them is a practical question with real consequences, and the answer should rest on inspectable evidence rather than demos.