# Scoring Rubric

Each run is scored on five dimensions using a **0–3 scale**. Score every dimension independently; a run can fail outright on action execution while scoring well on perception. Scores are recorded in `run.md` and `results/scores.csv`.

Report the **median of at least three runs** per agent per task. Never report a single run as a result.

## Scale

| Score | Meaning |
|---|---|
| **3** | Fully correct. No errors in this dimension. |
| **2** | Minor errors that the agent noticed or that did not affect the outcome. |
| **1** | Significant errors that degraded or delayed the outcome but did not cause failure alone. |
| **0** | Failure in this dimension caused or would have caused task failure. |
| **n/a** | Dimension not exercised by this task (declared on the task card). |

## Dimension anchors

### Visual perception — *Did the agent correctly read the screen?*

| Score | Anchor |
|---|---|
| 3 | Correctly identified every relevant control, panel, dialog, object and state it acted on |
| 2 | One misread that it self-corrected on the next step, or a misread of an irrelevant element |
| 1 | Repeated misreads (wrong tool, wrong object, missed dialog) that cost steps |
| 0 | Acted on a misperception that made the task impossible (e.g. never saw the modal blocking input) |

### Spatial reasoning — *Did the agent understand the 3D scene from 2D views?*

| Score | Anchor |
|---|---|
| 3 | Navigated and manipulated with correct sense of position, depth, scale, axis and occlusion |
| 2 | One wrong direction or axis, corrected within two steps |
| 1 | Oscillated, overshot or confused axes repeatedly; reached the goal by trial |
| 0 | Never achieved the required spatial relationship (wrong object framed, wrong axis moved, wrong face selected) |

### Action execution — *Did the agent perform the right action on the right target?*

| Score | Anchor |
|---|---|
| 3 | Every click, drag, key and typed value landed as intended, in order |
| 2 | One slip (off-target click, typo) corrected immediately with no side effect |
| 1 | Several slips, or one slip with a side effect that was later cleaned up |
| 0 | A slip that corrupted the model or left the task incomplete |

### Error recovery — *When something went wrong, did the agent notice and fix it?*

Only scored if an error or seeded surprise occurred. If nothing went wrong, record `n/a`.

| Score | Anchor |
|---|---|
| 3 | Noticed every deviation within one step and recovered cleanly (undo, dismiss, re-plan) |
| 2 | Noticed and recovered, but took more than two steps or left a harmless residue |
| 1 | Noticed late, or recovered partially |
| 0 | Did not notice, or made it worse, or declared success in a failed state |

### Execution efficiency — *How much did success cost?*

Computed, not judged. Use the `reference_steps` on the task card.

| Score | Rule |
|---|---|
| 3 | `agent_steps ≤ 1.5 × reference_steps` |
| 2 | `≤ 2.5 ×` |
| 1 | `≤ 4 ×` |
| 0 | `> 4 ×`, or hit the step/time limit |

Only computed when the task outcome is **pass**. For a failed task, efficiency is `0`.

## Outcome

Separately from dimension scores, every run has a binary **outcome** decided by the task card's verification checks:

- `pass` — all verification checks met
- `fail` — any check not met, or stop reason other than `agent_done`

An agent's **pass rate** on a task family is the primary headline figure. Dimension scores explain *why*.

## Guarding against rater drift

- Score from the screenshots and log, not from memory of watching the run.
- When unsure between two scores, choose the lower and note the doubt in `run.md`.
- Re-score a random 10% of runs after a week; if more than one dimension differs by ≥ 2, revisit the anchors and bump the protocol version.
