Evaluations
Autonomy is earned, not assumed. Fourteen gold cases spanning all six outcomes run against the live agent in shadow mode: real tools, real guardrails, no stock movement, no buyer messages. This gate runs before any prompt, model or policy change ships.
The same 14 gold cases, run against three models on 17 Aug 2026.
| Model | Pass rate | Avg time | Avg cost | What happened |
|---|---|---|---|---|
| claude-sonnet-4-6 SHIPPED | 14/14 | ~32s | ~$0.033 | Follows the specification. The boundary case (11) held across three consecutive runs. |
| claude-opus-4-8 | 13/14 | ~27s | ~$0.030 | Failed case 11 by overriding the policy in force: made its own judgement call on an ambiguous request instead of following the specified behaviour. |
| claude-haiku-4-5 | 11/14 | ~23s | ~$0.027 | Called tools with wrong inputs and treated the empty results as facts (cases 1, 11); one run died mid-flight (14). |
Why Sonnet ships: Haiku fails on basics - wrong tool inputs become false facts, untrustworthy at any price. Opus is capable but argues - when it met ambiguity, it substituted its own judgement for the stated policy; in this product, policy belongs to the seller, not the model. Sonnet does the job as written.
Cost estimates price all input tokens at the uncached rate; with prompt caching enabled, true cost is at or below these figures.