R
B2B Reorder Agent

Evaluations

Autonomy is earned, not assumed. Fourteen gold cases spanning all six outcomes run against the live agent in shadow mode: real tools, real guardrails, no stock movement, no buyer messages. This gate runs before any prompt, model or policy change ships.

The same 14 gold cases, run against three models on 17 Aug 2026.

ModelPass rateAvg timeAvg costWhat happened
claude-sonnet-4-6 SHIPPED14/14~32s~$0.033Follows the specification. The boundary case (11) held across three consecutive runs.
claude-opus-4-813/14~27s~$0.030Failed case 11 by overriding the policy in force: made its own judgement call on an ambiguous request instead of following the specified behaviour.
claude-haiku-4-511/14~23s~$0.027Called tools with wrong inputs and treated the empty results as facts (cases 1, 11); one run died mid-flight (14).

Why Sonnet ships: Haiku fails on basics - wrong tool inputs become false facts, untrustworthy at any price. Opus is capable but argues - when it met ambiguity, it substituted its own judgement for the stated policy; in this product, policy belongs to the seller, not the model. Sonnet does the job as written.

Cost estimates price all input tokens at the uncached rate; with prompt caching enabled, true cost is at or below these figures.

Execution is disabled on this demo. Want to see it run live? Contact me at anutishri.com