DelegationBench: Measuring When AI Agents Should Ask Before Acting
Shiva Pochampally
Don't trust single-score delegation metrics. Test your agent across matched scenarios and measure consistency, not just agreement. Report phrasing sensitivity and tool-vs-judgment gaps separately. When rules matter, state them explicitly—models follow explicit rules almost perfectly.
AI agents that act autonomously—sending emails, editing files, making purchases—need to decide when to proceed and when to check with the user. Existing evaluation methods score models by showing them proposed actions and comparing their decisions to human labels.
Method: A 156-scenario benchmark with matched pairs that change single features (stakes, reversibility, who requested it) exposes three failures in current delegation scoring. A simple keyword rule outperforms eight of ten models on agreement scores, yet changes its decision in only 9 of 48 matched pairs. Equivalent phrasings shift how often models act by up to 52.5 percentage points. Every model asks less often when using tools than when judging proposed actions—despite identical stakes.
Caveats: Benchmark uses hypothetical scenarios, not real user tasks with actual consequences.
Reflections: Can models learn to maintain consistent delegation policies across equivalent phrasings without explicit rule statements? · Do tool-vs-judgment gaps persist when models have extended interaction history with a user? · What delegation consistency is acceptable for deployment—should matched pairs always yield identical decisions?