Automation Anywhere
Enterprise AI Agent Evaluation and Benchmarking
Pages
11
Time to read
18 mins
Publication
Language
English
Pages
11
Time to read
18 mins
Publication
Language
English
This technical white paper discusses the evaluation of enterprise AI agents through two primary research components: the τ-bench performance evaluation and the implementation of Process Reasoning Engine (PRE) with Context Intelligence. The paper outlines the results of the τ-bench evaluation, where Automation Anywhere's goal-based agents, particularly aa-agent-v1, achieved the highest scores at every pass level. The evaluation framework utilized a dual-metric approach, measuring both Task Success and Trajectory Accuracy, to assess the agents' overall performance across multiple service domains, including Airline, Retail, Telecom, and Banking. Furthermore, the paper presents findings from the proprietary GBA-Bench, which assesses agents on enterprise-specific workflows, highlighting advancements in goal completion and trajectory accuracy when using context intelligence. It concludes by detailing the structure and necessity of these evaluations for ensuring reliable performance in enterprise automation tasks, emphasizing the importance of both accuracy and speed in agent execution.