Digits
Evaluation of LLMs and Digits AGL for Accounting
Pages
15
Time to read
19 mins
Publication
Language
English
Pages
15
Time to read
19 mins
Publication
Language
English
This technical report presents a comprehensive evaluation of Digits’ specialized machine learning architecture, Digits Agentic General Ledger® (Digits AGL®), in comparison to 14 frontier general-purpose large language models (LLMs) for accounting tasks. The study assesses the accuracy, latency, reliability, and token usage of these models, focusing on their performance in transaction categorization. Key findings indicate that while no general-purpose LLM achieved greater than 81% accuracy in one-shot classification, several models surpassed human accuracy with zero hallucinations. The report details the methodology used, including a dataset of 2,000 financial transactions and the establishment of a human baseline through a majority-vote system among outsourced accountants. It also outlines the operational costs associated with using frontier LLMs in conjunction with agent harness architectures. The findings highlight the significant advantages of Digits AGL® over evaluated LLMs, particularly in terms of accuracy and operational efficiency.