Shift Technology
Comparison of Large Language Models in Insurance
Pages
5
Time to read
6 mins
Publication
Language
English
Pages
5
Time to read
6 mins
Publication
Language
English
This technical report presents a comparison of various Large Language Models (LLMs) applied to specific insurance use cases. The document outlines the methodology used by Shift Technology's data science and research teams, which devised four test scenarios to evaluate the performance of 16 publicly available LLMs. These scenarios include information extraction from airline invoices, property repair quotes, dental invoices, and document classification related to travel insurance claims. The report details the evaluation metrics of coverage and accuracy, highlighting how the models performed across different tasks. Additionally, it discusses the advancements in LLM technology, introducing six new models while removing two from evaluation. The findings indicate that GPT4o achieved the highest aggregate performance score, followed closely by GPT4o-Mini and Claude3.5 Sonnet. The report concludes that as LLM technology evolves, price may become a critical factor in model selection, particularly until more significant performance differences are established among the models tested.