Technical University of Munich
Automated Evaluation Framework for BIM-QA Systems
Pages
8
Time to read
26 mins
Publication
Language
English
Pages
8
Time to read
26 mins
Publication
Language
English
This document is a research article that presents an automated evaluation framework designed for Building Information Modeling Question Answering (BIM-QA) systems. The framework addresses the lack of standardized methods for cross-study comparison in the BIM-QA field, which has been hindered by the use of custom benchmarks and non-shared datasets. The authors introduce a multi-dimensional evaluation pipeline that utilizes large language model (LLM) judges to assess BIM-QA answers based on five quality criteria. The study validates the framework through inter-rater agreement analysis involving both LLM judges and human experts on a set of question-answer pairs. Results indicate that LLM judges achieve higher inter-rater reliability compared to human experts, suggesting that this automated approach can facilitate standardized evaluation across different BIM-QA systems. The paper outlines the necessity for a reproducible evaluation method to enhance trust in LLM-based systems for professional use and discusses the implications of these findings for future research in the field.