FIZ Karlsruhe
End-to-end Information Extraction from Archival Records
Pages
9
Time to read
38 mins
Publication
Language
English
Pages
9
Time to read
38 mins
Publication
Language
English
This technical report focuses on the challenges and advancements in semi-structured document understanding, particularly regarding key information extraction (KIE) from historical archival records. It highlights the limitations of traditional information extraction methods, which often rely on manual processes, making them inefficient for the diverse and complex nature of archival documents, such as digitized index cards. The report introduces BZKOpen, an annotated dataset for enhancing KIE in historical contexts. It evaluates the performance of various Multimodal Large Language Models (MLLMs) in extracting information from these records, specifically assessing models like InternVL2.0, InternVL2.5, and GPT-4o-mini. The findings indicate that larger models do not always yield better performance, with the open-source InternVL2.5-38B achieving the best results. The report also discusses prompt engineering strategies for optimizing model performance in real-world applications and emphasizes the need for more diverse datasets to further explore the capabilities of MLLMs in key information extraction.