NanoNets
Rethinking Document Information Extraction Datasets
Pages
24
Time to read
65 mins
Publication
Language
English
Pages
24
Time to read
65 mins
Publication
Language
English
This technical report presents K2Q, a new collection of datasets aimed at improving key information extraction (KIE) tasks within visually rich document understanding (VRDU). The report outlines the limitations of existing prompt-response datasets, which often rely on simplistic and uniform templates that do not adequately reflect the complexity of real-world queries. K2Q addresses this need by converting five existing KIE datasets into a diverse prompt-response format with over 300,000 questions. This method employs a suite of more than 100 bespoke templates tailored to specific KIE datasets, thereby enhancing the complexity and robustness of the generated questions. The report provides empirical comparisons of generative model performance on K2Q against simpler templates, demonstrating a significant improvement in model performance with the diverse templates. The findings advocate for future research on data quality and the training methodologies of generative models, emphasizing the essential role that high-quality datasets play in advancing VRDU capabilities.