NanoNets
Visual Caption Restoration Task and Dataset
Pages
22
Time to read
65 mins
Publication
Language
English
Pages
22
Time to read
65 mins
Publication
Language
English
This document is a conference paper that introduces the Visual Caption Restoration (VCR) task, which challenges vision-language models to restore partially obscured texts within images using pixel-level hints. The paper outlines the limitations of existing approaches that primarily rely on optical character recognition (OCR) and masked language modeling, which do not effectively address the complexities of text embedded in images. The authors develop a synthetic image generation pipeline to create a dataset called VCR-WIKI, consisting of 2.11 million English and 346,000 Chinese entities. The dataset features varying levels of caption visibility to adjust task difficulty. Empirical evaluations reveal that current vision-language models significantly underperform compared to human benchmarks in the VCR task, highlighting the need for innovative model architectures. The paper aims to stimulate further research in the field by releasing the VCR-WIKI dataset and accompanying code, facilitating advancements in understanding and interpreting multimedia content.