Statistical Machine Translation
Zero-Shot Joint Decoding Method for Multimodal Translation
Pages
15
Time to read
45 mins
Publication
Language
English
Pages
15
Time to read
45 mins
Publication
Language
English
This technical report presents a novel zero-shot ensembling strategy for integrating diverse models during the decoding phase in machine translation tasks. The approach addresses the challenges posed by differences in vocabularies among models, which limit the effectiveness of traditional ensemble methods. The proposed method allows for word-level re-ranking during decoding without requiring additional training or task-specific data. It enables the ranker model to influence the decoding process by predicting when a word is completed, thereby improving translation quality in multimodal scenarios. The report outlines the algorithm's design, which includes an online re-ranking mechanism that enhances the integration of information from different models. The effectiveness of the method is demonstrated through experiments, indicating that it successfully combines the strengths of translation and vision models, resulting in context-aware translations that are more accurate. The findings contribute to the field of natural language processing by providing a flexible solution for multimodal translation tasks.