Statistical Machine Translation
Evaluation of Automatic Metrics for Machine Translation
Pages
35
Time to read
86 mins
Publication
Language
English
Pages
35
Time to read
86 mins
Publication
Language
English
This technical report presents the findings from the WMT24 Metrics Shared Task, which focused on evaluating the performance of automatic metrics for machine translation (MT). The task specifically assessed LLM-based translations generated during the WMT24 General MT Shared Task. The evaluation aimed to determine the effectiveness of existing metrics in accurately assessing these translations. Human assessments were conducted using Multidimensional Quality Metrics (MQM) to provide a robust benchmark. The report details the methodology, including the collection of human quality ratings and the introduction of a challenge set subtask designed to test the metrics' ability to identify various translation errors. The analysis covers three language pairs: English to Spanish, Japanese to Chinese, and English to German, and discusses the performance of fine-tuned neural metrics. The findings indicate that these metrics continue to perform well, even for LLM-based translations, and highlight the importance of testing metrics across diverse linguistic phenomena and domains.