Statistical Machine Translation
Leveraging Machine Translation for Toxicity Classification
Pages
16
Time to read
39 mins
Publication
Language
English
Pages
16
Time to read
39 mins
Publication
Language
English
This research article presents an empirical exploration of multilingual toxicity detection using machine translation (MT) systems. The study compares translation-based classification pipelines against language-specific and multilingual classifiers across 17 languages with varying resource levels. The findings indicate that translation-based methods consistently outperform out-of-distribution classifiers in 81.3% of cases, particularly benefiting low-resource languages. The analysis reveals that traditional classifiers surpass large language model (LLM) judges, especially for languages with limited resources. Additionally, the study examines the impact of MT-specific fine-tuning on LLMs, noting that while it reduces refusal rates, it may adversely affect detection accuracy for low-resource languages. The article concludes with practical recommendations for developing scalable multilingual content moderation systems, emphasizing the significance of translation quality and resource levels in enhancing toxicity detection capabilities.