Statistical Machine Translation
Shortcomings of LLMs for Low-Resource Translation
Pages
23
Time to read
61 mins
Publication
Language
English
Pages
23
Time to read
61 mins
Publication
Language
English
This technical report investigates the capabilities of pretrained large language models (LLMs) in translating text from a low-resource language, specifically Southern Quechua, into a high-resource language, Spanish. The study conducts a series of experiments utilizing various types of context retrieved from a constrained database of pedagogical materials, including dictionaries and grammar lessons. The report outlines the methods used for evaluation, including both automatic and human assessments of model outputs, and details the ablation studies that manipulate context type, retrieval methods, and model types. Findings indicate that while smaller LLMs can utilize prompt context for zero-shot low-resource translation, the effectiveness varies based on context type and retrieval method. The report also discusses the limitations of LLMs in translation tasks for the majority of the world’s languages, highlighting ethical concerns and the challenges faced in developing effective low-resource machine translation systems.