Machine Translation

Using software to translate text from one language into another. Modern systems are built with machine learning, usually trained on large collections of parallel text — the same content in two languages — from which the system learns to map one to the other. For widely spoken languages with abundant parallel data, the results can be fluent and useful.

For low-resource languages, including most Indigenous languages, the picture is harder. There is little parallel text to learn from, and many of these languages have grammatical features rarely seen in the languages NLP is built around, so quality is often much lower. Because the output is probabilistic, a translation can read smoothly while being wrong — confidently producing something that is not what the source said, which is especially risky when few readers can catch the error. For this reason machine translation is usually best treated as a draft that a fluent speaker reviews, rather than a finished product.

It is also worth asking what the translation is for and who it serves. A system that translates into a community's language is a different proposition from one that translates out of it, and the data needed to build either raises data sovereignty and consent questions. Machine translation is the central task in the AmericasNLP Shared Tasks, which take up exactly these challenges for Indigenous languages of the Americas.

Updated
Supported By the National Science Foundation Award 2542375.