Scots Wikipedia
For years, a large fraction of the Scots-language Wikipedia was written and edited by a single person who did not actually speak Scots. They worked by translating English roughly one word at a time using a dictionary, producing text that looked like Scots but was, to fluent speakers, a distorted approximation of it. The scale was enormous (tens of thousands of articles) and it went largely unnoticed until a reader documented it publicly in 2020 (The Guardian).
The most immediate and obvious harm is to people learning the language, but this also has implications for AI. Scots Wikipedia was part of the training data for widely used multilingual models, including Google's mBERT, which learned from the top ~100 Wikipedias by size (Scots among them).
It is not surprising that models trained on flawed data produce flawed output. When a 2025 MIT Technology Review investigation tested Google Translate and ChatGPT on languages like Fulfulde and Greenlandic, it found clear failures, such as a wrong month name in Fulfulde and an inability to count to ten in Greenlandic. It's not clear to what extent the flaws in Wikipedia are responsible for these models' poor performance, but for many of these languages, Wikipedia is the main and sometimes the only substantial source of training text.
What went right
- The community caught it and mobilized. Scots speakers surfaced the problem, documented it, and generated a wide public discussion about how to prevent it from recurring.
What went wrong
- An outsider became the de facto authority on a language they did not speak. Good intentions didn't prevent harm: the sheer volume of one non-speaker's work crowded out and misrepresented the language.
- The flawed data fed straight into AI models. Nothing checked it against fluent-speaker judgment before models like mBERT trained on it.