Examples
Instead of prescribing rules which cannot possibly fit the wide heterogeneity of community goals, this wiki works with concrete examples: tools, collaborations, and approaches, each with its upsides and downsides. Through these examples, we also attempt to put terms into a concrete context. We hope these examples can be a resource for communities, researchers, and practitioners to have more productive conversations about this topic.
All
| Name | Genre | Stance | Themes | Summary |
|---|---|---|---|---|
| AmericasNLP Shared Tasks | Case | Promising | Capacity-building, Data sovereignty | A recurring research competition at the workshop on NLP for Indigenous Languages of the Americas. |
| Chatbots With Human Likeness | Case | Mixed | Consent and likeness, Representation and misrepresentation | Systems that give a tool a real person's face, voice, or words can be powerful and personal, but also raise important questions. |
| Choctaw Nation and Oracle | Case | Promising | Data sovereignty, Digitization and legacy materials, Capacity-building | A Nation-led collaboration with a commercial vendor to build a Choctaw translation model, with the language data kept under Choctaw control. |
| Community-Trained Transcription | Case | Promising | Capacity-building, Digitization and legacy materials | A short workshop taught language workers to build their own speech-to-text systems. |
| Image Generation For Language Learning | Case | Mixed | Representation and misrepresentation, Consent and likeness, Data quality and review | Using AI image generation to make visuals for language-learning materials, like flashcards and children's books. |
| OpenAI and Indigenous Languages | Case | Cautionary | Data sovereignty, Consent and likeness | OpenAI's Whisper was trained on Māori and Hawaiian recordings scraped without community consent. |
| Putting Data Online | Case | Mixed | Data sovereignty, Digitization and legacy materials | Putting community language materials online, weighing access against protection and consent. |
| Scots Wikipedia | Case | Cautionary | Data quality and review, Representation and misrepresentation | A non-speaker wrote a huge amount of inaccurate Scots-language Wikipedia content, producing a flawed resource that later fed AI training. |
| Glosbe | Tool | Mixed | Data quality and review, Data sovereignty | A large crowd-sourced, well-documented, multilingual dictionary that can be a valuable or risky source of data for language technology. |
| LLM-RBMT | Tool | Promising | Data sovereignty, Data quality and review | A translation system that constrains a language model to a hand-built grammar via constrained decoding, so every output is guaranteed grammatical by construction. |
| Mukurtu CMS | Tool | Mixed | Data sovereignty, Sustainability and succession, Digitization and legacy materials | A free, open-source content management system built for Indigenous communities. |
| Ojibwe Chat | Tool | Cautionary | Data quality and review, Representation and misrepresentation | A public Ojibwe translation chatbot that confidently returns incorrect output. |
| Transkribus | Tool | Mixed | Digitization and legacy materials, Data sovereignty | A widely used handwritten- and printed-text recognition platform run by a European cooperative that is strong on user ownership and EU data protection, but its cloud service is permitted to use uploaded material to improve its models. |
Tools
| Name | Technology | Stance | Summary |
|---|---|---|---|
| Glosbe | Dictionary | Mixed | A large crowd-sourced, well-documented, multilingual dictionary that can be a valuable or risky source of data for language technology. |
| LLM-RBMT | MT, LLM | Promising | A translation system that constrains a language model to a hand-built grammar via constrained decoding, so every output is guaranteed grammatical by construction. |
| Mukurtu CMS | Archive/CMS | Mixed | A free, open-source content management system built for Indigenous communities. |
| Ojibwe Chat | MT, LLM | Cautionary | A public Ojibwe translation chatbot that confidently returns incorrect output. |
| Transkribus | OCR | Mixed | A widely used handwritten- and printed-text recognition platform run by a European cooperative that is strong on user ownership and EU data protection, but its cloud service is permitted to use uploaded material to improve its models. |
Cases & patterns
| Name | Themes | Stance | Summary |
|---|---|---|---|
| AmericasNLP Shared Tasks | Capacity-building, Data sovereignty | Promising | A recurring research competition at the workshop on NLP for Indigenous Languages of the Americas. |
| Chatbots With Human Likeness | Consent and likeness, Representation and misrepresentation | Mixed | Systems that give a tool a real person's face, voice, or words can be powerful and personal, but also raise important questions. |
| Choctaw Nation and Oracle | Data sovereignty, Digitization and legacy materials, Capacity-building | Promising | A Nation-led collaboration with a commercial vendor to build a Choctaw translation model, with the language data kept under Choctaw control. |
| Community-Trained Transcription | Capacity-building, Digitization and legacy materials | Promising | A short workshop taught language workers to build their own speech-to-text systems. |
| Image Generation For Language Learning | Representation and misrepresentation, Consent and likeness, Data quality and review | Mixed | Using AI image generation to make visuals for language-learning materials, like flashcards and children's books. |
| OpenAI and Indigenous Languages | Data sovereignty, Consent and likeness | Cautionary | OpenAI's Whisper was trained on Māori and Hawaiian recordings scraped without community consent. |
| Putting Data Online | Data sovereignty, Digitization and legacy materials | Mixed | Putting community language materials online, weighing access against protection and consent. |
| Scots Wikipedia | Data quality and review, Representation and misrepresentation | Cautionary | A non-speaker wrote a huge amount of inaccurate Scots-language Wikipedia content, producing a flawed resource that later fed AI training. |
Cautionary tales
| Name | Genre | Themes | Summary |
|---|---|---|---|
| Ojibwe Chat | Tool | Data quality and review, Representation and misrepresentation | A public Ojibwe translation chatbot that confidently returns incorrect output. |
| OpenAI and Indigenous Languages | Case | Data sovereignty, Consent and likeness | OpenAI's Whisper was trained on Māori and Hawaiian recordings scraped without community consent. |
| Scots Wikipedia | Case | Data quality and review, Representation and misrepresentation | A non-speaker wrote a huge amount of inaccurate Scots-language Wikipedia content, producing a flawed resource that later fed AI training. |