Examples

Instead of prescribing rules which cannot possibly fit the wide heterogeneity of community goals, this wiki works with concrete examples: tools, collaborations, and approaches, each with its upsides and downsides. Through these examples, we also attempt to put terms into a concrete context. We hope these examples can be a resource for communities, researchers, and practitioners to have more productive conversations about this topic.

All
13 rows
NameGenreStanceThemesSummary
AmericasNLP Shared TasksCasePromisingCapacity-building, Data sovereigntyA recurring research competition at the workshop on NLP for Indigenous Languages of the Americas.
Chatbots With Human LikenessCaseMixedConsent and likeness, Representation and misrepresentationSystems that give a tool a real person's face, voice, or words can be powerful and personal, but also raise important questions.
Choctaw Nation and OracleCasePromisingData sovereignty, Digitization and legacy materials, Capacity-buildingA Nation-led collaboration with a commercial vendor to build a Choctaw translation model, with the language data kept under Choctaw control.
Community-Trained TranscriptionCasePromisingCapacity-building, Digitization and legacy materialsA short workshop taught language workers to build their own speech-to-text systems.
Image Generation For Language LearningCaseMixedRepresentation and misrepresentation, Consent and likeness, Data quality and reviewUsing AI image generation to make visuals for language-learning materials, like flashcards and children's books.
OpenAI and Indigenous LanguagesCaseCautionaryData sovereignty, Consent and likenessOpenAI's Whisper was trained on Māori and Hawaiian recordings scraped without community consent.
Putting Data OnlineCaseMixedData sovereignty, Digitization and legacy materialsPutting community language materials online, weighing access against protection and consent.
Scots WikipediaCaseCautionaryData quality and review, Representation and misrepresentationA non-speaker wrote a huge amount of inaccurate Scots-language Wikipedia content, producing a flawed resource that later fed AI training.
GlosbeToolMixedData quality and review, Data sovereigntyA large crowd-sourced, well-documented, multilingual dictionary that can be a valuable or risky source of data for language technology.
LLM-RBMTToolPromisingData sovereignty, Data quality and reviewA translation system that constrains a language model to a hand-built grammar via constrained decoding, so every output is guaranteed grammatical by construction.
Mukurtu CMSToolMixedData sovereignty, Sustainability and succession, Digitization and legacy materialsA free, open-source content management system built for Indigenous communities.
Ojibwe ChatToolCautionaryData quality and review, Representation and misrepresentationA public Ojibwe translation chatbot that confidently returns incorrect output.
TranskribusToolMixedDigitization and legacy materials, Data sovereigntyA widely used handwritten- and printed-text recognition platform run by a European cooperative that is strong on user ownership and EU data protection, but its cloud service is permitted to use uploaded material to improve its models.
Created · Updated
Supported By the National Science Foundation Award 2542375.