Raina Heaton

Associate Professor of Native American Studies · University of Oklahoma

Raina Heaton is a linguist at the University of Oklahoma, where she is an associate professor in Native American Studies and associate curator of the Native American Languages Collection at the Sam Noble Oklahoma Museum of Natural History. Her work documents and describes endangered languages and supports revitalization, including Kaqchikel and other Mayan languages, the isolate Tunica, and Enenlhet in Paraguay.

In Their Own Words

Info

"In their own words" sections represent the voices of the people who contributed to this wiki. They are not necessarily the views of the wiki as a whole.

The harms of generative AI have often been dismissed in the name of "progress" or "inevitability," but as many people have now written, including Indigenous authors, commercial AI is both digital colonialism and, at its roots, an embodiment of the centuries-old colonial enterprise. Perpetuation of and complicity in this is not something any of us should dismiss. The list of concrete harms is long, but it includes environmental harm (and environmental racism), the spread of incorrect information, the removal of actual sources of authority from the conversation, misrepresentation, data theft from the obvious to the covert (such as the use of entered material in irretrievable training datasets), predatory publishing (junk dictionaries and the like), and the unwanted standardization and leveling of linguistic diversity.

Doing no harm is only the start. Doing some good means a lot of things, but at least: relational accountability and transparency, accurate representation, correct and vetted information both going in and coming out, consideration for environmental impacts, local control of inputs, outputs, and products, and sustainability.

For Tribes and nations, the most important step is to develop explicit policy around AI, not just for language but for everything it is being used for, and to lead the relational collaborations with industry. The groundwork for data sovereignty has to be laid at that level, because it is very difficult to negotiate or be the decision-maker for as an individual.

For learners, teachers, speakers, and enthusiasts, our networks and populations need to be educated about the risks of commercial LLMs. Most importantly for language, I'll repeat the take-away from Michael and Caroline Running Wolf's presentation at the Algonquian conference a few years back: never input your languages into commercial LLMs.

For computer scientists and industry, communities need working products that people will actually use, not proofs of concept, or comparisons of the efficiency of different models, or conceptually interesting but impractical tools. Outside academics collaborating with communities should expect to be making finished, working, real-world products; to be doing at least some language-specific development to make that happen; and to understand that figuring out what is actually needed and practical is not accomplished in one conversation. Don't assume that making a process more efficient is actually valuable or wanted: the Cherokee DAILP project wanted to create learning through the manual transcription process itself. And don't assume that making machine-readable datasets is valuable or wanted, because inputting good data still has costs.

For anyone planning or evaluating a project, the first question is whether the work actually requires an LLM or "AI" in the first place. Local models should be used wherever possible, run on locally-controlled infrastructure, and explicitly address Indigenous data sovereignty (see the CARE Principles). If it is necessary to use non-local models, applicants should expect to explicitly justify why.

It is also worth remembering how much useful collaboration with computer scientists has little to do with generative AI at all: keyboards and fonts, spellcheckers, predictive text and other collocation tools, making resources accessible (text-to-speech, "translate this webpage"), semi-automating Word-of-the-Day sharing, appropriate image generation for textbook exercises, help migrating obsolete formats and databases (like Shoebox to FLEx), orthography converters and finite-state transducers for dictionaries, pipelines for structured data, screening for personal or copyrighted material, and metadata generation and systematization for archives.

Created · Updated
Supported By the National Science Foundation Award 2542375.