OpenAI and Indigenous Languages

Whisper is OpenAI's open-weight speech-to-text model. Its training data included recordings in Indigenous languages: by OpenAI's own paper, 1,381 hours of te reo Māori and 338 hours of ʻōlelo Hawaiʻi, collected by scraping audio from the web. No Māori or Hawaiian people were involved in building it, and the communities whose recordings were used were not asked.

In January 2023, Te Hiku Media (a Māori organization that builds its own language technology) published a critique titled "OpenAI's Whisper is another case study in Colonisation" (Keoni Mahelona, Gianna Leoni, Suzanne Duncan, and Miles Thompson). They described the use of te reo Māori as taonga (treasure) taken without consent, and said they would never put a model like Whisper into production because doing so would violate their Kaitiakitanga License and the trust of the people who contributed recordings.

This is the same extractive pattern as scraping in general, applied to languages communities have spent generations working to revitalize. For a people who survived the suppression of their language, finding it ingested into a Silicon Valley product without a say is a familiar kind of harm.

What went wrong

  • Collecting is not consent. The recordings were taken from the web without the knowledge or permission of the communities they came from. Who decides what gets collected, where it goes, and what it is used for is a data sovereignty question, and here the communities were not part of the decision.
  • No community in the loop. A model can be released as "open" and still be built entirely without the people whose data it depends on. Openness is not the same as accountability to the source community.
  • The output can mislead. On te reo Māori, Whisper's word error rate was far worse than a community-built model (about 73% versus 38% for Te Hiku Media's own system in their reported comparison; IEEE Spectrum). A tool that looks like it "supports" a language while getting most of it wrong can do damage of its own.

Silver linings

  • Communities can build the alternative, on their own terms. Te Hiku Media trains its models only on material contributed with full consent, draws on an archive of 30-plus years of recordings, and lets contributors keep ownership of their data (NVIDIA).
  • Some of OpenAI's funding has gone to Native-led work. Through its People-First AI Fund grantees include Native-led organizations such as the Tribal Education Departments National Assembly, which runs AI-literacy work tied to tribal sovereignty.
Created · Updated
Supported By the National Science Foundation Award 2542375.