Image Generation For Language Learning

Using AI image generation to make the visuals that language-learning materials need: a picture for each of thousands of vocabulary words, illustrations for a children's book, flashcards. A participant at the workshop raised it as a real possibility, since commissioning an artist for that much is infeasibile, especially for a small program.

Upsides

  • Scale and cost. Far cheaper and faster than an illustrator, which can make a picture-for-every-word feasible for a program that could not otherwise afford it.
  • Context for learners. A picture conveys a word's meaning directly, without routing through a dominant language.

Downsides

General-purpose models often lean on stereotypes, and prompt engineering is not always effective. Asked for "a Native American boy going to school," Dall-E returns a child in a feathered headdress:

imagegen-native-boy-no-prompt.png

"A Native American boy going to school," with no further instruction.

Telling it to use "casual school clothing, without traditional elements" turns the clothes a bit more casual but keeps the headdress anyway:

imagegen-native-boy-prompt-engineered.png

The same request, prompt-engineered to avoid traditional dress.

Using image generation for curriculum development has to be done carefully. Left unchecked, it can amount to misrepresentation and digital redface at scale, landing on the learners forming their first impressions of the language. The bias sits in the training data: state-of-the-art models inherit the stereotypes in the datasets they learn from, so the headdress above is not reliably a prompt away. That same training data is the deeper bind: community artists often do not want their work absorbed into these models, but leaving it out only makes the output more stereotyped, a question of ownership and consent with no easy answer.

Updated
Supported By the National Science Foundation Award 2542375.