Data Scraping

Collecting material automatically from across the web and other digital sources instead of gathering it by hand. Modern AI makes it far more capable than it used to be: agentic systems that navigate many sources and pull what they find, semantic search that locates material by meaning rather than exact keywords, and summarization that turns a large pile of recovered text into something a person can work through. Whether it helps or harms depends almost entirely on whose material is being collected, for what, and with whose involvement.

The language documentation field has a long history of extractive research, so a great deal of language material already sits outside the communities it came from, scattered across universities, libraries, and the web, often unknown to the people whose language it documents (see Jared Coleman's experience recovering a grammar built from his great-grandmother's words). Scraping can help find that material and make it usable for the community it is about. The same automation can also repeat the extraction, gathering material without a community's involvement or knowledge, which is a data sovereignty and consent question however the collecting is done (see OpenAI and Indigenous Languages).

Two risks follow once scraped data is in hand. A scraped collection can look authoritative while being, in places, simply wrong: large web corpora advertise a "clean" subset per language, but for lower-resource languages that subset is often still full of errors. And bad data spreads, since a scraped dataset can become training data for the next model; where there are few speakers to catch an error, misinformation can stick and propagate (see Scots Wikipedia).

Created · Updated
Supported By the National Science Foundation Award 2542375.