What's in a Word—Breaking Texts and Making Data
Event Details
Date
Tuesday, October 20, 2026
Time
1-2:30 p.m.
Location
231 Memorial Library
Description
Before text can be analyzed computationally, it has to be taken apart — split into pieces, normalized, stripped of noise. This session introduces the preprocessing behind nearly all text analysis: tokenization, lemmatization, stemming, and stop word removal. We'll also distinguish this sense of "token" from how the word is used in large language models, a difference that reveals a lot about both. You'll leave understanding not just how to prepare text, but what each choice quietly decides.
Cost
Free
Contact
Accessibility
We value inclusion and access for all participants and are pleased to provide reasonable accommodations for this event. Please email brady.krien@wisc.edu to make a disability-related accommodation request. Requests should be made by Tuesday, October 6, 2026, though reasonable effort will be made to support late accommodation requests.