Content

Speaker:

Jaewook (Jake) Lee

Abstract:

Learning vocabulary is a key part of language learning, yet unfamiliar words and characters can be difficult to remember. Keyword mnemonics are widely used to support vocabulary retention: they link a word’s sound or a character’s visual components to familiar keywords and then connect those keywords to the target meaning through a memorable verbal or visual cue. For example, "alleviate" can be linked to the keywords "a, leaf, he, ate," evoking the image of a hungry person eating a leaf to alleviate his hunger. For the kanji 休 ("rest"), a Japanese character of Chinese origin, the character can be decomposed into the components 人 ("person") and 木 ("tree"), evoking the image of a person resting beside a tree.
 
Such mnemonics appear in books, shared flashcard decks, and online communities where learners also create and exchange their own. However, creating effective mnemonics remains time-consuming, and what works well may vary across learners. Learners may differ in the keywords they find familiar or intuitive, the connections they find memorable, and the verbal or visual cues they prefer. These preferences can guide personalization once they are known, but learners who are just beginning to create their own mnemonics may not yet know what works for them or where to start. Rules derived from how other learners connect keywords to target meanings may provide a useful starting point before individual preferences become clear. In this proposal, I investigate how large language models (LLMs) can support the development of effective and adaptive mnemonic methods for vocabulary learning.

First, I investigate how keyword mnemonics can be generated and evaluated automatically for English vocabulary learning. I develop an overgenerate-and-rank approach in which an LLM generates multiple candidate keywords and verbal cues, and the candidates are then ranked using psycholinguistically motivated criteria. I evaluate the resulting mnemonics through both automated measures of imageability and coherence and a human study with English teachers and learners. The results show that LLM-generated mnemonics are comparable to human-authored ones in imageability, coherence, and perceived usefulness. However, substantial variation in learners’ judgments suggests that effective mnemonics are highly individual, motivating personalized mnemonic generation.

Second, I investigate how mnemonic generation can be personalized to individual learners. Focusing on kanji learning, I first conduct formative interviews to examine how learners differ in their preferences for component keywords, how those keywords are combined into mnemonics, and narrative style, as well as in how they interpret the semantic and visual properties of kanji components. Based on these findings, I develop an LLM-powered interactive system that allows learners to revise component keywords, select narrative features such as personal references or wordplay, and adjust mnemonic length. I evaluate personalized LLM-generated mnemonics against human-authored mnemonics in a within-subjects user study. I find that personalized mnemonics receive higher average preference ratings, while immediate recall is comparable and delayed recall outcomes vary.

Finally, I propose an interpretable framework that learns recurring rules for mnemonic construction from mnemonics created and shared by learners in online kanji mnemonic-sharing communities. The framework will learn how learners connect component keywords to kanji meanings through mnemonic stories, while modeling how these rules vary across learners and kanji. Then, I will generate mnemonics for new learners based on these rules. I will evaluate whether the generated mnemonics align with learner-authored mnemonics and follow the learned rules for connecting component keywords to kanji meanings. 

Overall, this thesis details how LLMs can support mnemonic generation for vocabulary learning through both sound-based and shape-based associations, from automatically generating and evaluating mnemonics to personalizing them for individual learners and learning recurring construction rules from learner-authored mnemonics. Together, these approaches advance the design of adaptive mnemonic systems and suggest how LLMs can support vocabulary learning across languages and writing systems.

Advisor: 

Andrew Lan