Content

Speaker:

Hunter McNichols

Abstract:

Large language models (LLMs) are being deployed increasingly in education technologies and ushering in a new wave of AI-mediated learning. While these models demonstrate impressive capabilities in correctly answering increasingly complex questions across many fields, teaching involves far more than being correct -- it requires understanding how a person could be incorrect and guiding them towards a new understanding. In order to leverage LLM-based technologies to scale personalized education and improve learning outcomes, the pedagogical capabilities of these models need to be evaluated and improved. Furthermore, the behavior patterns of learners in AI-mediated learning settings need to be understood to properly evaluate and improve future LLM-based learning technology. In this thesis, I investigate the capabilities of LLMs to understand student behavior and provide appropriate feedback in a variety of education domains. In addition, I report the results of a year-long case study into authentic student-AI conversational behavior.

First, I introduce a flexible LLM-based method for classifying student errors in open-ended algebra problems that generalizes beyond rigid, rule-based systems. Second, I present StudyChat, a large dataset of authentic student interactions with an LLM tutor during programming assignments in a university-level computer science course, and analyze how dialogue behaviors in these interactions relate to learning outcomes. Third, I investigate whether LLMs can replicate the feedback mechanisms of an intelligent tutoring system, finding that they capture the surface form of feedback but fail to generalize to previously unseen student errors. Finally, I present the Gumbel Machine, a controllable decoding approach that generates counterfactual revisions of student writing that improve rubric scores while preserving the student's original voice. Taken together, these contributions reveal both the promise of LLMs as educational tools and the structural limitations that must be addressed for AI-mediated learning to be reliably beneficial.

Advisor:

Andrew Lan