PhD Dissertation Proposal: Rumeng Li, AI for Aging: Longitudinal Disease Risk Modeling from Clinical Data
Content
Speaker:
Abstract:
Population aging is increasing the burden of aging-associated chronic diseases, many of which develop gradually over years before formal diagnosis. Characterizing disease risk earlier in the disease course may support more timely clinical evaluation, monitoring, and intervention planning. Electronic health records (EHRs) provide longitudinal observations from routine clinical care, creating an opportunity to identify emerging risk patterns and track their progression over time. However, many early signals of disease progression are not captured in structured fields alone; instead, they are fragmented across years of unstructured clinical narratives, making them difficult to extract, integrate, and interpret at scale. This dissertation addresses this limitation by developing AI and clinical natural language processing (NLP) methods that transform longitudinal narrative data into scalable and interpretable models of disease risk.
Alzheimer’s disease (AD) serves as the primary application domain for this work because it is a major and growing public health challenge among aging populations. It is also often preceded by gradual cognitive, behavioral, functional, and social changes that are documented during routine clinical care. Through this application, the dissertation develops methods for longitudinal signal identification, scalable supervision, and interpretable temporal risk assessment from real-world clinical data.
First, this dissertation investigates whether early disease-related signals can be identified from longitudinal clinical notes at a population scale. Using large-scale data from the Veterans Health Administration (VHA), we demonstrate that symptom trajectories extracted from clinical narratives diverge years before formal AD diagnosis and provide complementary predictive value beyond structured EHR features alone.
Building on these signals, this dissertation explores large language model (LLM)-based approaches for scalable supervision through synthetic clinical data generation. By integrating expert clinical knowledge with real-world population structure and longitudinal disease progression patterns, we generate clinically plausible synthetic data to support scalable training of downstream models while reducing reliance on exhaustive manual annotation.
Next, this dissertation develops multi-agent reasoning frameworks for interpretable longitudinal risk assessment. By synthesizing heterogeneous evidence across time and clinical domains, these frameworks transform fragmented longitudinal observations into more cohesive, transparent, and clinically meaningful risk estimates.
Finally, this dissertation examines the generalizability of longitudinal clinical AI across diverse populations and healthcare systems. Specifically, it investigates how social and behavioral determinants of health (SBDH) documented in clinical narratives contribute to disease risk and evaluates how narrative-based models generalize across healthcare settings with different patient demographics, documentation practices, and care environments.
In summary, this dissertation establishes a framework for transforming fragmented clinical narratives into scalable and interpretable models of longitudinal disease risk. By integrating longitudinal signal extraction, scalable supervision, and temporal reasoning over real-world clinical data, this work advances AI methods for earlier and more clinically grounded risk modeling. Although validated primarily in AD, the proposed methods are designed to generalize to broader aging-related and chronic disease settings.
Advisor:
Hong Yu