PhD Dissertation Proposal: Bryn Reimer, Machine Learning for Molecules
Content
Speaker:
Abstract:
Leveraging biological knowledge to improve model performance
I am interested in how we can use publicly available data to improve models for predicting drug resistance in Mycobacterium tuberculosis. However, recent results in transcriptomics have shown that deep-learning-based foundation models do not yet outperform simple linear baselines with carefully chosen input data (Ahlmann-Eltze, Huber, and Anders, Nature Methods 2025; Wong et al, Bioinformatics 2025). In my review paper (currently in preparation), I perform a cross-bacterial-species analysis of experimentally-validated genetic mechanisms of resistance in order to find biological hypotheses to drive genomic feature inclusion criteria. This hypothesis-driven work has already led to our discovery of a new drug resistance associated gene in Mycobacterium tuberculosis, rlmN (Computational and Structural Biotechnology Journal, 2026).
Understanding the limitations of deep learning models in sparse data spaces
Computational chemists refer to the space of all possible molecules as chemical space, and with a small number of reasonable assumptions about drug-likeness, the drug-like chemical space is estimated to comprise approximately 1060 molecules. Obviously, despite vast amounts of human effort to characterize chemical space and understand the properties and activities of small molecules, only a fraction of a fraction of molecules have been assessed on only a fraction of a fraction of all possible endpoints. The work of the computational chemist in a drug design context is to try to work from incomplete data to predict which drug-like molecules will have favorable properties and activities, thereby yielding a safe and effective therapeutic. My work with my undergraduate supervisee, Ivy, focuses on how we can assess generalizability for deep learning models that predict chemical properties --- where generalizability refers to the performance of deep learning models when assessed on molecules that are “unlike” the training set, by some measure. Current methods attempt to use various data-splitting mechanisms, such as scaffold splitting or UMAP splitting, to simulate out-of-distribution prediction performance. In our work (currently in preparation), we show that common data-splitting mechanisms that are thought to be a good proxy for out-of-distribution performance in fact still have significant cross-split overlap between the test and train distributions.
Advisor:
Anna Green