PhD Thesis Defense: Lingdong Wang, Deep Learning for Resource-Efficient Video Streaming
Content
Speaker:
Abstract:
Traditional video streaming treats the network bitstream as the primary carrier of source-specific information: a server compresses a video, the client decodes it, and improving quality generally requires transmitting more data. Deep learning changes this paradigm by introducing a learned prior at the client and making computation a controllable system resource. Although a pretrained model does not contain the target video itself, it captures general knowledge of visual appearance, semantics, motion, and geometry that can help reconstruct information omitted from the bitstream when guided by compact source-specific evidence. This dissertation investigates how learned priors and neural computation can enable resource-efficient video streaming across media representation, compression, adaptive delivery, reconstruction, and rendering. It develops methods that represent RGB and RGB-depth videos using compact semantic, visual, and geometric conditions; coordinate network transmission with client-side neural processing; concentrate reconstruction computation where it contributes most to perceived quality; and construct deployable volumetric media from readily available monocular video. Collectively, these methods reduce communication and computation burdens while improving perceptual quality, temporal consistency, and overall quality of experience. This dissertation argues that future video systems should be designed not solely around transporting compressed pixels, but around the coordinated use of transmitted evidence, knowledge stored in learned models, and adaptive scheduling.
Advisors:
Ramesh Sitaraman and Mohammad Hajiesmaili