Projects to Supervise

Dynamical Gaussian processes for Sequential Data

Supervisors: Vidhi Lalchand, Neil D. Lawrence

Gaussian process dynamical models (state-space models) builds upon a long line of work combining Gaussian processes (GP) with latent variable models for unsupervised learning tasks. Specifically, we narrow our focus on modelling high-dimensional sequential data ubiquitous in nature. The dynamics in observed space are captured by a smoothly evolving latent variable indexed by time and governed by a latent Gaussian process prior. The idea behind the project is to develop a scalable algorithm for sequential data which does not rely on holding the complete sequence in memory but can process the time-series or sequence in chunks. We also want to amortise the model with a suitable encoder like a recurrent neural network, LSTM or an Attention-based transformer.

The model is basically an extension of the model proposed here: https://gregorygundersen.com/blog/2020/07/24/gpdm/

Required reading:

[1] Gaussian process dynamical model: https://gregorygundersen.com/blog/2020/07/24/gpdm/

[2] Variational GPDM: https://proceedings.neurips.cc/paper/2011/file/af4732711661056eadbf798ba191272a-Paper.pdf

[3] Variational GPSSM: https://proceedings.neurips.cc/paper/2014/file/139f0874f2ded2e41b0393c4ac5644f7-Paper.pdf

[4] Recurrent Gaussian processes with SVI: https://arxiv.org/abs/1511.06644

[5] GPVAE for interpretable latent dynamics: http://proceedings.mlr.press/v118/pearce20a/pearce20a.pdf

Latent Probabilistic Models for Unsupervised Structure Learning in Massively Missing Data

Supervisors: Vidhi Lalchand, Neil D. Lawrence

Real-world datasets often contain entries with missing elements e.g. in a medical dataset, a patient is unlikely to have taken all possible diagnostic tests. Variational Autoencoders (VAEs) are popular generative models often used for unsupervised learning. Despite their widespread use it is unclear how best to apply VAEs to datasets with missing data.

In this projects we intend to explore a novel incarnation of a Gaussian process latent variables which can work seamlessly with missing data, the architecture would include a pointNet encoder which will encode every partial data point as a set with indicator variables (capturing the dimension identity) and a classical Gaussian process decoder.

The focus would be on assessing the quality of uncertainty calibration, structure learning in latent space, and reconstruction quality on previously unseen data.