Projects to Supervise

Monitoring Aerodynamic Performance Real Time for Cycling

Supervisor: Neil D. Lawrence

A cyclist’s aerodynamic position has a very strong affect on their performance. In professional cycling, extensive wind tunnel testing is used to hone a cyclist’s performance. Such testing is, however, highly expensive. In this project you will use machine learning techniques alongside the physics of cycling to estimate the aerodynamic performance of a cyclist real time whilst on a bicycle. By combining an anemometer, a power meter and an understanding of the rider’s kinetic and gravitational potential energy the power loss due to aerodynamic drag can be estimated. Software for the project will be written according to the principles of open data science.

Note that field experiments will require access to a road bicycle (power loss due to rolling resistance on a mountain bicycle is too large) and some form of GPS device (for preliminary experiments a smart phone is likely sufficient).

This project will suit students with strong analytical (mathematical) skills.

Projects on Agentic System Implementation (theme)

Supervisors: Christian Cabrera Jojoa, Neil D. Lawrence

Prerequisite: L172 Information, Energy and Intelligence (IEI), or equivalent preparation in information theory, maximum entropy, and information geometry.

The research group is interested in how we engineer multi-agent systems for reliability and transparancy. We have developed the DOAgent framework for delivering such multi-agent systems. This list of Part III / MPhil projects apply information-topographic ideas to engineered multi-agent and organisational systems using DOAgent. The projects listed serve as ideas, we will run a maximum of two or three projects in this space.

Projects on Information Topography (theme)

Supervisor: Neil D. Lawrence

Prerequisite: L172 Information, Energy and Intelligence (IEI), or equivalent preparation in information theory, maximum entropy, and information geometry.

Information topography is the geometry of how information flows through a complex system. These Part III / MPhil projects develop the mathematical foundations of that geometry — from Fisher conductance and open inaccessible games to geometries of agency and good-regulator conditions — with L172 IEI as the shared prerequisite. Related group programme: Information Topography.

An Academic Scholar System

Supervisors: Eric Meissner, Neil D. Lawrence

Academic publications systems such as Microsoft Academic and Google Scholar track individual authors, their papers and who they’ve cited. Semantic scholar makes its data freely available. In this project you will absorb the semantic scholar data to provide a scalable cloud based scholar service.

Animation by Machine Learning with Motion Capture Data

Supervisor: Neil D. Lawrence

This project is about using Machine Learning to model human motion for animation, with a particular focus on the demands of computer games. The idea is to learn what natural motion looks like, and then combine it with constraints to develop an animation sequence. The constraints could be animator imposed, or imposed by the computer game. A typical scenario might be that the player’s character is required to interact with a character in the game, for example a player might be given an object in the game. The constraint could be that the hands of the player touch the hands of the character giving the object. With the current approach to animation (looking up a library of motion capture data) such a constraint is very difficult to fulfill as the required motion won’t exactly match a sequence in the library. By modelling natural motion through machine learning, we should be able to generate a new sequence to satisfy the constraint.

The models of motion will be developed using Gaussian processes. In particular the project will make use of the “Gaussian Process Latent Variable Model” (Lawrence, 2003) which has already shown a lot of promise in this domain, and “Latent Force Models” (Alvarez et al, 2009) a recently developed approach to learning based on physics and probabilistic models.

The project will involve a large amount of mathematics, in particular advanced linear algebra and calculus.

Automatic Discovery of Trade-off Between Accuracy, Privacy and Fairness for ML models

Supervisors: Andrei Paleyes, Neil D. Lawrence

When machine learning models are deployed to solve real world problems, they are often trained on sensitive data, e.g. healthcare or financial records. Practitioners need to ensure fairness and privacy of the resulting model. Often privacy and fairness guarantees may only be achieved through sacrificing accuracy (as classically measured). Usually both privacy and fairness are set as fixed constraints, and the exact effect of such constraints on accuracy is unclear. This project proposes to develop a procedure of automatic discovery of the trade-off between these three metrics.

Automatic Discovery of Trade-off Between Utility and Energy Comsumption of ML models

Supervisors: Andrei Paleyes, Neil D. Lawrence

The use of machine learning models both in academia and industry is on the rise. And so is the environmental impact of ML models. While deploying a model to production, it is important to be able to estimate its carbon footprint and to understand the costs involved in running it. Balancing performance and carbon efficiency of ML models becomes critical to ensure that the benefits of ML are maximized while minimizing its environmental costs. This project proposes to develop a procedure that automatically discovers and quantifies this trade-off.

Data Driven Mechanistic Models for Systems Biology

Supervisor: Neil D. Lawrence

Decision Support for Cities Management

Supervisors: Christian Cabrera Jojoa, Neil D. Lawrence

Smart cities are urban environments where computing devices generate considerable amounts of heterogeneous data. Cities’ authorities need sophisticated platforms to manage and analyse such data before using it. In this project you will provide a flexible, scalable, and real time decision-support tool for city managers. This tool will manage data from heterogeneous sources and provide meaningful insights to city managers to support their decisions.

Dynamical Gaussian processes for Sequential Data

Supervisors: Vidhi Lalchand, Neil D. Lawrence

Gaussian process dynamical models (state-space models) builds upon a long line of work combining Gaussian processes (GP) with latent variable models for unsupervised learning tasks. Specifically, we narrow our focus on modelling high-dimensional sequential data ubiquitous in nature. The dynamics in observed space are captured by a smoothly evolving latent variable indexed by time and governed by a latent Gaussian process prior. The idea behind the project is to develop a scalable algorithm for sequential data which does not rely on holding the complete sequence in memory but can process the time-series or sequence in chunks. We also want to amortise the model with a suitable encoder like a recurrent neural network, LSTM or an Attention-based transformer.

The model is basically an extension of the model proposed here: https://gregorygundersen.com/blog/2020/07/24/gpdm/

Required reading:

[1] Gaussian process dynamical model: https://gregorygundersen.com/blog/2020/07/24/gpdm/

[2] Variational GPDM: https://proceedings.neurips.cc/paper/2011/file/af4732711661056eadbf798ba191272a-Paper.pdf

[3] Variational GPSSM: https://proceedings.neurips.cc/paper/2014/file/139f0874f2ded2e41b0393c4ac5644f7-Paper.pdf

[4] Recurrent Gaussian processes with SVI: https://arxiv.org/abs/1511.06644

[5] GPVAE for interpretable latent dynamics: http://proceedings.mlr.press/v118/pearce20a/pearce20a.pdf

Edge Service Placement based on Large Language Models (LLMs)

Supervisors: Christian Cabrera Jojoa, Neil D. Lawrence

Edge computing architectures propose placing software services closer to end users. This distributed placement can enable super-low latency, data-intensive applications that can benefit domains as diverse as virtual reality, gaming, and healthcare. The decision of what services to deploy in which edge is an optimisation problem called service placement. Solutions to the service placement problem must consider latency requirements and resource constraints while assigning services to edge servers in an automatic fashion. Exact, approximation, heuristics, and meta-heuristic algorithms are traditional approaches to solving such an optimisation problem. This project proposes to explore the capabilities of Large Language Models (LLMs) to make the placement decisions. The main idea is to replace current algorithms with a LLM-based agent.

Gesture Recognition using Kinect and Python

Supervisor: Neil D. Lawrence

Kinect cameras provide true image and an associated depth image. In this project the focus will be on data from the Gesture Recognition Challenge for kinect: http://www.kaggle.com/c/GestureChallenge/. The student will participate in the challenge using state of the art machine learning techniques with the assistance of the Sheffield Machine Learning group. A gesture recognizer for the Kinect would enable a large range of new interfaces between the human and computer. Software for the project will be written according to the principles of open data science.

This project will suit students with strong analytical skills, there will be a focus on linear algebra and probabilistic inference in the software.

Homeostatic Regulators in Multi-Agent Environments

Supervisors: Joery de Vries, Neil D. Lawrence

Prerequisite: L172 Information, Energy and Intelligence (IEI), or equivalent preparation in information theory, maximum entropy, and information geometry.

Conant and Ashby’s good regulator theorem concerns a regulator that holds an outcome steady against a disturbance. Its success criterion is the entropy of the outcome, which is different from the classical notion of reward in reinforcement learning: it is a concave objective over the occupancy polytope, so an optimal single regulator is deterministic. This project asks what happens when the disturbance is another regulator. Several agents share an environment and each minimises the entropy of its own outcome under its own reference measure. From any agent’s viewpoint the other agents are structured, adaptive disturbances. Refinements of the theorem, notably Wentworth’s, say the regulator must carry a posterior over its disturbance, thus the notion of “model” that Conant and Ashby’s theorem implies is a posterior over the other agents’ policies. We will try to answer whether this posterior is necessary, and whether the joint problem is Nash. The working hypothesis is that competing regulators partition the state space into per-agent stable niches, which remains to be verified experimentally.

How fair is fairness? A multiverse analysis of fairness definitions

Supervisors: Samuel J. Bell, Neil D. Lawrence

As machine learning is increasingly deployed into high-stakes settings such as healthcare and justice, it’s essential our models behave in a fair and unbiased manner. Unfortunately this isn’t the case, as evidenced by a growing number of papers showing significant bias and disparate performance (e.g. [1-3]). These failures demonstrate the importance of auditing both datasets and models to identify biases, harmful associates, and performance gaps across groups.

However, in order to audit, we have a series of choices to make about how we measure fairness [4]. Do we define fairness as parity betweens similar individuals, or between groups based on demographic attributes? Do we want parity of accuracy, or loss, false negatives, or false positives? Which dataset do we evaluate on? If we are using demographic data, how are we defining socially-constructed terms like race; are we relying on self-identified or perceived definitions of gender? We could go on, and each of these decisions may impact the resulting outcome of the fairness audit.

First introduced in psychology, the multiverse analysis [5] is a principled framework for evaluating the robustness and reproducibility of claims in light of the multitude of possible choices a researcher can make along the way [6]. The multiverse entails systematically re-evaluating our analyses at each leaf of the decision tree, then transparently reporting and visualizing the results. To make this tractable in a machine learning search space Bell et al. [6] rely on a Gaussian Process surrogate and Bayesian experimental design for efficient exploration of the multiverse of choices.

This project will apply a multiverse analysis to machine learning fairness, and systematically evaluate the conclusions of previous fairness audits in light of different definitions of fairness, different evaluation protocols, and different evaluation datasets. There is also scope to extend the project to a dataset audit or to a fairness intervention (e.g. gDRO [7]).

If you’re interested in the project, get in touch with Samuel Bell (sjb326@cam.ac.uk) for an informal chat.

[1] Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. FAccT.

[2] De Vries, T., Misra, I., Wang, C., & Van der Maaten, L. (2019). Does object recognition work for everyone?. CVPR.

[3] Bolukbasi, T., Chang, K. W., Zou, J. Y., Saligrama, V., & Kalai, A. T. (2016). Man is to computer programmer as woman is to homemaker? Debiasing word embeddings. NeurIPS.

[4] Jacobs, A., Z., & Wallach, H. (2021). Measurement and Fairness. FAccT.

[5] Steegen, S., Tuerlinckx, F., Gelman, A., & Vanpaemel, W. (2016). Increasing transparency through a multiverse analysis. Perspectives on Psychological Science.

[6] Bell, S. J., Kampman, O. P., Dodge, J., & Lawrence, N. D. (2022). Modeling the Machine Learning Multiverse. NeurIPS.

[7] Sagawa, S., Koh, P. W., Hashimoto, T. B., & Liang, P. (2019). Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization. ICLR.

Improving Probabilistic Models for Machine Learning in Science

Supervisors: Aditya Ravuri, Neil D. Lawrence

Each of the six following projects involves understanding and extending an existing probabilistic model commonly used in a scientific context to improve usability and model understanding. Please email me (ar847@cam.ac.uk) if interested.

Interpretable Machine Learning for Intensive Care Decision Support

Supervisors: Christian Cabrera Jojoa, Neil D. Lawrence

Intensive care units (ICUs) generate vast volumes of patient data, yet clinicians often lack tools to translate this data into timely, trustworthy decisions. The aICU project, which aims to support safe, interpretable, and clinically meaningful decision-making by establishing a standardised pipeline for developing, evaluating, and deploying AI in critical care. This project proposes to reproduce existing machine learning models that address specific ICU problems (e.g., mortality prediction, sepsis detection, ventilator weaning, or length-of-stay estimation) and then investigate interpretability methods to make the model predictions understandable to clinicians.

Interpretable Multi-Agent Systems with DOAgent

Supervisors: Christian Cabrera Jojoa, Neil D. Lawrence

Multi-agent systems (MAS) and Self-Adaptive Systems (SAS) are used across robotics, resource management, and autonomous computing, yet understanding why agents make particular decisions remains an open challenge. When multiple agents interact through shared environments, the resulting behaviour is difficult to trace, attribute, and explain. DOAgent is a Python library that addresses this gap by treating shared data as the primary interface between agents, automatically recording decisions, state transitions, and contributions so that agent behaviour can be analysed after execution. This project proposes to reproduce an existing multi-agent or self-adaptive system from the literature using DOAgent, and then explore interpretability approaches on the recorded agent interactions.

Latent Probabilistic Models for Unsupervised Structure Learning in Massively Missing Data

Supervisors: Vidhi Lalchand, Neil D. Lawrence

Real-world datasets often contain entries with missing elements e.g. in a medical dataset, a patient is unlikely to have taken all possible diagnostic tests. Variational Autoencoders (VAEs) are popular generative models often used for unsupervised learning. Despite their widespread use it is unclear how best to apply VAEs to datasets with missing data.

In this projects we intend to explore a novel incarnation of a Gaussian process latent variables which can work seamlessly with missing data, the architecture would include a pointNet encoder which will encode every partial data point as a set with indicator variables (capturing the dimension identity) and a classical Gaussian process decoder.

The focus would be on assessing the quality of uncertainty calibration, structure learning in latent space, and reconstruction quality on previously unseen data.

Learning Depth Perception using Kinect and Python

Supervisor: Neil D. Lawrence

Kinect cameras provide true image and an associated depth image. These two images are providing different information, yet a human can infer depth directly from an image. This project will focus on using machine learning techniques building on the machine learning groups python code to see what can be learnt about depth from images. The ultimate aim will be to reconstruct the depth in a real image by learning about depths from information provided by the Kinect camera. Software for the project will be written according to the principles of open data science.

This project will suit students with strong analytical skills, there will be a focus on linear algebra and probabilistic inference in the software.

Learning to Learn by Denoising Diffusion Optimisers

Supervisors: Francisco Vargas, Neil D. Lawrence

The seminal paper Learning to Learn by Gradient Descent By Gradient Descent [1] explores learning optimisers (rather than using SGD or Adam) through an RNN architecture which models time steps as training steps and an RL based loss.

In this project we aim to achieve a similar product but rather than RNNs and RL we would like to explore the closely linked diffusion models [5,6] and stochastic control. In particular there is already theoretical work that motivates the use of these methodologies for global optimisation [2], what now remains is to explore it in practice.

We expect the student to lightly adapt the methods in [3,4] to the optimisation setting in [2] (This simply amounts to exploring low temperatures in the artificial target distribution induced by the loss function). Furthermore, exploring tasks in engineering as in [1] might require being creative about the inductive biases baked into the NN parametrisation that we are working with.

The nature of this project will be mostly ML-engineering and playing around / being creative about NN inductive biases. That said if the student is more mathematically/theory oriented we could explore extending the results in [2] (which apply to the method in [4] only) to the method in in [3] using the mixing rates in [7] and maybe comparing to things like [8] (last bit would be very bonus/extra type of thing).

As always the student should explore simple 1d and 2d toy optimisation examples to assess the validity of the method before moving onto real world examples.

[1] https://arxiv.org/pdf/1606.04474.pdf 

[2] https://arxiv.org/abs/2111.00402

[3] https://openreview.net/forum?id=8pvnfTAbu1f

[4] https://arxiv.org/abs/2111.15141

[5] https://arxiv.org/abs/2006.11239

[6] https://arxiv.org/abs/2006.11239

[7] https://arxiv.org/pdf/2209.11215.pdf

[8] https://arxiv.org/abs/1707.06618

Machine Learning Methodologies for Personalized Health

Supervisor: Neil D. Lawrence

It is now possible to have several perspectives on a patient. mHealth provides information derived from mobile phones. Full genotyping of patients is becoming affordable, providing information about genetic background. The phenotype of disease is becoming better characterised than ever before. Techniques such as transcriptome analysis allow a highly detailed characterization of the state of a tissue. Finally the UK government’s midata initiative (and similar initiatives elsewhere) may eventually allow patients to provide information about their consumer spending habits as well as social network behaviour.

These different data modalities need to be combined into one model of the patients well being. There are major challenges with doing this: models need to be applied across millions of patients and for any given patient many information modalities will be missing. Addressing data of this type requires new machine learning methodologies. This project will focus on combining data from different modalities within the same probabilistic model.

Machine Learning for Fitness Monitoring with Python

Supervisor: Neil D. Lawrence

Technologies that were previously only available to elite athletes are becoming widespread. Now casual athletes can buy systems that monitor pace, heart rate and other information for under 300 pounds. This project will build a software tool for analysis of data of this type. Software for the project will be written according to the principles of open data science.

This project will suit students with strong analytical skills, there will be a focus on linear algebra and probabilistic inference in the software.

Machine Learning for Modelling Formula One Races

Supervisors: Hanbo Li, Ryan Daniels, Neil D. Lawrence

The machine learning group is working with one of the leading forumla one teams in analysis of data generated in Formula One races with the aim of improving strategy. With this aim we are running one or more projects this year focussed on Formula One data. Formula one is a data intensive sport, information about the location of each team’s car during the race is provided to the teams. Optimization of pit stop strategy can make the difference between winning and losing the race. There are commercial confidentiality issues over which areas will be studied, but interested students can discuss these areas directly with the supervisors.

Motion Capture Data Modelling in Python

Supervisor: Neil D. Lawrence

The Python programming language is becoming a de facto standard for implementation of machine learning algorithms. This project will develop tools for modelling of motion capture data in the Python programming language. Based on existing tools in MATLAB, the aim of the project will be to port the underlying machine learning techniques to the more powerful Python programming language. The end aim is to provide a simple tool for animators to model motion capture data and create new animations for computer games or the film industry.

A previous project has built a visualization framework for motion capture data, this year’s project will focus on more advanced machine learning methodologies. The student will work with an ongoing software development framework being developed by the machine learning group. Software for the project will be written according to the principles of open data science.

This project will suit students with strong analytical skills, there will be a focus on linear algebra and probabilistic inference in the software.

Multi-objective optimisation of cloud infrastructure

Supervisors: Andrei Paleyes, Neil D. Lawrence

Machine learning (ML) and optimisation techniques are increasingly used to help solve decision-making problems that would be difficult or time-consuming to address manually. One such problem is the configuration of cloud infrastructure, where many deployment parameters can affect several competing objectives at the same time. This project investigates the use of multi-objective optimisation to automatically explore different cloud infrastructure configurations defined through Infrastructure-as-Code templates. Our aim will be to build a fully automated system that identifies a range of Pareto-optimal configurations that represent different trade-offs between the objectives being considered. Such a system can help reduce the time and cost required to create efficient cloud deployments while providing a better understanding of the available configuration choices.

Optimisation Benchmarks for Online RL

Supervisors: Pierre Thodoroff, Christian Cabrera Jojoa, Neil D. Lawrence

Reinforcement Learning algorithms have been applied in different domains. Now there is a growing interest in applying RL algorithms to optimisation problems. Such algorithms are a better alternative to produce near-optimal solutions in dynamic environments, compared against exact or approximation algorithms. Benchmarks are a key element in the development and evaluation of novel algorithms as they enable a standardised comparison of these algorithms’ performance. In this project you will provide a set of optimisation benchmarks to evaluate online RL algorithms.

Protein Folding Explantions via Diffusion Bridge Score Matching

Supervisors: Neil D. Lawrence, Francisco Vargas, Pierre Thodoroff

Technical Title: Sampling Transition Paths Between Molecular Conformations Using Diffusion Bridges and Score Matching.

Recent advances in Schrodinger Bridges [1,2] have enabled to learn stochastic mappings between 2 probability distributions (p(x) and q(x)) such that the stochastic map (which is modelled by a diffusion / SDE) is regularised by some prior process (whether it be computational or physical).

This project seeks to explore these methodologies in particular the simpler case studied in [3] where both the source and the target distributions are modeled as point masses (dirac delta functions / measures). We seek to apply the approach in [3] to sampling physically meaningful transition paths between two protein configurations as done in [4]. Ideally, we would aim to start working with simple/toy proteins and then move on to larger scale tasks where one of the protein configurations is a flat amino-acid and proteins produced by alpha fold, the end product would be to generate a video which gives a physically plausible folding process for alphafold [5] predictions.

An example Timeline of the project could be:

  1. Get the codebase of [4] working and reproduce results on simple proteins.
  2. Extend the work in [3] to the underdampened version. (Francisco can help with this)
  3. Apply the extensions and adaptions of [3] to work on the simple proteins of [4] and compare to the method in [4].
  4. Consider enhancements / extensions, would a full Schrödinger bridge work better here?
  5. If time allowing, pick some of the most recently exciting discoveries from alphafold [5] that have a known potential and try and see if we can get it working.

Point 5. is an “if time allows” type of objective and I predict most the time will be spent in 3., successful completion of 3 could lead to a publication at a top venue whilst 5. could have a broader impact on the field.

Ideally a good background in the following could be very helpful for this project:

  1. Timeseries models (Kalman filters, AR processes, Gaussian processes).
  2. Introductory calculus (ODEs, Basic PDEs, limits).
  3. Probability Theory (Limit Theorems, Change of Variables, basic concentration inequalities e.g., Markov/Chebyshev)
  4. Variational Inference (MFVI, Amortised VI, Deep hierarchical latent variable models)


[1] https://arxiv.org/pdf/2106.01357.pdf
[2] https://arxiv.org/pdf/2106.02081.pdf
[3] https://arxiv.org/pdf/2207.02149.pdf
[4] https://arxiv.org/pdf/2111.07243.pdf
[5] https://alphafold.ebi.ac.uk/

Scikic: The Artificial Psychic

Supervisor: Neil D. Lawrence

In this project you will contribute to an artificial psychic called Scikic (scikic.org). The artificial psychic works by querying a user on preferences about life (e.g. movies) and making predictions about what type of person the user is. Scikic consists of a front end (a web interface or a mobile app), and a back end (an information engine). At the moment Scikic isn’t a very good artificial psychic (its information engine is a little rusty, it doesn’t have enough data), but over time Scikic will be able to make good predictions about people using only a little information. Software for the project will be written according to the principles of open data science.

The project is a collaboration with the start up company CitizenMe.

This project could suit students with strong analytical skills: for the inference engine there will be a focus on linear algebra and probabilistic inference in the software. However, we also need students with a good knowledge of web interfaces and a flair for design.

Self-Adaptive Systems and Large Language Models (LLMs)

Supervisors: Christian Cabrera Jojoa, Neil D. Lawrence

Software systems are increasingly complex and include different actors and components interacting in dynamic environments. Maintaining such systems is a difficult task where human intervention is not feasible. Autonomous computing has explored approaches to optimise systems’ performance by changing their structure, behaviour, or environment variables. These approaches rely on feedback loops that accumulate knowledge from the system interactions to inform autonomous decision-making. However, this knowledge is often limited, constraining the systems’ interpretability and adaptability. This project proposes to explore the capabilities of Large Language Models (LLMs) for self-adaptive systems. The main idea is to replace current autonomous RL-agents with LLM-based agents to make self-adaptive decisions.

Unconventional AI and explainable AI

Supervisors: Soumya Banerjee, Neil D. Lawrence

These twelve projects in unconvential and explainable AI would be supervised by Soumya Banerjee.

Verifying Identity

Supervisor: Neil D. Lawrence

How can you tell if a user is who they say they are? How can you tell if they are a real person or a bot? Can you do it without having the user reveal their inforamtion to you. An individual has the right to privacy, but what if they abuse that right to commit fraud? In this project (in collaboration with a start up company) we will consider how machine learning can be used to balance the need of the individual for privacy agains the need of society to be able to validate identity. Our aim is to build distributed user indenity validation systems that do not require the user to reveal personal information. We will do this by designing intelligent, machine learning based, agents that validate a user’s information locally on the telephone. The project may involve collaboration with a London based start up company operating in this area.

This project will suit students with strong analytical skills, there will be a focus on linear algebra and probabilistic inference in the software.

What Does a Good Regulator Need to Know?

Supervisors: Joery de Vries, Neil D. Lawrence

Prerequisite: L172 Information, Energy and Intelligence (IEI), or equivalent preparation in information theory, maximum entropy, and information geometry.

Conant and Ashby’s (1970) good regulator theorem says a successful regulator must be a model of its system. Which model depends on what the regulator observes: Wentworth’s (2021) “gooder regulator” for instance requires the belief state. Since a good regulator objective minimises the entropy of a regulated outcome this adds a secondary dependence during learning due to concavity of the optimization problem. Similar to convex RL, it can be solved by a sequence of linear rewards built from the occupancy of the outcome features. Although the agent converges to a single deterministic policy, during learning its representation must support every reward in the sequence. Therefore, reusing what it learned under earlier rewards while staying focused on what the objective makes relevant is crucial. For instance, the successor features of the outcome suffice for this. Despite much work on state abstraction, self-predictive representations and sensorimotor world models, it is unclear what a good regulator needs to represent while it learns. This project investigates what acting and learning require for good regulators in the language of state abstractions of Li, Walsh and Littman (2006) and of Ni et al. (2024), and what combination of latent world-model loss delivers all aspects.

What Does a Good Regulator Need to Know?

Supervisors: Joery de Vries, Neil D. Lawrence

Prerequisite: L172 Information, Energy and Intelligence (IEI), or equivalent preparation in information theory, maximum entropy, and information geometry.

Conant and Ashby’s (1970) good regulator theorem says a successful regulator must be a model of its system. Which model depends on what the regulator observes: Wentworth’s (2021) “gooder regulator” for instance requires the belief state. Since a good regulator objective minimises the entropy of a regulated outcome this adds a secondary dependence during learning due to concavity of the optimization problem. Similar to convex RL, it can be solved by a sequence of linear rewards built from the occupancy of the outcome features. Although the agent converges to a single deterministic policy, during learning its representation must support every reward in the sequence. Therefore, reusing what it learned under earlier rewards while staying focused on what the objective makes relevant is crucial. For instance, the successor features of the outcome suffice for this. Despite much work on state abstraction, self-predictive representations and sensorimotor world models, it is unclear what a good regulator needs to represent while it learns. This project investigates what acting and learning require for good regulators in the language of state abstractions of Li, Walsh and Littman (2006) and of Ni et al. (2024), and what combination of latent world-model loss delivers all aspects.

Why 10 random seeds? Exploring a heteroskedastic multiverse

Supervisors: Samuel J. Bell, Neil D. Lawrence

A multiverse analysis, first introduced in psychology [1], is a principled toolkit for evaluating the robustness and generality of scientific claims. As researchers, we make numerous choices through the course of an investigation that could influence the final result. Applying the multiverse analysis to machine learning, Bell et al. [2] systematically evaluate how researcher choices (e.g., hyperparameters, evaluation dataset) can fundamentally change the conclusions drawn. For example, Bell et al. evaluate the role of hyperparameters in optimizer choice, and evaluate whether large batch training necessarily leads to a drop in generalisation performance.

In a deep learning setting, the multiverse is large and contains continuous dimensions, making it challenging to rigorously explore. To overcome this, Bell et al. implement a simple Gaussian Process surrogate and turn to Bayesian experimental design to sample experimental configurations to evaluate. This has the added benefit of allowing us to quantify our confidence in different regions of the multiverse.

A simplifying assumption in Bell et al.’s approach is that the multiverse is homoskedastic: that each point has the same variance. This is unlikely to be the case when working with neural networks. This project will extend the underlying surrogate to account for heteroskedasticity. One possible approach to this is to explicitly model the noise as a function of the inputs using a separate Gaussian Process [3].

This approach has an important consequence. In deep learning, it is typical to evaluate models using the mean performance over a number (usually 10) runs with different random seeds. But why 10? Whether this is an appropriately large sample size for robust inference, i.e. to confidently claim model A outperforms model B, remains to be seen (for an example in NLP, see [4]). By accounting for heteroskedasticity in a multiverse analysis, we can make an informed decision about how many runs are appropriate, and make a principled trade-off between re-running the same configuration and sampling a new point in the multiverse [5].

If you’re interested in the project, get in touch with Samuel Bell (sjb326@cam.ac.uk) for an informal chat.

[1] Steegen, S., Tuerlinckx, F., Gelman, A., & Vanpaemel, W. (2016). Increasing transparency through a multiverse analysis. Perspectives on Psychological Science.

[2] Bell, S. J., Kampman, O. P., Dodge, J., & Lawrence, N. D. (2022). Modeling the Machine Learning Multiverse. NeurIPS.

[3] Goldberg, P., Williams, C., & Bishop, C. (1997). Regression with input-dependent noise: A Gaussian process treatment. NeurIPS.

[4] Card, D., Henderson, P., Khandelwal, U., Jia, R., Mahowald, K., & Jurafsky, D. (2020, November). With Little Power Comes Great Responsibility. EMNLP.

[5] Binois, M., Huang, J., Gramacy, R. B., & Ludkovski, M. (2019). Replication or exploration? Sequential design for stochastic simulation experiments. Technometrics.