Our research develops methods for extracting reliable information from scientific data under realistic conditions, with a focus on data-efficient machine learning, uncertainty quantification, and scalable algorithms. We are particularly interested in settings with expensive, heterogeneous or uncertain data, approximate models, and limited computational resources.
Data-efficient and multifidelity machine learning
High-quality scientific data can be expensive to generate or acquire. Multifidelity methods make it possible to combine information from data and models of different accuracy and computational cost. Together with sampling and active learning, these methods allow new data points to be selected where they are most useful and expensive training data to be used more efficiently. A major application area in our group is molecular machine learning and quantum chemistry, where accurate training data can be particularly expensive to generate.
Uncertainty quantification and Bayesian methods

Scientific predictions are often affected by uncertainty in observations, models, and assumptions. Our work on uncertainty quantification addresses these uncertainties in data-driven and simulation-based predictions. For inverse problems, Bayesian methods provide a framework for combining uncertain or indirect observations with mathematical or simulation models to infer quantities that cannot be observed directly. A current application in our group is climate reconstruction, where indirect and noisy observations are combined with forward models and statistical assumptions.
Scalable kernel methods and algorithms

Kernel-based methods, including Gaussian processes and kernel-based surrogate models, provide a mathematically well-understood framework for approximation, learning, and uncertainty modelling. Their computational and memory cost, however, can become limiting for larger data sets and problems. Our work therefore aims to reduce these costs through structured kernel approximations, low-rank and hierarchical methods, and parallel and hardware-aware algorithms. Efficient implementations and research software are an important part of making these methods scalable, reproducible, and usable in scientific applications.
Scientific applications driving methodological development

Many of these methods are developed in close collaboration with researchers from other disciplines. Current applications include molecular modelling and quantum chemistry, climate reconstruction, and system modelling. In these collaborations, application-specific constraints often lead to new methodological questions: which data is informative, which model assumptions are appropriate, which uncertainties need to be represented explicitly, and what computational effort is justified.