Skip to main content

Research Seminars

Our research seminars bring leading academics to the LSE Department of Statistics throughout the year. They showcase a broad range of methodological and applied scholarship in statistical science, and strengthen our global research connections.

The 'Statistics and Data Science Seminar Series' forms the backbone of our research seminar programme, and these usually take place on Monday afternoons in the Leverhulme Library (COL 6.15).

Research seminars are open to all LSE staff and students, as well as selected invited guests. Come along, exchange ideas, build collaborations, and engage with researchers from Europe, the US, and worldwide.

Find the details of upcoming and past seminars below.

2026

Speaker: Professor Rina Foygel Barber

Date: 06/07/2026

Abstract: We study the problem of coincidence detection: given data from multiple sensors, each individual sensor may have a high false positive rate when searching for events or signals among background noise, but we can reduce the rate of false positives by searching for simultaneous detections in multiple sensors. The coincidence detection problem arises across many applications, and this particular work is motivated by applications in astrophysics, where the aim is to detect astrophysical events such as gravitational waves. A commonly used technique in that field is "time-shifting": the timeline of one data stream is randomly shifted relative to another, to determine whether nearly simultaneous detections might occur frequently simply due to random chance. This talk will present a theoretical analysis of the time-shifting methodology, examining whether it is able to offer false positive rate control in settings where the data streams may each have high temporal dependence, as is common in time series data. This work is joint with Ruiting Liang, Samuel Dyson, and Daniel Holz.

Biography: My research focuses on the theoretical foundations of statistical problems in estimation, prediction, and inference. In many modern settings, classical methods may not be reliable due to high dimensionality, failure of model assumptions, or other challenges. I work on distribution-free inference methods such as conformal prediction, and on developing hardness results to establish what types of questions can or cannot be solved with distribution-free methods. I am also interested in multiple testing methods, in algorithmic stability, and shape-constrained inference. I also collaborate on modeling and optimization problems in image reconstruction for medical imaging.

Speaker:Ying (Alison) Cheng, PhD

Date: 11/06/2026

Abstract: Quality control in assessment is of critical importance to ensure the reliability and validity of the scores produced by assessments, and the replicability and reproducibility of research findings. Various methods have been developed to detect behavior anomalies in assessments, such as disengagement, rapid guessing, cheating, or shifts in behavior. This talk presents a statistical framework for detecting and addressing such anomalies through outlier detection, change point analysis, and robust estimation methods within latent variable models. We begin with person-fit statistics, which identify individuals whose response patterns deviate from model expectations, and discuss their connections to likelihood-based diagnostics. We then introduce change point analysis to detect within-person shifts in response behavior over the course of an assessment. To mitigate the impact of anomalous responses, we examine robust estimation approaches that reduce sensitivity to contamination. These methods are discussed in the context of item response theory (IRT) and related models. Given the growing interest in there’s growing interest in multi-modal data and passive sensing data in psychological and educational research, in developing and applying these methods, we consider the use of response time data alongside response accuracy to improve detection and interpretation.

Speaker: Dr Mengchu Li

Date: 06/07/2026

Abstract:L earning from distributed and heterogeneous data is central to modern data science. Recent advances in learning with multi-source data have shown that effectively integrating information across related datasets can significantly improve algorithmic performance. However, heterogeneity across datasets, in terms of sample size, distributional shift, and data quality, poses fundamental challenges in determining the optimal strategy for aggregating information. Moreover, sharing potentially sensitive information across distributed units raises serious privacy concerns.

In this talk, I will discuss two related projects addressing these challenges. The first studies federated transfer learning under privacy constraints. We introduce a notion of federated differential privacy, which protects each local data set without assuming a trusted central server, and characterise the statistical costs of privacy and heterogeneity across several statistical problems. The second project focuses on robust multi-task learning against adversarial contamination. We show that several existing regularisation-based approaches suffer from a dimension-dependent contamination error and are therefore statistically suboptimal. Motivated by this gap, we develop a computationally efficient filtering-based method that achieves near-optimal statistical performance over a broad range of model parameters.

Biography: Dr Mengchu Li is an Assistant Professor in the School of Mathematics, University of Birmingham. His research focuses on statistical theory and methodology for heterogeneous data analysis, especially in high-dimensional settings and under privacy constraints. Particular topics include change point analysis, transfer learning, robust statistics and differential privacy.

Speaker: Dr. Song

Date: 25/06/2026

Biography: Testing for concurrent effects arises from many applications such as mediation analysis, replicability analysis, and construction of causal knowledge graphs. When a null hypothesis involves multiple parameters, such a task of hypothesis testing is analytically of great difficulty. Classical procedures, including the Sobel's test and the MaxP test, are known to be conservative in the type I error control, and a unified methodology to address this issue is an important open problem in the statistical literature. In this paper, we propose a new solution using renewable estimating functions (REFs) to construct a pivotal test statistic that is invariant to regions of the null parameters. Consequently, the proposed method can overcome conservatism and achieve proper control of type I error and moreover improve statistical power. The key analytic behind our methodology resembles stochastic gradient descent, which incrementally decouples null cases to translate a test for complex composite nulls into that for a simple single null. We establish a full theoretical framework for this new methodology, including key large-sample properties. Through extensive simulation studies and real-world data applications, we demonstrate that the proposed method achieves superior finite-sample performance compared to existing alternatives. This is a joint work with Drs. Canyi Chen and Ling Zhou.

Biography: Dr. Song is Professor of Biostatistics at the University of Michigan School of Public Health, Ann Arbor. He received his PhD in Statistics from the University of British Columbia, Vancouver, Canada in 1996. He has published over 250 peer-reviewed papers and graduated 28 PhD students and trained 6 postdoc research fellows. Dr. Song's current research interests include data integration, distributed inference, high-dimensional data analysis, longitudinal data analysis, mediation analysis, and spatiotemporal modeling. He is AAAS Fellow, IMS Fellow, ASA Fellow and Elected Member of the International Statistical Institute. Dr. Song now serves as Area Editor of the Annals of Applied Statistics (Medicine, EHR and Smart Health), Associate Editor of the Journal of American Statistical Association (both T&M and ACS), Journal of the Royal Statistical Society Series B (Statistical Methodology) and the Journal of Multivariate Analysis. He was Interim Chair and Associate Chair for Research of the U-M Biostatistics and is now chairing the ASA Interest Group StatsUpAI and the IMS Committee on Fellows

Speaker: Dr Sander Beckers

Date: 21/05/2026

Abstract: Recent work by Chatzi et al. and Ravfogel et al. has developed, for the first time, a method for generating counterfactuals of probabilistic Large Language Models. Such counterfactuals tell us what would – or might – have been the output of an LLM if some factual prompt x had been x* instead. The ability to generate such counterfactuals is an important necessary step towards explaining, evaluating, and eventually improving, the behavior of LLMs.

I argue, however, that the existing method rests on an ambiguous interpretation of LLMs: it does not interpret LLMs literally, for the method involves the assumption that one can change the implementation of an LLM’s sampling process without changing the LLM itself, nor does it interpret LLMs as intended, for the method involves explicitly representing a nondeterministic LLM as a deterministic causal model. I here present a much simpler method for generating counterfactuals that is based on an LLM’s intended interpretation by representing it as a nondeterministic causal model instead. The advantage of this simpler method is that it is directly applicable to any black-box LLM without modification, as it is agnostic to any implementation details.

The advantage of the existing method, on the other hand, is that it directly implements the generation of a specific type of counterfactuals that is useful for certain purposes, but not for others. I clarify how both methods relate by offering a theoretical foundation for reasoning about counterfactuals in LLMs based on their intended semantics, thereby laying the groundwork for novel application-specific methods for generating counterfactuals.

Biography: Sander is a postdoctoral fellow at the Department of Statistical Science, where he does research as part of the Causality in Healthcare and AI (CHAI) hub, under the direction of Ricardo Silva. Ever since obtaining his PhD in Computer Science at the University of Leuven in 2016, Sander has worked on causal and ethical issues across the fields of philosophy and AI. He was previously employed at Cornell University, University of Amsterdam, University of Tuebingen, LMU Munich, and University of Utrecht, across various departments and specialties.

Speaker: Prof. Mats Julius Stensrud

Date: 13/05/2026

Abstract:

Many studies aim to estimate treatment effects on outcomes that are defined only for individuals who experience a post-treatment event. For example, the effect of cancer therapies on quality of life is only well defined among individuals who are alive. Similar complications arise in evaluations of educational or labor-market interventions, where outcomes such as final exam scores or wages are defined only for individuals who remain in school or become employed. A naive comparison of outcomes conditional on such post-treatment events generally lacks a causal interpretation, even when treatment is randomly assigned.

In this talk, I discuss causal contrasts for outcomes conditional on post-treatment events, including principal stratum effects and conditional separable effects. I then derive identification results for these estimands and discuss their interpretation and the conditions under which they might, in principle, be falsified. I illustrate the relevance of these results for clinical trials of cancer therapies and vaccines, and I conclude by revisiting the classical birth weight paradox in epidemiology.

Biography: Prof. Mats Julius Stensrud — Chair of Biostatistics at EPFL (since 2020) and Director of the Doctoral School in Mathematics at EPFL (since 2025) — is a medical doctor, Dr. philos., and Associate Professor of Statistics. He holds degrees in mathematics, medicine, and statistics, and has professional experience as a clinical physician. Before joining EPFL in 2020, he was a Fulbright Scholar and Kolokotrones Fellow at the Harvard T. H. Chan School of Public Health and a postdoctoral researcher at the University of Oslo.

His research focuses on causal inference and statistical methodology. He develops methods to address challenges in analyzing longitudinal data, such as time to event outcomes. His work puts emphasis on scientifically meaningful targets, which are studied under assumptions that are transparent and testable. One part of his work clarifies the causal interpretation of classical survival analyses and gives the theory of separable effects, a class of causal parameters that disentangle direct and indirect effects. While working on these problems, he also became interested in questions about mechanisms, including mediation analysis. His work in this domain has been recognized with the Arthur Linder Prize (2023), the Lambert Award (2023), and the Sverdrup Prize (2024).

Speaker: Timothy Sudijono

Date: 27/05/2026

Abstract: We show that feedforward neural networks generalize on low complexity data, suitably defined. Given data generated from a simple programming language, the minimum description length (MDL) feedforward neural network which interpolates the data generalizes with high probability. We define this simple programming language, along with a notion of description length of such networks. We provide several examples on basic computational tasks, such as checking the primality of a natural number. Extensions to noisy data are also discussed, suggesting that MDL neural network interpolators can demonstrate tempered overfitting. This is joint work with Sourav Chatterjee.

Biography: I'm a 5th year PhD student advised by Sourav Chatterjee interested in neural networks, empirical Bayes, and causal inference.

Speaker: Professor Fadoua Balabdaoui

Date: 20/05/2026

Abstract: Consider the regression problem where the response and the covariate are unmatched. Under this scenario we do not have access to pairs of observations from their joint distribution, but instead we have separate data sets of responses and covariates, possibly collected from different sources. We study this problem assuming that the regression function is linear and the noise distribution is known or can be estimated. We introduce an estimator of the regression vector based on deconvolution (the DLSE) and demonstrate its consistency and asymptotic normality under parametric identifiability. Under non-identifiability of the regression vector but identifiability of the distribution of the predictor, we construct an estimator of the latter based on the DLSE and show that it converges to the true distribution of the predictor at the parametric rate in the Wasserstein distance of order 1. We illustrate the theory with several simulation results.

This talk is based on my joint work with Mona Azadkia, Antonio di Noia and Cecile Durot

Biography: F. Balabdaoui is currently an Adjunct Professor at the Seminar for Statistics at ETH of Zurich. Her main research interests are non-parametric and shape-constrained estimation, mixture models, unmatched regression, prediction-powered inference, asymptotic theory and empirical processes.

Speaker: Professor Michael Wolf

Date: 13/05/2026

Abstract: Accurate inflation forecasting is essential for economic policy, financial markets, and broader societal stability. In recent years, machine-learning methods, particularly the random forest, have shown strong potential for improving forecasting accuracy, often outperforming traditional benchmarks. Building on this success, our paper adapts the hedged random forest (HRF) of \cite{beck:kozbur:wolf:hrf} to the task of forecasting inflation. Unlike the standard random forest, the HRF assigns non-equal, and potentially negative, weights to individual trees to enhance forecasting performance. We develop customized estimators for the HRF’s two key inputs: the mean and covariance matrix of the vector of forecast errors corresponding to individual trees. An extensive empirical analysis demonstrates that our approach consistently outperforms the standard random forest.

Abstract:

Michael Wolf is a Professor of Econometrics and Applied Statistics at the University of Zurich and a Senior Fellow at ADIA Lab, Abu Dhabi. Heh holds a Ph.D. in Statistics from Stanford University. Before joining the University of Zurich's Department of Economics, he held positions at the University of California, Los Angeles (UCLA), Universidad Carlos III de Madrid, and Universitat Pompeu Fabra in Barcelona.

His research interests include resampling-based inference, multiple-testing procedures, shrinkage estimation of large covariance matrices, financial econometrics, and machine learning. His work has been published in leading journals such as The Annals of Statistics, Biometrika, Econometrica, Journal of the American Statistical Association, and The Review of Financial Studies.

Speaker: Professor Min-ge Xie

Date: 06/05/2026

Abstract: Rapid advances in data science demand fundamentally new statistical frameworks for inference problems that go beyond the classical setup -- especially when parameters are discrete or non-numerical, or data types are irregular and non-numerical. This talk presents an effective, likelihood-free, wide-reaching, and principled framework, called the repro samples method, for tackling such challenges. The development leverages repro samples (i.e., synthetic data generated to mimic observations or noise) together with Fisher inversion techniques to enable rigorous inference. In this talk, we specifically focus on inference problems that simultaneously involve mixed types of discrete/nonnumerical and continuous parameters, covering a wide range of irregular challenges arising in modern data science.

A step-by-step case study on inference in high-dimensional regression that accommodates model selection uncertainty is provided, including the construction of confidence sets for the unknown sparse model, for a single or any collection of regression coefficients, or for both the sparse model and regression coefficients jointly. The effectiveness of the method is illustrated in a model-free high-dimensional binary classification setting through simulations and a real-data analysis of single-cell RNA-seq data from mouse bone marrow–derived dendritic cells.

Biography: Min-ge Xie, PhD, is a Distinguished Professor at Rutgers, The State University of New Jersey. Dr. Xie received his PhD in Statistics from the University of Illinois at Urbana-Champaign and his BS in Mathematics from the University of Science and Technology of China.

He is the current Editor of The American Statistician and a co-founding Editor-in-Chief of the New England Journal of Statistics in Data Science. He is a Fellow of the ASA, a Fellow of the IMS, an elected member of the ISI, and a Fulbright Scholar. His research interests include the theoretical foundations of statistical inference and machine learning uncertainty quantification, fusion learning, finite- and large-sample theory, and parametric and nonparametric methods. He is the Director of the Rutgers Office of Statistical Consulting and has extensive interdisciplinary research experience collaborating with biomedical researchers, computer scientists, engineers, and scientists in other fields.

Speaker: Professor Yulan He

Date:

Abstract: As large language models (LLMs) become more capable at reasoning, a paradox emerges: improving their reasoning can make them simultaneously more intelligent and less trustworthy. This talk will explore this paradox from three areas. First, we address the problem of faithfulness. Through a probabilistic inference framework that leverages task-specific and lookahead rewards, and through Concise-SAE, a training-free method for identifying and editing instruction-relevant neurons, we show how to ensure that LLMs’ explanations and decisions remain grounded in input context and intent. Next, we address efficiency in reasoning.

NOVER (No-Verifier Reinforcement Learning) enables incentive-based reasoning without requiring external verifiers, while CODI (Continuous Chain-of-Thought via Self-Distillation) compresses explicit natural language reasoning into a continuous latent space, achieving efficiency and robustness without compromising accuracy. Finally, we delve into the tension between capability and alignment. Through the discovery of Reasoning-Induced Misalignment (RIM), we reveal how improving reasoning capabilities can negatively impact model alignment and safety. Our studies outline a trajectory toward trustworthy reasoning models.

Biography: Yulan He is a Professor in Natural Language Processing at King’s College London. She is currently holding a prestigious 5-year UKRI Turing AI Fellowship. Her recent research focused on addressing the limitations of Large Language Models (LLMs), aiming to enhance their reasoning capabilities, robustness, and explainability. She has published over 300 papers on topics such as self-evolution of LLMs, mechanistic interpretability, and LLMs for educational assessment and health. She received several prizes and awards for her research, including an SWSA Ten-Year Award, a CIKM Test-of-Time Award, and was recognised as an inaugural Highly Ranked Scholar by ScholarGPS. She served as the General Chair for AACL-IJCNLP 2022 and a Program Co-Chair for various conferences such as ECIR 2024, CCL 2024, and EMNLP 2020. Her research has received support from the EPSRC, Royal Academy of Engineering, EU-H2020, Innovate UK, British Council, and industrial funding.

Speaker: Professor Emma Brunskill

Date: 13/03/2026

Abstract: From medicine to marketing to social sciences, the promise of tailoring interventions to individual characteristics is undeniable. However, personalization often comes with costs— from logistical challenges to lack of shared context to concerns about fairness. In addition, personalized decision policies can be more fragile, because they typically require more data to learn accurately compared to identifying a single best intervention for all. In this talk I’ll introduce a new statistical estimator that quantifies, given historical data, if there is evidence that a personalized intervention policy provides significantly superior expected outcomes compared to deploying the best single overall intervention. We present results across four diverse datasets to highlight the wide range of settings where quantifying the impact of personalization can be helpful, and the strength of our proposed estimator over prior related approaches.

Joint work with Zhaoqi Li.

Biography: Emma Brunskill is a tenured professor in the Computer Science Department at Stanford University where she and her lab aim to create AI systems that learn from few samples to robustly make good decisions. Brunskillhas received multiple awards including a Rhodes Scholarship, a NSF CAREER award, an Office of Naval Research Young Investigator award, a Microsoft Faculty Fellow award, and an alumni impact award from the computer science and engineering department at the University of Washington. She is a fellow of the Association for the Advancement of AI (AAAI) and Brunskill served as co-program chair of the International Conference on Machine Learning (ICML 2023). Brunskill and her lab have received 10 best paper nominations and awards across their AI and machine learning research and their AI for education research.

Speaker: Professor Xiaowen Dong

Date: 06/03/2026

Abstract: The increasing availability of graph-structured data motivates a new type of optimisation problems over graph-based functions, i.e., searching for the graph or node that maximises the value of an underlying function. Such optimisation problems are challenging due to the search space that is discrete and high-dimensional, as well as the underlying function that is often black-box and expensive to evaluate. In this talk, I will provide several examples on how Bayesian optimisation can be used to optimise graph-based functions defined on graphs, node set of a graph, and node subsets of a graph. These are enabled by generalising Gaussian processes to graph-structured data, and they demonstrate the promise in combining probabilistic and geometric reasoning for analysing complex functions or solving machine learning tasks. Practical applications include automated machine learning, epidemiological source identification, and social influence maximisation.

Biography: Xiaowen Dong is an associate professor in the Department of Engineering Science at the University of Oxford, where he is a member of the Machine Learning Research Group. Prior to joining Oxford, he was a postdoctoral associate in the MIT Media Lab, and received his PhD degree from the Swiss Federal Institute of Technology (EPFL). His main research interests concern signal processing and machine learning techniques for analysing data with complex structures, and their applications in social, urban, and financial network analysis.

Speaker: Professor Gesine Reinert

Date: 12/12/2025

Abstract: Synthetic data are increasingly used in computational statistics and machine learning. Some applications relate to privacy concerns, to data augmentation, and to method development. A particular interest lies in anomaly detection.

Synthetic data should reflect the underlying distribution of the real data, being faithful but also showing some variability. In this talk we focus on networks as a data type, such as networks of transactions between agents. This data type poses additional challenges due to the complex dependence which it often represents.

The talk will present a new idea for synthetic network generation. It will also include a statistical method for assessing their quality. Theoretical guarantees for both, the quality assessment and the data generation, are based on Stein's method. The talk will touch on these guarantees. It will conclude with some ideas for non-network data generation.

Biography: Professor Gesine Reinert is a Research Professor at the Department of Statistics, Oxford, and a Professorial Fellow at Keble College. She is also a Fellow of the Alan Turing Institute in London, the national institute for data science, with headquarters at the British Library.

Her current and main research interests are in network statistics and to investigate such networks in a statistically rigorous fashion. Often this will require some approximation. Her preferred tool for such approximations is Stein’s method, an excellent method to derive distances between the distributions of random quantities. A good part of her research concerns the further development of Stein’s method.

The general area of Prof Reinert’s research falls under the category of Applied Probability while many of the problems and examples she studies come from Network Science and Computational Biology (or Bioinformatics).

Speaker: Dr Adam Sykulski

Date: 28/11/2026

Abstract: Spectral techniques are widely used for a number of reasons in time series analysis and spatial statistics, for example 1) to understand the frequency or wavenumber content via the power spectral density, 2) to speed up computation of statistical operations such as matrix inversion via Fast Fourier Transforms, and 3) to perform signal compression. With increasingly big data sources, the variance of spectral methods can be reduced by smoothing across multiple measurements and realisations, thus reducing the error in stochastic environments with noisy signals. In this talk I will show, however, that there is a real issue of statistical bias which will not always reduce with increasing signal sizes, and indeed will become the dominant source of error meaning that if the bias is ignored then one would become increasingly confident in something that is increasingly wrong as the signal lengths grow!

I will show the prevalence and significance of this bias, and how it can be removed, in two contexts: 1) estimating the power spectral density in nonparametric estimates of the power spectral density, and 2) estimating parameters of stochastic processes in the frequency domain using the Whittle Likelihood. The bias correction is designed to work with both spatial and time series data, as well as data that might be non-Gaussian, subject to missingness, or non-stationary. I will present motivating applications in oceanography and the geosciences.

Biography: Adam is an Associate Professor and Reader in Statistics. Adam's research is in time series, spatial and spatio-temporal statistics, with an application focus in environmental, climate and ocean sciences. Adam has supervised 7 PhD students to completion (with a further 7 under current supervision), 5 post-doctoral researchers and 21 Master’s project students. Adam has published 35 peer-reviewed papers in statistics journals such as Biometrika (x3), the Journal of the Royal Statistical Society (Series B and Series C, x3), Journal of Time Series Analysis, and Spatial Statistics (x2), and broader scientific journals such as the Journal of Geophysical Research, IEEE Transactions on Signal Processing (x2), and the European Journal of Operational Research. In 2024 Adam was elected to be a Council member and Trustee of the Royal Statistical Society from 2025-2028, where he was previously the Discussion Papers Editor and Discussion Meetings Secretary for the Journals of the Royal Statistical Society from 2021-2024.

Speaker: Professor Christophe Giraud

Date:

Abstract:

In many high-dimensional problems, the best polynomial-time estimators fall short of the information-theoretic limits that are provably attainable without computational constraints. The low-degree polynomial framework has emerged as a powerful tool for analyzing the fundamental capabilities of polynomial-time algorithms. Building on recent advances in the study of low-degree lower bounds, we will explore two notable high-dimensional phenomena:

  1. Complex structures – We will show that polynomial-time algorithms may fail to leverage certain low-dimensional yet complex structures.
  2. Beyond predictions from statistical physics – We will highlight cases where the computational barriers suggested by tools from statistical physics do not remain valid in specific high-dimensional regimes.

Biography: Christophe Giraud is a Professor at the Institut de Mathématiques d’Orsay, Université Paris-Saclay. After an early career in probability theory and mathematical physics, his research interests shifted toward fundamental problems in statistics and machine learning, with a particular focus on high-dimensional settings. He serves as an Action Editor for the Journal of Machine Learning Research (JMLR), and as an Associate Editor for the Journal of the European Mathematical Society (JEMS). He is also an Adjunct Professor at Cornell University. He is the author of Introduction to High-Dimensional Statistics, with editions published in 2015 and 2021.

Speaker: Dr 顾雨琦

Date:

Abstract: Overlapping clustering aims to assign data points to multiple clusters and has wide applications in fields including educational cognitive diagnosis, recommendation systems, network community detection, etc. The combinatorial nature of this problem creates prohibitive computational challenges for recovering the overlapping cluster memberships. We propose a scalable geometry-inspired spectral method for overlapping clustering with solid statistical guarantees. Geometrically, we reveal that the population singular subspace embeddings of the exponentially many overlapping clusters form a nice “crystal” structure of a parallelotope, which is a parallelogram generalized to higher dimensions. By exploiting this geometry and a mild identifiability condition, we gracefully turn the challenging combinatoric optimization problem of finding 2^K overlapping clusters into a geometric problem of locating K major edges of a parallelotope. Methodologically, we propose an edge hunting plus hard thresholding procedure to estimate the overlapping clusters and other model parameters. We call our method CRYSTAL: Cluster-geometry-inspired Spectral Algorithm. Theoretically, in the high-dimensional regime, we prove exact recovery guarantees for the overlapping clusters. We also establish asymptotic normality of the spectral estimators, which facilitates statistical inference and uncertainty quantification for the parameters. We demonstrate our method by applying it to an international educational assessment dataset to perform cognitive diagnosis, as well as an ecology dataset to perform overlapping clustering of species.

Biography: I am an Assistant Professor in the Department of Statistics at Columbia University. I am also a member of the Data Science Institute. Before joining Columbia in 2021, I spent a year as a postdoc at Duke University, mentored by David B. Dunson. In 2020 I received a Ph.D. in Statistics from the University of Michigan, advised by Gongjun Xu.

In 2015 I received a B.S. in Mathematics from Tsinghua University. My first name can be pronounced as /ju:-tʃi:/. My name in Chinese is 顾雨琦. My research develops statistical theory and methods for uncovering latent structure in modern complex data. A unifying theme is to make latent structure and representation learning identifiable, interpretable, computationally scalable, and statistically reliable.

• Identifiable deep generative models and causal representation learning: I study identifiability, latent graph discovery, and causal representation learning in nonlinear probabilistic graphical models with latent structures.

• High-dimensional statistical inference for latent structure: The high dimensionality and latent structure pose double statistical challenges. I develop spectral, tensor, and likelihood-based methods for mixture, mixed-membership, and nonlinear low-rank representation problems, with finite-sample theory and uncertainty quantification.

• Latent variable models for psychometrics, heterogeneous data, and AI evaluation: I propose principled latent variable models for educational, psychological, biomedical, and language-model data, including cognitive diagnosis, item response theory, and psychometric frameworks for evaluating large language models (LLMs).

Speaker: Professor Yasser Roudi

Date:

Abstract:

In the standard setup for learning, a learning machine, whether a neural newtork, statistical model, biological agent, attempts to lear structure from data generated from the environment. In a closed-loop learning, on the other hand, model parameters are repeatedly estimated from data generated from the model itself, another model with a similar structure. Given the possibility that large neural network models may, in the future, be primarily trained with data generated by artificial neural networks themselves, understanding the long-term behaviour of learning machines under closed-loop learning is receiving a lot of attention.

In this talk, we study this process in the case of models that belong to exponential families. We show that in this case, maximum likelihood estimation of the parameters results in a process that converges to absorbing states that amplify initial biases present in the data. This outcome may, however, be prevented if the data contains at least one data point generated from a ground truth model, by relying on maximum a posteriori estimation or by introducing regularisation. We also discuss similar results in other learning machines, specifically the Hopfield model and Restricted Boltzmann Machines.

Ref: https://arxiv.org/abs/2506.20623

Biography: Professor Yasser Roudi is a Professor of Disordered Systems in the Department of Mathematics, King’s College London. He received his PhD from SISSA (International School for Advanced Study) in 2005, and before that read physics at Sharif University of Technology in Tehran. Prior to joining KCL, he worked at the Kavli Institute for Systems Neuroscience in Norway, NORDITA, University College London and the Institute for Advance Study. He is a member of Royal Norwegian Society of Sciences and Letters and has received several awards for his research including the Eric Kandel Young Neuroscientist prize in 2015.

Speaker: Professor Primoz Skraba

Date:

In this talk I will focus on universal behavior exhibited by scale invariant geometric functionals. I will focus on two cases: In the first half of the talk will concerned with persistent homology and topological data analysis. First discovered experimentally, a particular functional of persistent homology was found to exhibit universal behavior, in that its distribution did not depend on the input distribution. This led to the formulation of three conjectures relating to this phenomenon, the first of which we call weak universality, we proved. In this talk I will present the setting of persistent homology, an outline of the main ideas of the proof (which is phrased for general scale invariant functionals) and applications.

In the second half of the talk I will talk about more recent work which is concerned with intrinsic dimensionality estimation. I will discuss how universality leads to not only a rigorous theoretical basis for dimensionality estimation but also improved results in practice.

Biography: Primoz Skraba is a Professor in Applied and Computational Topology. His research is broadly related to data analysis with an emphasis on topological data analysis. Generally, the problems he considers span both theory and applications. On the theory side, the areas of interest include stability and approximation of algebraic invariants, stochastic topology (the topology of random spaces), and algorithmic research. On the applications side, he focuses on combining topological ideas with machine learning, optimization, and other statistical tools. Other applications areas of interest include visualization and geometry processing.

He received a PhD in Electrical Engineering from Stanford University in 2009 and has held positions at INRIA in France and the Jozef Stefan Institute, the University of Primorska, and the University of Nova Gorica in Slovenia, before joining Queen Mary University of London in 2018. He is also currently a Fellow at the Alan Turing Institute.

Speaker: Dr Justin Weltz

Date: 21/11/2025

Abstract: Across the social and behavioral sciences, there is interest in how relationships shape individual welfare. However, in many contexts, collecting data on social networks is difficult because the relevant population is inaccessible with conventional sampling methodologies. This talk will focus on two methods for studying the social networks of these “hidden” or “hard-to-reach” populations. First, we will discuss respondent-driven sampling (RDS), which is widely used to study marginalized or stigmatized populations by incentivizing study participants to recruit their social connections. The success and efficiency of RDS can depend critically on the nature of the incentives, including their number, value, call to action, etc. Standard RDS uses an incentive structure that is set a priori and held fixed throughout the study. Thus, it does not make use of accumulating information on which incentives are effective and for whom. We propose a reinforcement learning (RL) based adaptive RDS study design in which the incentives are tailored over time to maximize cumulative utility during the study.

We show that these designs are more efficient, cost-effective, and can generate new insights into the social structure of hidden populations. Second, we will address hard-to-reach populations in community social network surveys. In many settings, only specific members of a household, such as the household head, can be accessed and queried about their social connections. This makes a complete census impossible and results in an incomplete sampling frame for the social network. To remedy this issue, we explore how questions that address household behavior, such as the exchange of household goods, can be leveraged to infer missing information about community members who cannot be sampled.

Biography:

Speaker:

Dr. Yuting We

Professor Yuxin Chen

Date: 14/11/2025

Abstract: The denoising diffusion probabilistic model (DDPM) has become a cornerstone of generative AI. While sharp convergence guarantees have been established for DDPM, the iteration complexity typically scales with the ambient data dimension of target distributions, leading to overly conservative theory that fails to explain its practical efficiency. This has sparked recent efforts to understand how DDPM can achieve sampling speed-ups through automatic exploitation of intrinsic low dimensionality of data. This talk explores two key scenarios: (1) For a broad class of data distributions with intrinsic dimension k, we prove that the iteration complexity of the DDPM scales nearly linearly with k, which is optimal under the KL divergence metric; (2) For mixtures of Gaussian distributions with k components, we show that DDPM learns the distribution with iteration complexity that grows only logarithmically in k. These results provide theoretical justification for the practical efficiency of diffusion models.

Abstract: Large language models are capable of in-context learning, the ability to perform new tasks at test time using a handful of input-output examples, without parameter updates. We develop a universal approximation theory to elucidate how transformers enable in-context learning. For a general class of functions (each representing a distinct task), we demonstrate how to construct a transformer that, without any further weight updates, can predict based on a few noisy in-context examples with vanishingly small risk. Unlike prior work that frames transformers as approximators of optimization algorithms (e.g., gradient descent) for statistical learning tasks, we integrate Barron's universal function approximation theory with the algorithm approximator viewpoint. Our approach yields approximation guarantees that are not constrained by the effectiveness of the optimization algorithms being mimicked, extending far beyond convex problems like linear regression. The key is to show that (i) any target function can be nearly linearly represented, with small ℓ1-norm, over a set of universal features, and (ii) a transformer can be constructed to find the linear representation akin to solving Lasso at test time.

Biography: Dr. Yuting Wei is an Associate Professor in the Statistics and Data Science Department at the Wharton School, University of Pennsylvania. Prior to that, Dr. Wei spent two years at Carnegie Mellon University as an assistant professor and one year at Stanford University as a Stein's Fellow. She received her Ph.D. in statistics at the University of California, Berkeley. She was the recipient of the 2025 Gottfried E. Noether Early Career Scholar Award, Google Research Scholar Award, NSF Career award, and the Erich L. Lehmann Citation from the Berkeley statistics department. Her research interests include high-dimensional and non-parametric statistics, reinforcement learning, and diffusion models.

Biography: Yuxin Chen is currently a professor of statistics and data science and of electrical and systems engineering at the University of Pennsylvania. Before joining UPenn, he was an assistant professor of electrical and computer engineering at Princeton University. He completed his Ph.D. in Electrical Engineering at Stanford University and was also a postdoc scholar at Stanford Statistics. His current research interests include high-dimensional statistics, machine learning theory, and optimization. He has received the Alfred P. Sloan Research Fellowship, the SIAM Activity Group on Imaging Science Best Paper Prize, the ICCM Best Paper Award (gold medal), and was selected as a finalist for the Best Paper Prize for Young Researchers in Continuous Optimization. He has also received the Princeton Graduate Mentoring Award.

Speaker: Professor David Firth

Date: 21/10/2025

Abstract: A composition vector describes the relative sizes of parts of a thing. Some important modern application areas are microbiome analysis, time-use analysis and archaeometry (to name just three). We develop model-based analysis of composition, through the first two moments of measurements on their original scale. In current applied work the most-used route to compositional data analysis, following an approach introduced by the late John Aitchison in the 1980s, is based on contrasts among log-transformed measurements. The quasi-likelihood model framework developed here provides a general alternative with several advantages. These include robustness to secondary aspects of model specification, stability when there are zero-valued or near-zero measurements in the data, and more direct interpretation. Linear models for log-contrast transformed data are replaced by generalized linear models with logit link, and variance-covariance estimation is straightforward via suitably standardized residuals. (joint work with Fiona Sammut, University of Malta: https://arxiv.org/abs/2312.10548 )

Biography: David Firth is Emeritus Professor of Statistics at the University of Warwick. He is a Fellow of the British Academy, a former Editor of JRSS-B, and a recipient of the RSS Guy Medal in Silver. His work ranges from statistical theory to a variety of applications which include exit-polling at UK general elections.

Speaker: Professor Omiros Papaspiliopoulos

Date: 22/10/2026

Abstract: This talk relates to a book currently being written in collaboration with political scientist Max Goplerud (University of Texas at Austin) and focuses on variational inference and its applications to large-scale inference for generalized bilinear mixed models.

This model structure is pervasive across the applied sciences and includes, as special cases, generalized linear mixed models, matrix factorization models, and item response theory. Professor Papaspiliopoulos will present a synthesis of methodological, theoretical, and software work within this framework, and link these developments to other results in the field.

Biography: Omiros Papaspiliopoulos is a Full Professor at Bocconi University and previous to this he has held faculty positions at UPF in Barcelona and Warwick University, postdoctoral positions at Lancaster and Oxford, and has also worked in Berlin, Osaka, Paris, Madrid and Lima. He is co-editor of Biometrika since 2018, and has been an AE for JRSSB, Statistics and Computing and SIAM Journal of Uncertainty Qunatification. He received the Guy Medal in 2010 by the Royal Statistical Society, and the DeGroot prize in 2021 for his book “An Introduction to Sequential Monte Carlo” with Nicolas Chopin. He has extensive experience at designing and directing postgraduate and undergraduate programs, and at executive education.

Speaker: Dr Wei Pan

Date: 03/10/2026

Abstract: This talk explores how machine learning can enable adaptive reactive intelligence in robotic systems, addressing the challenge that robots need both reactive and deliberative capabilities. Dr Pan will present research on reinforcement learning, system identification, and multi-agent coordination, validated across platforms from quadrotors to spacecraft rendezvous.

2025

Speaker: Christophe Giraud (Université Paris-Saclay)

Date: 12/12/2025

Speaker: Yuqi Gu (Columbia University)

Date: 09/12/2025

Speaker: Yasser Roudi (King’s College London)

Date: 05/12/2025

Speaker: Kelly Zhang (Imperial College London)

Date: 27/06/2025

Abstract: Increasingly, recommender systems are tasked with improving users' long-term satisfaction. In this context, we study a content exploration task, which we formalize as a bandit problem with delayed rewards. There is an apparent trade-off in choosing the learning signal: waiting for the full reward to become available might take several weeks, slowing the rate of learning, whereas using short-term proxy rewards reflects the actual long-term goal only imperfectly. First, we develop a predictive model of delayed rewards that incorporates all information obtained to date. Rewards as well as shorter-term surrogate outcomes are combined through a Bayesian filter to obtain a probabilistic belief. Second, we devise a bandit algorithm that quickly learns to identify content aligned with long-term success using this new predictive model. We prove a regret bound for our algorithm that depends on the Value of Progressive Feedback, an information theoretic metric that captures the quality of short-term leading indicators that are observed prior to the long-term reward. We apply our approach to a podcast recommendation problem, where we seek to recommend shows that users engage with repeatedly over two months. We empirically validate that our approach significantly outperforms methods that optimize for short-term proxies or rely solely on delayed rewards, as demonstrated by an A/B test in a recommendation system that serves hundreds of millions of users.

Speaker: Fanghui Liu (University of Warwick)

Date: 20/06/2025

Abstract: In this talk, I will discuss some fundamental questions in modern machine learning:

- What is a suitable model capacity of a modern machine learning model?

- How to precisely characterise the test risk under such a model capacity?

- What is the corresponding function space induced by such a model capacity?

- What are the fundamental limits of statistical/computational learning efficiency within space?

My talk will partly answer the above questions, through the lens of norm-based capacity control. By deterministic equivalence, we provide a precise characterisation of how the estimator’s norm concentrates and how it governs the associated test risk. Our results show that the predicted learning curve admits a phase transition from under- to over-parameterisation, but no double descent behaviour, and reshapes scaling laws as well. Additionally, I will talk about the path-norm based capacities and the induced Barron spaces to understand the fundamental limits of statistical efficiency, particularly in terms of sample complexity and dimension dependence—highlighting key statistical-computational gaps. Talk is based on https://arxiv.org/abs/2502.01585, https://arxiv.org/abs/2404.18769

Speaker: Björn Andersson (University of Oslo)

Date: 06/06/2025

Abstract: Full-information maximum likelihood estimation of latent variable models with mixed data, such as a combination of categorical, count and continuous observed variables, requires computing integrals without an explicit solution. In the literature, approaches based on numerical quadrature (with or without dimension-reduction), adaptive Gauss-Hermite quadrature, Laplace approximations, or stochastic approximations have been described. Here, we discuss usage of Laplace approximations for marginal maximum likelihood estimation of generalized linear latent variable models, which support any response variable distribution in the exponential family. We outline how to efficiently implement first- and second-order Laplace approximations for these models and discuss the computational complexity in relation to the model structure. We also compare the finite-sample estimation properties and computational efficiency of these estimators to alternative methods with a set of commonly used model structures. Lastly, we outline future applications of Laplace approximations within the method of extended variational approximation.

Speaker: Hans-Georg Muller (University of California)

Date: 29/05/2025

Abstract: The underlying probability measure of random objects, i.e., metric-space valued data, can be assessed by distance profiles that correspond to one-dimensional distributions of probability mass falling into balls of increasing radius. These distance profiles have various applications, including the quantification of centrality. In the presence of Euclidean (vector) predictors, we propose to use conditional average transport costs to transport a given distance profile to all other distance profiles to obtain conditional conformity scores. Utilizing the split conformal algorithm these can be used to construct conditional prediction sets with asymptotic conditional validity. Applications include network data from New York taxi trips and compositional data on energy sourcing of U.S. states. This talk is based on joint work with Hang Zhou (Davis).

Speaker: Jane-Ling Wang (University of California)

Date: 29/05/2025

Abstract: In this talk, we present two applications of deep neural networks (DNNs) for functional data. Traditional methods require dimension reduction via pre-selected basis expansions, which may not be optimal. We propose an adaptive approach using a DNN with a basis-layer, where hidden units act as basis functions through micro neural networks. This architecture focuses on relevant information, improving dimension reduction and outperforming other DNNs in classification and regression tasks. Additionally, we demonstrate using transformers to effectively impute irregularly and sparsely observed functional data. Based on joint work with Ju-Sheng Hong, Jonas Mueller and Juwen Yao.

Speaker: Nicola Gnecco (Imperial College London)

Date: 28/03/2025

Abstract:Classical methods for quantile regression fail in cases where the quantile of interest is extreme and only few or no training data points exceed it. Asymptotic results from extreme value theory can be used to extrapolate beyond the range of the data, and several approaches exist that use linear regression, kernel methods or generalized additive models. Most of these methods break down if the predictor space has more than a few dimensions or if the regression function of extreme quantiles is complex. We propose a method for extreme quantile regression that combines the flexibility of random forests with the theory of extrapolation. Our extremal random forest (ERF) estimates the parameters of a generalized Pareto distribution, conditional on the predictor vector, by maximizing a local likelihood with weights extracted from a quantile random forest. We penalize the shape parameter in this likelihood to regularize its variability in the predictor space. Under general domain of attraction conditions, we show consistency of the estimated parameters in both the unpenalized and penalized case. Simulation studies show that our ERF outperforms both classical quantile regression methods and existing regression approaches from extreme value theory. We apply our methodology to extreme quantile prediction for U.S. wage data. This is joint work with Edossa Merga Terefe and Sebastian Engelke.

Speaker: Xi Chen (NYU Stern School of Business)

Date: 24/03/2025

Abstract: This talk has two parts. The first part is on digital privacy in personalized pricing. When involving personalized information, how to protect the privacy of such information becomes a critical issue in practice. In this talk, we consider a dynamic pricing problem with an unknown demand function of posted prices and personalized information. By leveraging the fundamental framework of differential privacy, we develop a privacy-preserving dynamic pricing policy, which tries to maximize the retailer revenue while avoiding information leakage of individual customers' information and purchasing decisions. This is joint work with Prof. Yining Wang and Prof. David Simchi-Levi. The second part introduces the concept of using blockchain to create a decentralized computing market for any AI training/fine-tuning. We introduce the concept of incentive-security that incentivizes rational trainers to behave honestly for their best interest. We design a Proof-of-Learning mechanism with computational efficiency, a provable incentive-security guarantee, and controllable difficulty. Our research also proposes an environmentally friendly verification mechanism for blockchain systems, allowing existing proof-of-work computations to be used for AI services, thus achieving useful proof-of-work.

Biography: Xi Chen is a professor and Andre-Meyer Faculty Fellow at Stern School of Business, New York University, who is also an affiliated professor at Computer Science and Center for Data Science. Before that, he was a Postdoc in the group of Prof. Michael Jordan at UC Berkeley and obtained his Ph.D. from the Machine Learning Department at Carnegie Mellon University. He studies high-dimensional machine learning, online learning, large-scale stochastic optimization, and applications to operations management and FinTech. Recently, he started a new research line on blockchain technology and decentralized finance. He is an IMS Fellow, recipient of COPSS Leadership Award, NSF Career Award, The World’s Best 40 under 40 MBA Professor by Poets & Quants, and Forbes 30 under 30 in Science. Take a look at this website too.

Speaker: Matteo Farnè (Università di Bologna)

Date: 21/03/2025

Abstract: This talk provides a comprehensive overview of the estimation framework for high-dimensional spectral density matrices via nuclear norm plus 1 norm penalization under the assumption of an underlying low rank plus sparse structure, which naturally occurs when the data follow an approximate dynamic factor model with a sparse residual autocovariance. The underlying assumptions allow for non-pervasive latent dynamic eigenvalues and a prominent residual autocovariance pattern. In that context, existing approaches based on principal components may lead to misestimate the number of factors. On the contrary, minimizing a quadratic loss under a nuclear norm plus 1 norm constraint, which controls the latent rank and the residual sparsity pattern via two specific threshold parameters, proves to be an effective optimization strategy. In this time-dependent setting, we propose a new estimator of high-dimensional spectral density matrices, called ALgebraic Spectral Estimator (ALSE). The quadratic loss function requires as input the classical smoothed periodogram estimator and prompts a consequent choice of the two threshold parameters. We prove consistency of ALSE as both the dimension and the sample size diverge to infinity, as well as the recovery of latent rank and residual sparsity pattern with probability one. We then propose the UNshrunk ALgebraic Spectral Estimator (UNALSE), which is designed to minimize the Frobenius loss with respect to the pre-estimator while retaining the optimality of the ALSE, thus presenting the tightest error bound in minimax sense. On top of that, we prove that the ensuing estimators of dynamic factor loadings and scores via Bartlett’s and Thomson’s methods have the same minimax property, and we provide the asymptotic rates for those estimators. An interesting property of such dynamic factor score estimators is that they are entirely based on frequency-domain quantities, which allows to avoid a direct control on the ergodicity and the number of lags of the common component. When applying UNALSE to a standard U.S. quarterly macroeconomic dataset, we find evidence of two main sources of comovements: a real factor driving the economy at business cycle frequencies, and a nominal factor driving the higher frequency dynamics. Dynamic factor scores can be estimated by UNALSE spectral density matrices, thus depicting the behaviour of each of the two factors over time.

Speaker: Shuangning Li (University of Chicago Booth School of Business)

Date: 14/03/2025

Abstract: The field of causal inference develops methods for estimating treatment effects, often relying on the Stable Unit Treatment Value Assumption (SUTVA), which states that a unit’s outcome depends only on its own treatment. However, in many real-world settings, SUTVA is violated due to interference—where the treatment assigned to one unit influences the outcomes of others. Such interference can arise from social interactions among units or competition for shared resources, complicating causal analysis and leading to biased estimates. Fortunately, in many cases, interference follows structured patterns that can potentially be leveraged for more accurate estimation. In this paper, we examine and formalize two specific forms of structured interference—monotone interference and submodular interference—which we believe arise in many practical settings. We investigate how incorporating these structures can improve causal effect estimation. Our main contributions are (i) a set of bounds relating key interference estimands under these structural assumptions and (ii) new estimators that integrate these structures through constrained optimization. Since these constraints may introduce bias, we further develop debiasing techniques based on treatment regeneration and bootstrap methods to mitigate this issue. This is joint work (ongoing) with Kevin Han and Johan Ugander.

Biography: Shuangning Li is an Assistant Professor of Econometrics and Statistics at the University of Chicago Booth School of Business. Previously, she was a postdoctoral fellow in the Department of Statistics at Harvard University. She earned her Ph.D. in Statistics from Stanford University, advised by Professors Emmanuel Candès and Stefan Wager. Before that, she obtained a Bachelor of Science from the University of Hong Kong. Her research focuses on causal inference, multiple hypothesis testing, selective inference, and statistical reinforcement learning.

Speaker: Wen Zhou (NYU GPH)

Date: 07/03/2025

Abstract: In network analysis, noises and biases, which are often introduced by peripheral or non-essential components, can mask pivotal structures and hinder the efficacy of many network modeling and inference procedures. Recognizing this, identification of the core--periphery (CP) structure has emerged as a crucial data pre-processing step. While the identification of the CP structure has been instrumental in pinpointing core structures within networks, its application to directed weighted networks has been underexplored. Many existing efforts either fail to account for the directionality or lack the theoretical justification of the identification procedure. In this work, we seek answers to three pressing questions: (i) How to distinguish the informative and noninformative structures in weighted directed networks? (ii) What approach offers computational efficiency in discerning these components? (iii) Upon the detection of CP structure, can uncertainty be quantified to evaluate the detection? We adopt the signal-plus-noise model, categorizing different types of noninformative relational patterns, by which we define the sender and receiver peripheries. Furthermore, instead of confining the core component to a specific structure, we consider it complementary to either the sender or receiver peripheries. Based on our definitions on the sender and receiver peripheries, we propose spectral algorithms to identify the CP structure in directed weighted networks. Our algorithm stands out with statistical guarantees, ensuring the identification of sender and receiver peripheries with overwhelming probability. Additionally, we propose a hypothesis testing framework to infer CP structure upon detection. Our methods scale effectively for expansive directed networks. Implementing our methodology on faculty hiring network data revealed captivating insights into the informative structures and distinctions between informative and noninformative sender/receiver nodes across various academic disciplines. This is a joint work with Wenqin Du, Tianxi Li, and Lihua Lei.

Biography: Wen Zhou is an Associate Professor in the Department of Biostatistics at the School of Global Public Health. He received his Ph.D.s in Statistics and Applied Mathematics from the Iowa State University. His research focuses on developing theories and methods for network data analysis, high-dimensional statistics, machine learning, and causal inference. He is interested in applications within genomics, genetics, protein structure modeling, social science, and health policy. Wen serves on the editorial boards of the Statistica Sinica, Journal of Multivariate Analysis, Biometrics, as well as serves as the Editor-in-Chief of Journal of Biopharmaceutical Statistics. He is an elected member of the International Statistical Institute and has been elected as the WNAR program coordinator in 2024.

Speaker: Ruth Heller (Tel-Aviv University)

Date: 07/03/2025

Abstract: We begin by introducing the problem of multiple testing and selective inference, emphasizing thecentral roles of the false discovery rate (FDR) and the false coverage rate (FCR) in the analysis ofhigh-dimensional problems. We then address the important roles that the FDR and FCR may playin prediction problems, particularly in supervised learning tasks, including regression andclassification. In these contexts, conformal methods provide prediction sets for outcomes orlabels with finite-sample coverage guarantees for any machine learning predictor.Our focus is on cases where such prediction sets are constructed following a selection process.This selection process requires that the selected prediction sets be "informative" in a well-definedsense. We explore both classification and regression settings, where the analyst may defineinformative prediction sets as those that are sufficiently small, exclude null values, or satisfy otherappropriate monotone constraints. We introduce InfoSP and InfoSCOP, novel procedures thatprovide FCR control for informative prediction sets. We demonstrate the utility of these methodsthrough applications to both real and simulated data. While our primary focus is on the "batch"setting, we also address the ubiquitous "online" learning setting. Joint work with Cherief-Abdeltaif, B. , Gazin, U., Humbert, P., Marandon, A., and Roquain, E.

Speaker: Eleni Matechou (University of Kent)

Date: 21/02/2025

Abstract: Ecological surveys often track individuals or species to monitor time-varying processes such as migration patterns and changes in behavioural or life states. Mixture models are a suitable and flexible approach for analysing such data, and they have been used extensively in the field. In this seminar, parametric, nonparametric, and repulsive mixture models are discussed for different types of ecological data and results are presented for case studies on species monitored using standard ecological data, as well as data collected using new technologies.

Speaker: Solenne Gaucher (Centre de Mathématiques Appliquées at École Polytechnique)

Date: 14/02/2025

Abstract: Artificial intelligence (AI) is increasingly shaping the decisions that affect our lives—from hiring and education to healthcare and access to social services. While AI promises efficiency and objectivity, it also carries the risk of perpetuating and even amplifying societal biases embedded in the data used to train these systems. Algorithmic fairness aims to design and analyze algorithms capable of providing predictions that are both reliable and equitable. In this talk, I will introduce one of the main approaches to achieving this goal: statistical fairness. After outlining the basic principles of this approach, I will focus specifically on a fairness criterion known as "demographic parity," which seeks to ensure that the distribution of predictions is identical across different populations. I will then discuss recent results related to regression and classification problems under this fairness constraint, exploring scenarios where differentiated treatment of populations is either permitted or prohibited.

Biography: Solenne Gaucher is an assistant professor at the Centre de Mathématiques Appliquées at École polytechnique. Her work focuses on developing fair machine learning algorithms to adresse algorithmic biaises and mitigate their broader societal impact. Prior to this role, she completed a postdoctoral fellowship at the Center for Research in Economics and Statistics (CREST) at ENSAE Paris. Solenne holds a Ph.D. in mathematics from Université Paris-Saclay. She was awarded the "Young French Talent" prize by the L'Oréal-UNESCO Foundation For Women in Science.

Speaker: Karla Diaz Ordaz (UCL)

Date: 07/02/2025

Abstract: Instrumental variable methods are very popular in econometrics and biostatistics for inferring causal average effects of an exposure on an outcome where there is unmeasured confounding. However, their application for learning heterogeneous treatment effects, such as conditional average treatment effects (CATE), in combination with machine learning, is somewhat limited. A generic approach that allows the use of arbitrary machine learning algorithms can be based on the popular two-stage principle. We first "regress" the exposure on the instrumental variables (and pre-exposure covariates) and then learn the causal treatment effects by regressing the outcome on the predicted exposure. This is the approach of Foster and Syrgkanis (2023), referred to as IV-debiased machine learning (IV-DML). Unfortunately, the slow convergence rates of the data-adaptive estimators that affect the first-stage predictions propagate into the resulting CATE estimates. In view of this, we propose the IV-learner, inspired by infinite-dimensional targeted learning procedures (Vansteelandt 2023, van der Laan et al 2024). It strategically targets the first-stage predictions so they perform well in their ultimate task: CATE estimation. The resulting learner is easy to construct based on arbitrary, off-the-shelf algorithms. We study the finite sample performance of our proposal using simulations, and compare it to existing methods. We also illustrate it using a real data example. Joint work with Stijn Vansteelandt, Stephen O’Neill, Richard Grieve.

Speaker: Karla Diaz Ordaz (UCL)

Date: 31/01/2025

Abstract: To enable closed form conditioning, a common assumption in Gaussian process (GP) regression is independent and identically distributed Gaussian observation noise. This strong and simplistic assumption is often violated in practice, which leads to unreliable inferences and uncertainty quantification. Unfortunately, existing methods for robustifying GPs break closed-form conditioning, which makes them less attractive to practitioners and significantly more computationally expensive. In this work, we demonstrate how to perform provably robust and conjugate Gaussian process (RCGP) regression at virtually no additional cost using generalised Bayesian inference. RCGP is particularly versatile as it enables exact conjugate closed form updates in all settings where standard GPs admit them. To demonstrate its strong empirical performance, we deploy RCGP for problems ranging from Bayesian optimisation to sparse variational Gaussian processes.

Speaker: Peter Orbanz (UCL)

Date: 24/01/2025

Abstract: The term Gaussian universality refers to a class of results that are, loosely speaking, generalized central limit theorems (where, somewhat confusingly, the limit law is not necessarily Gaussian). They provide useful tools to study certain problems in machine learning. I will give a short overview of this idea and present two types of results: One are upper and lower bounds that map out where Gaussian universality is applicable and what rates of convergence one can expect. The other is the use of these techniques to obtain quantitative results on the effects of data augmentation in machine learning problems. This is joint work with KH Huang (Gatsby Unit) and M Austern (Harvard).

Speaker: Nicolas Verzelen (INRAE)

Date: 17/01/2025

Abstract: We investigate the existence of a fundamental computation-information gap for the problem of clustering a mixture of isotropic Gaussian in the high-dimensional regime, where the ambientdimension p is larger than the number n of points. The existence of a computation-information gap ina specific Bayesian high-dimensional asymptotic regime has been conjectured by Lesieur et al. (2016)based on the replica heuristic from statistical physics. We provide evidence of the existence of such agap generically in the high-dimensional regime p >n, by (i) proving a non-asymptotic low-degreepolynomials computational barrier for clustering in high-dimension, matching the performance of thebest known polynomial time algorithms, and by (ii) establishing that the information barrier forclustering is smaller than the computational barrier, when the number K of clusters is large enough.These results are in contrast with the (moderately) low-dimensional regime n> poly(p,K) where there isno computation-information gap for clustering a mixture of isotropic Gaussian. This is based on a jointwork with Bertrand Even and Christophe Giraud (Paris-Saclay).