M.Sc. Marlon Tobaben defends his PhD thesis “Privacy in Data-Efficient Deep Learning” on Friday the 26th of June 2026 at 12 in the University of Helsinki Physicum building, Auditorium E204 (Gustaf Hällströmin katu 2, 2nd floor). His opponent is Tenure Track Faculty Franziska Boenisch (CISPA Helmholtz Center for Information Security, Germany) and custos Professor Antti Honkela (University of Helsinki). The defence will be held in English.
The thesis of Marlon Tobaben is a part of research done in the Department of Computer Science and in the Trustworthy Machine Learning group at the University of Helsinki. His supervisor has been Professor Antti Honkela (University of Helsinki).
Privacy in Data-Efficient Deep Learning
Machine learning (ML) aims at learning from data and is being adopted to many aspects of our lives. While there are many risks of adapting ML models, the risk we consider in this dissertation is exposing an individual's or group's personal information contained in the training data of such models. We consider attacks on the training data and defences to mitigate them. We focus our study on data-efficient deep learning, where artificial neural networks are trained on relatively small training datasets through techniques such as transfer learning.
To study the privacy risk of ML models, we employ membership inference attacks (MIAs), which aim to infer if a particular datapoint is part of a dataset. To mitigate MIAs or other privacy attacks, we train our ML models using the formal privacy framework differential privacy (DP). Sufficiently strong DP training results in similar ML models, regardless of whether an individual's or group's information is contained in the training data, mitigating privacy attacks.
We introduce an experimental framework for few-shot transfer learning, which is used in both the first and second article. The framework focuses on image classification with pre-trained ML models that are fine-tuned in settings where only few-shot datasets, which contain a small number of examples per class, are available.
In our first article, we study the rate of change in vulnerability to MIAs when training dataset properties (examples per class and number of classes) change. We fit a regression model to the vulnerability data. Using a mapping from DP to MIA metrics, we estimate that for worst-case protection, extremely high numbers of examples per class are needed to achieve strong privacy.
In our second article, we study the utility impact of small numbers of examples per class and of distribution shift between pre-training and fine-tuning data. We also compare fine-tuning different parameter subsets of the pre-trained model. We find that surprisingly little data is required even under strong DP when the distribution shift is small, which is much less than what prior work has considered. When the distribution shift increases, more data is required and fine-tuning only the last layer is not sufficient.
In our third article, we introduce DP to the application of Document Visual Question Answering, where ML models respond to natural language questions about documents. We motivate the privacy protection for groups of documents due to shared sensitive information, adapt training algorithms to this protection, and train high utility models.
Availability of the dissertation
An electronic version of the doctoral dissertation will be available in the University of Helsinki open repository Helda at
Printed copies will be available on request from Marlon Tobaben: