Deep Learning under Data Scarcity: Methods for Sparse, Imbalanced, and Shifted Data

Abstract

The effectiveness of deep learning is highly sensitive to data availability and distribution. The performance can degrade when real-world data are limited, label imbalanced, incomplete, or unevenly distributed across the input space. In this work, data scarcity is understood not only as a lack of samples, but more broadly as an insufficiency of informative and usable data for reliable learning. These situations undermine the effectiveness of standard deep learning methods, which typically rely on large, representative, and stationary datasets. To investigate these issues, this dissertation examines two representative application domains where data scarcity and distribution shifts are especially common: financial risk assessment and weather nowcasting. Building on these challenges, the thesis develops robust deep learning approaches that maintain predictive performance under data scarcity. We begin by addressing the challenge of domain shifts and cold-start scenarios in finance, where scarce and imbalanced historical data limit a model’s generalization ability. To address this, we introduce TransCORALNet, a two-stream deep domain adaptation framework that combines correlation alignment loss, attention mechanisms, and synthetic data generation. This approach reduces distribution discrepancies while mitigating class imbalance, enabling performance stability even when target domains diverge from the source. We then turn to the problem of learning when data are distributed, isolated, or cannot be shared due to privacy, regulatory, or organizational constraints. To address this challenge, we develop an explainable federated learning framework, Trans-XFed, which integrates homomorphically encrypted model updates with a performance-based client selection strategy and transformer-based representation learning. This design enables decentralized entities to collaboratively train predictive models without exposing raw data while simultaneously providing interpretability mechanisms for decision support. Next, we focus on the challenge of predicting rare and extreme events, which data-driven models often struggle to represent accurately. We propose GA-SmaAt-GNet, a novel generative adversarial framework for extreme precipitation nowcasting. The model integrates precipitation masks to enhance predictive accuracy and incorporates an attention-augmented discriminator inspired by the Pix2Pix architecture. In addition, uncertainty analysis and Grad-CAM–based visual explanations provide insight into model behavior and highlight the spatial regions that drive predictions. Finally, we integrate multi-variable weather station data with radar data to enhance nowcasting performance. We introduce two architectures for multi-source data integration. SmaAt-fUsion extends the SmaAt-UNet framework by incorporating weather station data into the network bottleneck. SmaAt-Krige-GNet combines precipitation maps with weather station observations processed via Kriging to generate spatial maps, which are then integrated through a dual-encoder architecture. Together, these models enable deep networks to learn coherent spatial representations from sparsely observed. Complementing these methodological contributions, the dissertation places strong emphasis on interpretability and reliability. Self-attention analysis, gradient-based attribution methods, model-agnostic explanation techniques, and uncertainty estimation are employed to examine model behavior and assess predictive confidence under data scarcity. Comprehensive experiments demonstrate that appropriately designed deep learning models can maintain reliable performance despite limited data, class imbalance, domain shift, privacy, and rare events. Overall, this dissertation illustrates how deep learning can be adapted to data-scarce settings and provides practical guidance for developing dependable models when data are incomplete or unevenly distributed.

Keywords

Dataschaarste, Domeinadaptatie, Uitlegbare AI, Federated Learning, Extreme gebeurtenissen, Datafusie., Data Scarcity, Domain adaptation, Explainable, Federated learning, Extreme event, Data fusion.

Citation

Shi, J 2026, 'Deep Learning under Data Scarcity: Methods for Sparse, Imbalanced, and Shifted Data', Doctor of Philosophy, Universiteit Utrecht, Utrecht. https://doi.org/10.33540/3661