SIMBA: A robust and generalizable measure of data imbalance

Publication date

2025-12-12

Authors

Pivin-Bachler, JulieORCID 0009-0005-9624-5057ISNI 0000000524651833
van den Broek, EgonORCID 0000-0002-2017-0141ISNI 0000000395166232

Editors

Advisors

Supervisors

Document Type

Article
Open Access logo

License

cc_by

Abstract

Ranging from health to cybersecurity, real-world data are heavily imbalanced. Handling imbalance is among the formidable challenges of machine learning (ML), as it deteriorates ML’s performance, yielding biased results toward majority classes. However, finding an adequate measure to assess the impact of data imbalance is a field of research by itself. Following a review of the available imbalance measures, we introduce the status of imbalance (SIMBA), which considers data distribution and overlap, both of which are crucial to assess the impact of imbalance. SIMBA is benchmarked against seven imbalance measures on five ML models, 428 synthetic and 70 non-synthetic datasets from various domains. Resulting correlation coefficients between imbalance measures and classification performance and an analysis with 20 complexity measures prove that SIMBA consistently outperforms other measures. Overall, SIMBA accurately quantifies multiclass data imbalance and may help alleviate ML data imbalance challenges in the future.

Keywords

benchmark, data complexity, data distribution, data imbalance, domain generalization, feature importance, machine learning, measure, SIMBA, status of imbalance, survey, General Decision Sciences

Citation

Pivin-Bachler, J R & van den Broek, E L 2025, 'SIMBA : A robust and generalizable measure of data imbalance', Patterns, vol. 6, no. 12, 101395. https://doi.org/10.1016/j.patter.2025.101395