FIONA: Detecting Syntactical Outliers in Attributes with Categorical Values

Publication date

2025-04-17

Authors

Tsiamis, Thanos
Qahtan, HakimORCID 0000-0001-8254-1764ISNI 0000000492915493

Editors

Advisors

Supervisors

Document Type

Contribution to conference
Open Access logo

License

taverne

Abstract

Outlier detection is crucial for data cleaning, influencing analysis and decision-making. While numerical outlier detection is well-studied, identifying outliers in relational data with categorical attributes poses greater challenges due to difficulties in defining a suitable similarity measure. Current approaches for detecting categorical outliers are based on coding the categorical values as numerical values, using the frequency as an indicator of the outlierness score and extracting predefined syntactic structures of the values. In this paper, we propose FIONA (FInding Outliers iN Attributes) to detect outliers in attributes with categorical values. Since categorical values in the relational model usually follow specific syntactic structures, FIONA defines a similarity measure that can reveal the hidden patterns and identify a set of dominant patterns in the data. Values that do not conform to the dominating patterns are declared as outliers. In comparison to alternative tools, FIONA accurately identifies outliers and dominant patterns within datasets and provides a clear explanation for declaring a given value as an outlier.

Keywords

Categorical outliers, generalization tree, patterns, similarity measures, syntactic structure, Taverne

Citation

Tsiamis, T & Qahtan, H 2025, 'FIONA: Detecting Syntactical Outliers in Attributes with Categorical Values', Paper presented at 20th International Conference, ICDATA 2024, Held as Part of the World Congress in Computer Science, Las Vegas, 22/07/24 - 25/07/24 pp. 432-448. https://doi.org/10.1007/978-3-031-85856-7_32, conference