FIONA: Detecting Syntactical Outliers in Attributes with Categorical Values
Publication date
2025-04-17
Editors
Advisors
Supervisors
Document Type
Contribution to conference
Metadata
Show full item recordCollections
License
taverne
Abstract
Outlier detection is crucial for data cleaning, influencing analysis and decision-making. While numerical outlier detection is well-studied, identifying outliers in relational data with categorical attributes poses greater challenges due to difficulties in defining a suitable similarity measure. Current approaches for detecting categorical outliers are based on coding the categorical values as numerical values, using the frequency as an indicator of the outlierness score and extracting predefined syntactic structures of the values. In this paper, we propose FIONA (FInding Outliers iN Attributes) to detect outliers in attributes with categorical values. Since categorical values in the relational model usually follow specific syntactic structures, FIONA defines a similarity measure that can reveal the hidden patterns and identify a set of dominant patterns in the data. Values that do not conform to the dominating patterns are declared as outliers. In comparison to alternative tools, FIONA accurately identifies outliers and dominant patterns within datasets and provides a clear explanation for declaring a given value as an outlier.
Keywords
Categorical outliers, generalization tree, patterns, similarity measures, syntactic structure, Taverne
Citation
Tsiamis, T & Qahtan, H 2025, 'FIONA: Detecting Syntactical Outliers in Attributes with Categorical Values', Paper presented at 20th International Conference, ICDATA 2024, Held as Part of the World Congress in Computer Science, Las Vegas, 22/07/24 - 25/07/24 pp. 432-448. https://doi.org/10.1007/978-3-031-85856-7_32, conference