Understanding classification
Publication date
2009-01-14
Authors
Subianto, M.
Editors
Advisors
DOI
Document Type
Dissertation
Metadata
Show full item recordCollections
License
Abstract
In practical data analysis, the understandability of models plays an important role in their acceptance. In the data mining literature, however, understandability plays is hardly ever mentioned. If it is mentioned, it is interpreted as meaning that the models have to be simple. In this thesis we argue that understandability should be interpreted as explainable. That is, for predictive models we should be able to explain the predictions made by the model. For a classifier this means, e.g., that we can explain which features were decisive for the classification. For example, a loan was refused because the prospective client is unemployed. In this thesis we introduce various tools that provide such explanations for classifiers of arbitrary type. Visualization through Mosaic plots forms an important ingredient in these tools. Throughout the thesis, the tools are illustrated on a small, well-known, data set from the UCI repository, viz., the voting data set. In the final chapter, a real case study is presented on the discovery of small genes.
Keywords
Citation
Subianto, M 2009, 'Understanding classification', Doctor of Philosophy, Utrecht University.