Why good data analysts need to be critical synthesists: Determining the role of semantics in data analysis

Publication date

2017-07

Authors

Scheider, SimonORCID 0000-0002-2267-4810ISNI 0000000382824363
Ostermann, Frank
Adams, Benjamin

Editors

Advisors

Supervisors

Document Type

Article
Open Access logo

License

Abstract

In this article, we critically examine the role of semantic technology in data driven analysis. We explain why learning from data is more than just analyzing data, including also a number of essential synthetic parts that suggest a revision of George Box’s model of data analysis in statistics. We review arguments from statistical learning under uncertainty, workflow reproducibility, as well as from philosophy of science, and propose an alternative, synthetic learning model that takes into account semantic conflicts, observation, biased model and data selection, as well as interpretation into background knowledge. The model highlights and clarifies the different roles that semantic technology may have in fostering reproduction and reuse of data analysis across communities of practice under the conditions of informational uncertainty. We also investigate the role of semantic technology in current analysis and workflow tools, compare it with the requirements of our model, and conclude with a roadmap of 8 challenging research problems which currently seem largely unaddressed.

Keywords

Data driven analysis, Learning, Semantic Web, e-Science, Data science

Citation

Scheider, S, Ostermann, F & Adams, B 2017, 'Why good data analysts need to be critical synthesists : Determining the role of semantics in data analysis', Future Generation Computer Systems, vol. 72, pp. 11-22. https://doi.org/10.1016/j.future.2017.02.046