CONAN : Text Mining in the Biomedical Domain

Publication date

2006-09-11

Authors

Malik, R.

Editors

Advisors

Supervisors

DOI

Document Type

Dissertation
Open Access logo

License

Abstract

This thesis is about Text Mining. Extracting important information from literature. In the last years, the number of biomedical articles and journals is growing exponentially. Scientists might not find the information they want because of the large number of publications. Therefore a system was constructed that automatically extracts very important information from text. This information includes gene- and protein names, (point) mutations, protein-protein interactions and biologically interesting keywords like tissues or diseases. We named this system CONAN. In contrast to other text mining systems, CONAN does not introduce new algorithms, but takes well-known programs and combines them to get the best available results. Integration of new methods and algorithms is very easy. We show that with this approach, CONAN performs very well in all evaluations. We also show the various applications of CONAN, a command-line tool, a webserver and the integration of the data in a gene-interaction network.

Keywords

text mining, literature mining, knowledge discovery, interaction mining, interaction networks

Citation