Journal Article

·2006

An Approach for Word Categorization Based on Semantic Similarity Measure Obtained from Search Engines

Mehmet Fatih Amasyalı YTU

Abstract

Word categorization based on semantic similarity is a problem need to be solved for several natural language applications. A similarity measure is need for word categorization. In this study it is proposed that the semantic similarity between two Turkish words is in direct proportion to the number of pages which the words are located next to each other. Google and Yahoo search engines were used to find the number of pages. In the first attempt to verify the proposal, the experiments were done with small datasets. The average success ratio is 87%.

Keywords

Categorization Semantic similarity Computer science Similarity (geometry) Word (group theory) Natural language processing Turkish Information retrieval Measure (data warehouse) Artificial intelligence Text categorization Similarity measure Search engine Explicit semantic analysis Semantic search Semantic computing Data mining Semantic technology Semantic Web Mathematics Linguistics

Subject Areas

Text and Document Classification Technologies ·Artificial Intelligence ·Physical Sciences
Advanced Text Analysis Techniques ·Artificial Intelligence ·Physical Sciences
Natural Language Processing Techniques ·Artificial Intelligence ·Physical Sciences

Citations by Year