Journal Article

·2009

Measurement of turkish word semantic similarity and text categorization application

Mehmet Fatih Amasyalı YTU , Aytunç Beken YTU

Abstract

In literature, texts to be classified are generally represented in the large dimensional bag of words space in which every dimension equals to a word or ngram. In this study, firstly the words are placed in a semantic space. The word's coordinates in semantic spaces needs the similarity of the words according to their meanings. Harris states that two words' semantic similarity is related to the number of documents which the words are both in. We used his hypothesis for Turkish words. Firstly, we obtained word co-occurrence matrix from a Web corpus. Then, the numerical coordinates of the words are calculated by using multi dimensional scaling. Texts coordinates are obtained from word coordinates which passes in the texts. In our experiments, Turkish news texts are classified into 5 classes. We get more successful results than the traditional bag of words space. Our approach is not for only Turkish words/texts, but also for all other languages.

Keywords

Turkish Categorization Semantic similarity Similarity (geometry) Word (group theory) Natural language processing Artificial intelligence Computer science Dimension (graph theory) Multidimensional scaling Space (punctuation) Semantic space Text categorization Semantics (computer science) Linguistics Mathematics

Subject Areas

Advanced Text Analysis Techniques ·Artificial Intelligence ·Physical Sciences
Text and Document Classification Technologies ·Artificial Intelligence ·Physical Sciences
Natural Language Processing Techniques ·Artificial Intelligence ·Physical Sciences

Citations by Year

OpenAlex SDG Match

SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).

Quality Education 81%