Journal Article

·2010

Text2arff: Automatic feature extraction software for Turkish texts

Mehmet Fatih Amasyalı YTU , Feruz Davletov YTU , Ars Ian Torayew YTU , Umit Ciftci YTU

Abstract

Which features are the most important for the text classification tasks? In the automatic text categorization area, several studies seek answers to this question. In this paper, a feature extraction tool for Turkish texts (Text2arff) is presented. The toolbox automatically extracts several features such as the frequencies of the words and ngrams, word clustering, Latent semantic indexing etc. The features of the texts are saved in arff (WEKA) file format. The arff files can be used easily with WEKA machine learning library.

Keywords

Computer science Toolbox Artificial intelligence Turkish Cluster analysis Search engine indexing Feature extraction Natural language processing Categorization Word (group theory) Feature (linguistics) Software Information retrieval Programming language

Subject Areas

Text and Document Classification Technologies ·Artificial Intelligence ·Physical Sciences
Natural Language Processing Techniques ·Artificial Intelligence ·Physical Sciences
Advanced Text Analysis Techniques ·Artificial Intelligence ·Physical Sciences

Citations by Year

OpenAlex SDG Match

SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).

Quality Education 75%