Journal Article

·2012

Comparing the impacts of dimension reduction methods that use class labels on text classification

Göksel Biricik YTU

Abstract

Classification of datasets that contain samples with numerous features is known as a costly process in time and space. In order to overcome this problem, dimensionality reduction techniques like feature selection and feature extraction are proposed in literature. In this paper, we compare the impacts of abstract feature extraction method and other popular techniques that use class labels for dimensionality reduction on classification performances. For evaluation, we utilize two standard text datasets having high dimensional samples. We compare the impacts of selected methods on performance by applying them on selected datasets and testing on five different classifiers with different design approaches. Results show that using abstract feature extraction method for dimensionality reduction produces much better classification performance, when compared with other selected methods.

Keywords

Dimensionality reduction Computer science Feature extraction Artificial intelligence Pattern recognition (psychology) Class (philosophy) Feature selection Curse of dimensionality Reduction (mathematics) Dimension (graph theory) Process (computing) Data mining Feature (linguistics) Machine learning Mathematics

Subject Areas

Text and Document Classification Technologies ·Artificial Intelligence ·Physical Sciences
Face and Expression Recognition ·Computer Vision and Pattern Recognition ·Physical Sciences
Spam and Phishing Detection ·Information Systems ·Physical Sciences

OpenAlex SDG Match

SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).

Quality Education 53%