Journal Article

·2022 OPEN ACCESS

IMPACT OF N-STAGE LATENT DIRICHLET ALLOCATION ON ANALYSIS OF HEADLINE CLASSIFICATION

Zekeriya Anıl Güven , Banu Di̇ri̇ YTU , Tolgahan Çakaloğlu

Computer Science

Abstract

Data analysis becomes difficult with the increase of large amounts of data. More specifically, extracting meaningful insights from this vast amount of data and grouping them based on their shared features without human intervention requires advanced methodologies. There are topic modeling methods to overcome this problem in text analysis for downstream tasks, such as sentiment analysis, spam detection, and news classification. In this research, we benchmark several classifiers, namely Random Forest, AdaBoost, Naive Bayes, and Logistic Regression, using the classical LDA and n-stage LDA topic modeling methods for feature extraction in headlines classification. We run our experiments on 3 and 5 classes publicly available Turkish and English datasets. We demonstrate that n-stage LDA as a feature extractor obtains state-of-the-art performance for any downstream classifier. It should also be noted that Random Forest was the most successful algorithm for both datasets.

Keywords

Computer science Latent Dirichlet allocation Naive Bayes classifier Random forest Artificial intelligence Classifier (UML) Overfitting Extractor Topic model Machine learning Pattern recognition (psychology) Linear discriminant analysis Data mining Support vector machine Artificial neural network

Subject Areas

Text and Document Classification Technologies ·Artificial Intelligence ·Physical Sciences
Sentiment Analysis and Opinion Mining ·Artificial Intelligence ·Physical Sciences
Web Data Mining and Analysis ·Information Systems ·Physical Sciences

OpenAlex SDG Match

SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).

Life in Land 70%