Conference Article

·2022 OPEN ACCESS

Investigating Semi-Supervised Learning Algorithms in Text Datasets

H. Toprak Kesgin YTU , Mehmet Fatih Amasyalı YTU

2022 Innovations in Intelligent Systems and Applications Conference (ASYU)

Abstract

Using large training datasets enhances the generalization capabilities of neural networks. Semi-supervised learning (SSL) is useful when there are few labeled data and a lot of unlabeled data. SSL methods that use data augmentation are most successful for image datasets. In contrast, texts do not have consistent augmentation methods as images. Consequently, methods that use augmentation are not as effective in text data as they are in image data. In this study, we compared SSL algorithms that do not require augmentation; these are self-training, co-training, tri-training, and tri-training with disagreement. In the experiments, we used 4 different text datasets for different tasks. We examined the algorithms from a variety of perspectives by asking experiment questions and suggested several improvements. Among the algorithms, tri-training with disagreement showed the closest performance to the Oracle; however, performance gap shows that new semi-supervised algorithms or improvements in existing methods are needed.

Keywords

Computer science Oracle Generalization Machine learning Artificial intelligence Variety (cybernetics) Artificial neural network Image (mathematics) Contrast (vision) Training set Labeled data Pattern recognition (psychology) Data mining Mathematics

Subject Areas

Domain Adaptation and Few-Shot Learning ·Artificial Intelligence ·Physical Sciences
Text and Document Classification Technologies ·Artificial Intelligence ·Physical Sciences
Topic Modeling ·Artificial Intelligence ·Physical Sciences

OpenAlex SDG Match

SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).

Quality Education 47%