Journal Article

·2025

Creating a Large Clean Web Corpus for Turkish

Muhammet Uzun YTU , Yusuf Sait Erdem YTU , Tolga İzdaş YTU , Ömerhan Sancak YTU , Ahmed Zeer YTU , Elif İnce YTU , Atahan Uz YTU , Osama Shbib YTU , Eren Doğan YTU , H. Toprak Kesgin YTU ,

Abstract

In this study, it is aimed to create a high-quality dataset to improve the performance of Turkish language models and to investigate the effect of this dataset on language model training. The irregular, context-free and noisy structure of webbased data can negatively affect the success of large language models. To address this issue, a comprehensive cleaning and filtering process has been performed on the existing CulturaX dataset. In this process, page-based and content-based filtering methods have been applied to make the dataset more consistent and meaningful. In addition, supervised machine learning models have been trained using various feature sets such as FastText and BERT embeddings, and the contributions of these feature sets to model performance have been compared. Experimental results have shown that the model that uses FastText embeddings and heuristic features together has achieved the highest accuracy and F1 scores. The resulting 125 GB cleaned Turkish web corpus provided lower perplexity and higher accuracy rates in the training of language models. This study provides a significant contribution to the development of more reliable and effective large language models for Turkish, thus providing a solid foundation for use in language model research.

Keywords

Perplexity Turkish Language model Feature (linguistics) Heuristic Process (computing) Language identification Feature engineering Computer science Artificial intelligence Natural language processing

Subject Areas

Natural Language Processing Techniques ·Artificial Intelligence ·Physical Sciences
Web Data Mining and Analysis ·Information Systems ·Physical Sciences