Conference Article

·2022

Improving BERT Pre-training with Hard Negative Pairs

Yusuf Ziya Poyraz YTU , Miraç Tuğcu YTU , Mehmet Fatih Amasyalı YTU

2022 Innovations in Intelligent Systems and Applications Conference (ASYU)

Abstract

In this paper, we ran various experiments on BERT's pre-training tasks and observed their impact on a language model's success on different downstream tasks such as masked word prediction, sentiment analysis, named entity recognition and text classification. Also, an improvement method called Hard Negative Pairs (HNP) is suggested to increase the success of the Same Sentence Prediction (SSP) task. The goal of HNP is to pick negative pairs that are more similar to the original sentence. After that, two additional improvements for SSP are proposed. One of these methods aims to create shorter sequences while the other one's main goal is selecting a splitting point other than the middle of the sentence. Experiments were performed for these three different suggestions and their different combinations. The results show that the SSP model that uses HNP and smaller sequence generation methods has an improvement over the original SSP. The longest trained model on 1 GB data gets closer results to the BERTurk, even though BERTurk trained with 35 times more data.

Keywords

Computer science Sentence Task (project management) Word (group theory) Natural language processing Artificial intelligence Point (geometry) Training set Sequence (biology) Language model Sequence labeling Speech recognition Linguistics Mathematics

Subject Areas

Topic Modeling ·Artificial Intelligence ·Physical Sciences
Natural Language Processing Techniques ·Artificial Intelligence ·Physical Sciences
Text and Document Classification Technologies ·Artificial Intelligence ·Physical Sciences

Citations by Year

OpenAlex SDG Match

SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).

Quality Education 75%