Abstract
This study focuses on the task of Speech Emotion Recognition (SER) that utilizes only the acoustic modality to perform emotional analysis. The previous studies have been conducted on the SER domain showed that neutrality is the most difficult emotional state to recognize. Pointing to the incapabilities of the neutral emotional state recognition, we propose to employ the Log-Mel Spectrogram of the speech signals with an attentive utilization of the signal's auto-correlation features. The empirical studies conducted on two benchmark data sets, i.e. IEMOCAP and RAVDESS, proved the proposed approach with a major advancement on the neutral emotional state predictions compared to the previous studies. Although there was some inconsistency, improvements are observed on the other emotions by using the auto-correlation features. The inconsistency is discussed and associated with the sensitivity of the auto-correlation response to the fixed window length.
Keywords
Subject Areas
OpenAlex SDG Match
SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).