TY - GEN
T1 - Harnessing the Power of CNN-Transformer Encoders in Stress Speech Analysis
AU - Shabiyya, Syifa'Hukma H.
AU - Prasetio, Barlian Henryranu
AU - Widasari, Edita Rosana
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - Stress is the physiological response to mental, emotional, or physical stress, which varies between individuals. A survey by Ipsos Global showed that around 30% of respondents identified stress as a significant health issue. Some countries in Southeast Asia, such as Cambodia, have much higher rates of depression than the world average. In Indonesia, the stress rate reached 9.8% in 2018. This research focuses on Speech Stress Recognition (SSR), an automated method that recognizes stress levels through speech characteristic analysis. We use Mel-Frequency Cepstral Coefficients (MFCC) feature extraction and the CNN-Transformer Encoder model. Evaluation results on the SUSAS dataset showed an overall accuracy of 73.76%. When the classification results are viewed by gender, male data appears better at classifying stress levels than female data. To improve performance, we implemented the Voice Activity Detection method, which resulted in an accuracy of 81% for male and 69.23% for female. The findings of this research have potential applications in various fields, including mental health and emotion analysis in human communication.
AB - Stress is the physiological response to mental, emotional, or physical stress, which varies between individuals. A survey by Ipsos Global showed that around 30% of respondents identified stress as a significant health issue. Some countries in Southeast Asia, such as Cambodia, have much higher rates of depression than the world average. In Indonesia, the stress rate reached 9.8% in 2018. This research focuses on Speech Stress Recognition (SSR), an automated method that recognizes stress levels through speech characteristic analysis. We use Mel-Frequency Cepstral Coefficients (MFCC) feature extraction and the CNN-Transformer Encoder model. Evaluation results on the SUSAS dataset showed an overall accuracy of 73.76%. When the classification results are viewed by gender, male data appears better at classifying stress levels than female data. To improve performance, we implemented the Voice Activity Detection method, which resulted in an accuracy of 81% for male and 69.23% for female. The findings of this research have potential applications in various fields, including mental health and emotion analysis in human communication.
KW - CNN
KW - MFCC
KW - Speech Stress Recognition
KW - Stress
KW - Transformer Encoder
UR - https://www.scopus.com/pages/publications/85187212974
U2 - 10.1109/ICITCOM60176.2023.10442454
DO - 10.1109/ICITCOM60176.2023.10442454
M3 - Conference contribution
AN - SCOPUS:85187212974
T3 - Proceeding - International Conference on Information Technology and Computing 2023, ICITCOM 2023
SP - 147
EP - 151
BT - Proceeding - International Conference on Information Technology and Computing 2023, ICITCOM 2023
A2 - Chen, Hsing-Chung
A2 - Damarjati, Cahya
A2 - Blum, Christian
A2 - Jusman, Yessi
A2 - Kanafiah, Siti Nurul Aqmariah Mohd
A2 - Ejaz, Waleed
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2023 International Conference on Information Technology and Computing, ICITCOM 2023
Y2 - 1 December 2023 through 2 December 2023
ER -