TY - GEN
T1 - Speaker Recognition on Low Power Device Using Fully Convolutional QuartzNet
AU - Laksono, Blessius Sheldo Putra
AU - Prasetio, Barlian Henryranu
N1 - Publisher Copyright:
© 2023 ACM.
PY - 2023/10/24
Y1 - 2023/10/24
N2 - The need for a small and lightweight algorithm used for speaker recognition that can run on low-power devices is on the rise. This is mainly caused by security and privacy concerns of users with the use of their personal and biometric data. The speaker recognition task is mainly used as a biometric authentication, so an accurate model is also needed. The previous method uses feature engineering to extract features from raw audio files with heavy reliance on the training data and a dissimilarity between the training data and real-world implementation causes a significant decrease in its accuracy. We propose a Fully Convolutional QuartzNet as a deep learning approach to this problem. We achieved 84.6% accuracy when testing on a small subset DR-VCTK dataset with 30 classes and 56.40% accuracy on a small subset of the VoxCeleb dataset with fewer files for each of the 125 classes. The proposed model was also tested for binary speaker recognition, achieving 5.07% EER. We also achieve a small parameter count of only 33K parameters without sacrificing significant performance, and the proposed method can achieve its highest accuracy with only 53K parameters.
AB - The need for a small and lightweight algorithm used for speaker recognition that can run on low-power devices is on the rise. This is mainly caused by security and privacy concerns of users with the use of their personal and biometric data. The speaker recognition task is mainly used as a biometric authentication, so an accurate model is also needed. The previous method uses feature engineering to extract features from raw audio files with heavy reliance on the training data and a dissimilarity between the training data and real-world implementation causes a significant decrease in its accuracy. We propose a Fully Convolutional QuartzNet as a deep learning approach to this problem. We achieved 84.6% accuracy when testing on a small subset DR-VCTK dataset with 30 classes and 56.40% accuracy on a small subset of the VoxCeleb dataset with fewer files for each of the 125 classes. The proposed model was also tested for binary speaker recognition, achieving 5.07% EER. We also achieve a small parameter count of only 33K parameters without sacrificing significant performance, and the proposed method can achieve its highest accuracy with only 53K parameters.
KW - Artificial Intelligence on The Edge
KW - Edge Computing
KW - Fully Convolutional Network
KW - Speaker Recognition
KW - Time Channel Separable Convolution
UR - https://www.scopus.com/pages/publications/85182393695
U2 - 10.1145/3626641.3626946
DO - 10.1145/3626641.3626946
M3 - Conference contribution
AN - SCOPUS:85182393695
T3 - ACM International Conference Proceeding Series
SP - 619
EP - 624
BT - SIET 2023 - Proceedings of the 8th International Conference on Sustainable Information Engineering and Technology
PB - Association for Computing Machinery
T2 - 8th International Conference on Sustainable Information Engineering and Technology, SIET 2023
Y2 - 24 October 2023 through 25 October 2023
ER -