TY - GEN
T1 - Real-Time American Sign Language Interpretation Using Deep Convolutional Neural Networks
AU - Biswasa, Arghya
AU - Sa, Gaurav
AU - Nanda, Umakanta
AU - Sharma, Diksha
AU - Sharma, Lakhan Dev
AU - Kuswiradyo, Primatar
N1 - Publisher Copyright:
© 2023, The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd.
PY - 2023
Y1 - 2023
N2 - In spite of being the 4th most commonly used language in the United States, sign language is actively used by only 10–14% of the members of the speech and hearing impairment community and continues to be one of the most understudied areas. In compliance with recent advances in deep learning, this paper explores the possibility of using deep convolutional neural network to interpret American Sign Language in real-time. The paper aims to develop a CNN from scratch and train it using a dataset patterned to closely match the format of the classic MNIST dataset (28 × 28 pixel images with pixel values ranging from 0 to 255). The evaluation of the network shows that it outperforms all previous implementations surrounding this task with 89.6% accuracy during real-time user testing (99.7% accuracy on validation set). Such high accuracy measure and fast converging time can be attributed to modelling the training data as a dataframe of pixel values rather than using traditional images and including batch normalization, concepts which have not been employed in earlier implementations to the best of our knowledge. To make the solution as accessible as possible, we refrained from using sophisticated hardware like motion-tracking gloves and depth-sensing cameras and deployed the trained model as a multi-platform mobile application.
AB - In spite of being the 4th most commonly used language in the United States, sign language is actively used by only 10–14% of the members of the speech and hearing impairment community and continues to be one of the most understudied areas. In compliance with recent advances in deep learning, this paper explores the possibility of using deep convolutional neural network to interpret American Sign Language in real-time. The paper aims to develop a CNN from scratch and train it using a dataset patterned to closely match the format of the classic MNIST dataset (28 × 28 pixel images with pixel values ranging from 0 to 255). The evaluation of the network shows that it outperforms all previous implementations surrounding this task with 89.6% accuracy during real-time user testing (99.7% accuracy on validation set). Such high accuracy measure and fast converging time can be attributed to modelling the training data as a dataframe of pixel values rather than using traditional images and including batch normalization, concepts which have not been employed in earlier implementations to the best of our knowledge. To make the solution as accessible as possible, we refrained from using sophisticated hardware like motion-tracking gloves and depth-sensing cameras and deployed the trained model as a multi-platform mobile application.
KW - American sign language
KW - Deep convolutional neural networks
KW - Deep learning
KW - Flutter development
KW - Hearing impairment
UR - https://www.scopus.com/pages/publications/85164954003
U2 - 10.1007/978-981-99-1203-2_18
DO - 10.1007/978-981-99-1203-2_18
M3 - Conference contribution
AN - SCOPUS:85164954003
SN - 9789819912025
T3 - Lecture Notes in Networks and Systems
SP - 209
EP - 220
BT - Advances in Distributed Computing and Machine Learning - Proceedings of ICADCML 2023
A2 - Chinara, Suchismita
A2 - Tripathy, Asis Kumar
A2 - Li, Kuan-Ching
A2 - Sahoo, Jyoti Prakash
A2 - Mishra, Alekha Kumar
PB - Springer Science and Business Media Deutschland GmbH
T2 - 4th International Conference on Advances in Distributed Computing and Machine Learning, ICADCML 2023
Y2 - 15 January 2023 through 16 January 2023
ER -