TY - GEN
T1 - Mutual Information for Learning Context Representation on RNN-Attention Based Models in Open Domain Generative Chatbot
AU - Fauzulhaq, Alfirsa Damasyifa
AU - Bachtiar, Fitra Abdurrachman
N1 - Publisher Copyright:
© 2023 ACM.
PY - 2023/10/24
Y1 - 2023/10/24
N2 - Chatbot is an example of the application of Artificial Intelligence that can receive and answer questions automatically. Chatbots are widely used in various fields such as health, customer service, entertainment, education and others. There are two approaches to chatbot development, rule-based and generative. Rule-based chatbot has the advantage of being easy to develop and produces good answers but requires predefined rules that are defined manually. Generative chatbot can provide dynamic and natural answers and does not require predefined rules. However, the drawback of generative chatbot lies in the weak representation of sentence information and information bottleneck which results in loss of information or context. The main objective of this research is to get the best model for open domain generative chatbot in a predefined scenario and improve the performance of the model in terms of word information representation using SBERT Pretrained Word Embedding and reduce information loss in encoder bottleneck and output using Mutual Information. Based on the experimental results, LSTM with the addition of Bahdanau Attention achieved the best performance in all scenarios with the highest BLEU and BERT F1-Score. Whereas in the 50 and 100 (long) sequence scenarios, the addition of Mutual Information and SBERT can improve overall model performance for BLEU by 3.62% and 2.58% respectively and BERT Score by 3.16% and 5.10% respectively.
AB - Chatbot is an example of the application of Artificial Intelligence that can receive and answer questions automatically. Chatbots are widely used in various fields such as health, customer service, entertainment, education and others. There are two approaches to chatbot development, rule-based and generative. Rule-based chatbot has the advantage of being easy to develop and produces good answers but requires predefined rules that are defined manually. Generative chatbot can provide dynamic and natural answers and does not require predefined rules. However, the drawback of generative chatbot lies in the weak representation of sentence information and information bottleneck which results in loss of information or context. The main objective of this research is to get the best model for open domain generative chatbot in a predefined scenario and improve the performance of the model in terms of word information representation using SBERT Pretrained Word Embedding and reduce information loss in encoder bottleneck and output using Mutual Information. Based on the experimental results, LSTM with the addition of Bahdanau Attention achieved the best performance in all scenarios with the highest BLEU and BERT F1-Score. Whereas in the 50 and 100 (long) sequence scenarios, the addition of Mutual Information and SBERT can improve overall model performance for BLEU by 3.62% and 2.58% respectively and BERT Score by 3.16% and 5.10% respectively.
KW - Attention Mechanism
KW - Chatbot
KW - Deep Neural Network
KW - Natural Language Processing
KW - Sequence-to-sequence
UR - https://www.scopus.com/pages/publications/85182397697
U2 - 10.1145/3626641.3626926
DO - 10.1145/3626641.3626926
M3 - Conference contribution
AN - SCOPUS:85182397697
T3 - ACM International Conference Proceeding Series
SP - 112
EP - 118
BT - SIET 2023 - Proceedings of the 8th International Conference on Sustainable Information Engineering and Technology
PB - Association for Computing Machinery
T2 - 8th International Conference on Sustainable Information Engineering and Technology, SIET 2023
Y2 - 24 October 2023 through 25 October 2023
ER -