TY - GEN
T1 - Knowledge-Enriched Domain Specific Chatbot on Low-resource Language
AU - Perdana, Rizal Setya
AU - Adikara, Putra Pandu
AU - Indriati,
AU - Kurnianingtyas, Diva
N1 - Publisher Copyright:
© 2022 IEEE.
PY - 2022
Y1 - 2022
N2 - This study presents architecture in learning human-machine conversation using a human language known as chatbots. Given an utterance, a chatbot attempt to reply using human language by mimicking human intelligence in communication. Previous works in this domain were mostly implemented in English, thus it remains a problem when bringing to a low-resource language, e.g., Indonesian. The limited availability of human-to-human dialog history led to difficulty in the learning process of machine learning algorithms. Therefore, this study proposed a Knowledgeable Chatbot (KC), an enhanced pipeline that enables to receiving of transferred knowledge from another task. A data augmentation pipeline is proposed to handle the limited number of available. To deal with the low-resource language, this study proposed to incorporate a pre-trained language model to gain contextualized language understanding. As this research can be categorized as preliminary, extensive experiments are required to prove the effectiveness of each part. Standard automatic metrics for information retrieval and classification prove that KC excels in the ablation study.
AB - This study presents architecture in learning human-machine conversation using a human language known as chatbots. Given an utterance, a chatbot attempt to reply using human language by mimicking human intelligence in communication. Previous works in this domain were mostly implemented in English, thus it remains a problem when bringing to a low-resource language, e.g., Indonesian. The limited availability of human-to-human dialog history led to difficulty in the learning process of machine learning algorithms. Therefore, this study proposed a Knowledgeable Chatbot (KC), an enhanced pipeline that enables to receiving of transferred knowledge from another task. A data augmentation pipeline is proposed to handle the limited number of available. To deal with the low-resource language, this study proposed to incorporate a pre-trained language model to gain contextualized language understanding. As this research can be categorized as preliminary, extensive experiments are required to prove the effectiveness of each part. Standard automatic metrics for information retrieval and classification prove that KC excels in the ablation study.
KW - chatbots
KW - deep learning
KW - information retrieval
KW - language model
KW - natural language processing
KW - virtual assistants
UR - https://www.scopus.com/pages/publications/85140575848
U2 - 10.1109/EECCIS54468.2022.9902930
DO - 10.1109/EECCIS54468.2022.9902930
M3 - Conference contribution
AN - SCOPUS:85140575848
T3 - Proceedings - 11th Electrical Power, Electronics, Communications, Control, and Informatics Seminar, EECCIS 2022
SP - 310
EP - 315
BT - Proceedings - 11th Electrical Power, Electronics, Communications, Control, and Informatics Seminar, EECCIS 2022
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 11th Electrical Power, Electronics, Communications, Control, and Informatics Seminar, EECCIS 2022
Y2 - 23 August 2022 through 25 August 2022
ER -