TY - GEN
T1 - BERT Fine-Tuning for Sentiment Analysis on Indonesian Mobile Apps Reviews
AU - Nugroho, Kuncahyo Setyo
AU - Sukmadewa, Anantha Yullian
AU - Wuswilahaken Dw, Haftittah
AU - Bachtiar, Fitra A.
AU - Yudistira, Novanto
N1 - Publisher Copyright:
© 2021 ACM.
PY - 2021/9/13
Y1 - 2021/9/13
N2 - User reviews have an essential role in the success of the developed mobile apps. User reviews in the textual form are unstructured data, creating a very high complexity when processed for sentiment analysis. Previous approaches that have been used often ignore the context of reviews. In addition, the relatively small data makes the model overfitting. A new approach, BERT, has been introduced as a transfer learning model with a pre-trained model that has previously been trained to have a better context representation. This study examines the effectiveness of fine-tuning BERT for sentiment analysis using two different pre-trained models. Besides the multilingual pre-trained model, we use the pre-trained model that only has been trained in Indonesian. The dataset used is Indonesian user reviews of the ten best apps in 2020 in Google Play sites. We also perform hyper-parameter tuning to find the optimum trained model. Two training data labeling approaches were also tested to determine the effectiveness of the model, which is score-based and lexicon-based. The experimental results show that pre-trained models trained in Indonesian have better average accuracy on lexicon-based data. The specific Indonesian pre-trained model achieved the highest accuracy of 84%, with 25 epochs and 24 minutes of training time.
AB - User reviews have an essential role in the success of the developed mobile apps. User reviews in the textual form are unstructured data, creating a very high complexity when processed for sentiment analysis. Previous approaches that have been used often ignore the context of reviews. In addition, the relatively small data makes the model overfitting. A new approach, BERT, has been introduced as a transfer learning model with a pre-trained model that has previously been trained to have a better context representation. This study examines the effectiveness of fine-tuning BERT for sentiment analysis using two different pre-trained models. Besides the multilingual pre-trained model, we use the pre-trained model that only has been trained in Indonesian. The dataset used is Indonesian user reviews of the ten best apps in 2020 in Google Play sites. We also perform hyper-parameter tuning to find the optimum trained model. Two training data labeling approaches were also tested to determine the effectiveness of the model, which is score-based and lexicon-based. The experimental results show that pre-trained models trained in Indonesian have better average accuracy on lexicon-based data. The specific Indonesian pre-trained model achieved the highest accuracy of 84%, with 25 epochs and 24 minutes of training time.
KW - Apps review
KW - BERT fine-tuning
KW - Sentiment analysis
UR - https://www.scopus.com/pages/publications/85118877812
U2 - 10.1145/3479645.3479679
DO - 10.1145/3479645.3479679
M3 - Conference contribution
AN - SCOPUS:85118877812
T3 - ACM International Conference Proceeding Series
SP - 258
EP - 264
BT - Proceedings of 2021 International Conference on Sustainable Information Engineering and Technology, SIET 2021
PB - Association for Computing Machinery
T2 - 6th International Conference on Sustainable Information Engineering and Technology, SIET 2021
Y2 - 13 September 2021 through 14 September 2021
ER -