Skip to main navigation Skip to search Skip to main content

Classification of Protein Structure with A Variety of Code Features on the Same Index

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

There have been several studies that use protein structure data using various methods, the highest accuracy of some of these studies is a maximum of 80%. In general, there are three ways to change the character code data of a protein structure to be numeric, that is by calculating the number of each code, using the PAM matrix and by using the concept of probability for each code in the same index. In this study, the data and methods are the same as those done by Tawang with a maximum accuracy of 79.17% [1], namely the classification of protein structure data using the Naïve Bayes Classifier method. The difference with this research is the features used. For research conducted by Tawang, the feature used is the length of one protein structure (393 characters). In this study, the features are taken from the diversity of codes in the same index because the results of the observation of many protein structures have the same code in the same index and the average number of diverse features as many as 250 indexes. This study produces the best accuracy (more than 90%) using test data smaller or equal to 50. With training data greater than 200 of 753 data will produce accuracy that tends to be low (less than 60%), then changes in value prior probability which have the smallest value among other priors. The test results show that changes in the prior value can increase the accuracy by up to 5%.

Original languageEnglish
Title of host publication3rd International Conference on Sustainable Information Engineering and Technology, SIET 2018 - Proceedings
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages326-331
Number of pages6
ISBN (Electronic)9781538674079
DOIs
Publication statusPublished - 2 Jul 2018
Event3rd International Conference on Sustainable Information Engineering and Technology, SIET 2018 - Malang, Indonesia
Duration: 10 Nov 201812 Nov 2018

Publication series

Name3rd International Conference on Sustainable Information Engineering and Technology, SIET 2018 - Proceedings

Conference

Conference3rd International Conference on Sustainable Information Engineering and Technology, SIET 2018
Country/TerritoryIndonesia
CityMalang
Period10/11/1812/11/18

Keywords

  • Naïve Bayes
  • P53
  • Prior
  • Protein Structure

Fingerprint

Dive into the research topics of 'Classification of Protein Structure with A Variety of Code Features on the Same Index'. Together they form a unique fingerprint.

Cite this