TY - GEN
T1 - Temporal evolution of motion superpixel for video classification
AU - Yudistira, Novanto
AU - Kurita, Takio
N1 - Publisher Copyright:
© 2017 IEEE.
PY - 2017/7/19
Y1 - 2017/7/19
N2 - Superpixels are frequently used in pixel-level pre- processing for recognition purposes because they enable regional representation. Because dense motion contains more dynamic information for consecutive frames than pixels, it is beneficial for the task of video recognition. Correspondingly, we introduce motion superpixels in motion space, which are independently constructed per frame using superpixels extracted via energy driven sampling (SEEDS) and temporally tracked based on the nearest neighbourhood within consecutive flow fields. To define temporal information features, we track superpixels based on its central masses and descript the evolution of three motion based features namely histogram of flow (HOF), neighbourhood correlations, and centre of mass itself. Furthermore, we use the integrations of wavelet packet decomposition coefficients to achieve invariant scale and translation of sampled subsequences within a video sequence while also enrich features. To form sparseness of feature vectors, bag of features and generalized Histogram Kernel Support Vector Machine (HK-SVM) are used as learning algorithms. In our findings, firstly, various sizes of flow localisations are useful to improve recognition and secondly, spatio temporal dynamics of events can be covered by using wavelet packet either along space or time in which also improving accuracies. We demonstrate our results using the egocentric videos of JPL First-Person Interaction dataset and wild sport videos of UCF Sports dataset.
AB - Superpixels are frequently used in pixel-level pre- processing for recognition purposes because they enable regional representation. Because dense motion contains more dynamic information for consecutive frames than pixels, it is beneficial for the task of video recognition. Correspondingly, we introduce motion superpixels in motion space, which are independently constructed per frame using superpixels extracted via energy driven sampling (SEEDS) and temporally tracked based on the nearest neighbourhood within consecutive flow fields. To define temporal information features, we track superpixels based on its central masses and descript the evolution of three motion based features namely histogram of flow (HOF), neighbourhood correlations, and centre of mass itself. Furthermore, we use the integrations of wavelet packet decomposition coefficients to achieve invariant scale and translation of sampled subsequences within a video sequence while also enrich features. To form sparseness of feature vectors, bag of features and generalized Histogram Kernel Support Vector Machine (HK-SVM) are used as learning algorithms. In our findings, firstly, various sizes of flow localisations are useful to improve recognition and secondly, spatio temporal dynamics of events can be covered by using wavelet packet either along space or time in which also improving accuracies. We demonstrate our results using the egocentric videos of JPL First-Person Interaction dataset and wild sport videos of UCF Sports dataset.
UR - https://www.scopus.com/pages/publications/85027868725
U2 - 10.1109/CYBConf.2017.7985816
DO - 10.1109/CYBConf.2017.7985816
M3 - Conference contribution
AN - SCOPUS:85027868725
T3 - 2017 3rd IEEE International Conference on Cybernetics, CYBCONF 2017 - Proceedings
BT - 2017 3rd IEEE International Conference on Cybernetics, CYBCONF 2017 - Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 3rd IEEE International Conference on Cybernetics, CYBCONF 2017
Y2 - 21 June 2017 through 23 June 2017
ER -