Skip to main navigation Skip to search Skip to main content

Vision-text time series correlation for visual-to-language story generation

Research output: Contribution to journalArticlepeer-review

Abstract

Automatic generation of textual stories from visual data representation, known as visual storytelling, is a recent advancement in the problem of images-to-text. Instead of using a single image as input, visual storytelling processes a sequential array of images into coherent sentences. A story contains non-visual concepts as well as descriptions of literal object(s). While previous approaches have applied external knowledge, our approach was to regard the non-visual concept as the semantic correlation between visual modality and textual modality. This paper, therefore, presents new features representation based on a canonical correlation analysis between two modalities. Attention mechanism are adopted as the underlying architecture of the image-to-text problem, rather than standard encoder-decoder models. Canonical Correlation Attention Mechanism (CAAM), the proposed end-to-end architecture, extracts time series correlation by maximizing the cross-modal correlation. Extensive experiments on VIST dataset (http://visionandlanguage.net/VIST/dataset.html) were conducted to demonstrate the effectiveness of the architecture in terms of automatic metrics, with additional experiments show the impact of modality fusion strategy.

Original languageEnglish
Pages (from-to)828-839
Number of pages12
JournalIEICE Transactions on Information and Systems
VolumeE104D
Issue number6
DOIs
Publication statusPublished - 2021
Externally publishedYes

Keywords

  • Attention mechanism
  • Correlation analysis
  • Visual storytelling

Fingerprint

Dive into the research topics of 'Vision-text time series correlation for visual-to-language story generation'. Together they form a unique fingerprint.

Cite this