本文へ移動

論文 ·用例に日本語 ·未確認

Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms

Yasunori Ohishi Akisato Kimura Takahito Kawanishi Kunio Kashino David Harwath James Glass

刊行年
2020-04-09
言語
英語
OpenAlex
W3015300171
DOI
10.1109/icassp40776.2020.9053428
MAG
3015300171
URL
https://doi.org/10.1109/icassp40776.2020.9053428

要旨

We propose a trilingual semantic embedding model that associates visual objects in images with segments of speech signals corresponding to spoken words in an unsupervised manner. Unlike the existing models, our model incorporates three different languages, namely, English, Hindi, and Japanese. To build the model, we used the existing English and Hindi datasets and collected a new corpus of Japanese speech captions. These spoken captions are spontaneous descriptions by individual speakers, rather than readings based on prepared transcripts. Therefore, we introduce a self-attention mechanism into the model to better map the spoken captions associated with the same image into the embedding space. We hope that the self-attention mechanism efficiently captures relationships between widely separated word-like segments. Experimental results show that the introduction of a third language improves the average performance in terms of cross-modal and cross-lingual retrieval accuracy, and that the self-attention mechanism added to the model works effectively.

主題

この書誌の出所

  • openalex— W3015300171(2026-08-14取得)

引用

Yasunori Ohishi・Akisato Kimura・Takahito Kawanishi・Kunio Kashino・David Harwath・James Glass(2020-04-09) Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms pp. 4352-4356

Ohishi2020TrilingualSemanticEmbeddings
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON