本文へ移動

論文 ·用例に日本語 ·未確認

Models of Visually Grounded Speech Signal Pay Attention to Nouns: A Bilingual Experiment on English and Japanese

William N. Havard Jean‐Pierre Chevrot Laurent Besacier

刊行年
2019-04-17
言語
英語
OpenAlex
W2920166246
DOI
10.1109/icassp.2019.8683069
MAG
2920166246
URL
https://arxiv.org/pdf/1902.03052

要旨

We investigate the behaviour of attention in neural models of visually grounded speech trained on two languages: English and Japanese. Experimental results show that attention focuses on nouns and this behaviour holds true for two very typologically different languages. We also draw parallels between artificial neural attention and human attention and show that neural attention focuses on word endings as it has been theorised for human attention. Finally, we investigate how two visually grounded monolingual models can be used to perform cross-lingual speech-to-speech retrieval. For both languages, the enriched bilingual (speech-image) corpora with part-of-speech tags and forced alignments are distributed to the community for reproducible research.

主題

この書誌の出所

  • openalex— W2920166246(2026-08-14取得)

引用

William N. Havard・Jean‐Pierre Chevrot・Laurent Besacier(2019-04-17) Models of Visually Grounded Speech Signal Pay Attention to Nouns: A Bilingual Experiment on English and Japanese pp. 8618-8622

Havard2019ModelsVisuallyGrounded
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON