本文へ移動

論文 ·日本語 ·未確認

Simultaneous Adaptation of Acoustic and Language Models for Emotional Speech Recognition Using Tweet Data

Tetsuo Kosaka Kazuya Saeki Yoshitaka Aizawa Kato Masaharu Takashi Nose

刊行年
2024-02-29
収録
『IEICE Transactions on Information and Systems』 E107.D(3) pp. 363-373
出版
Institute of Electronics, Information and Communication Engineers
言語
英語
OpenAlex
W4392281383
DOI
10.1587/transinf.2023hcp0010
ISSN
0916-8532
URL
https://www.jstage.jst.go.jp/article/transinf/E107.D/3/E107.D_2023HCP0010/_pdf

要旨

Emotional speech recognition is generally considered more difficult than non-emotional speech recognition. The acoustic characteristics of emotional speech differ from those of non-emotional speech. Additionally, acoustic characteristics vary significantly depending on the type and intensity of emotions. Regarding linguistic features, emotional and colloquial expressions are also observed in their utterances. To solve these problems, we aim to improve recognition performance by adapting acoustic and language models to emotional speech. We used Japanese Twitter-based Emotional Speech (JTES) as an emotional speech corpus. This corpus consisted of tweets and had an emotional label assigned to each utterance. Corpus adaptation is possible using the utterances contained in this corpus. However, regarding the language model, the amount of adaptation data is insufficient. To solve this problem, we propose an adaptation of the language model by using online tweet data downloaded from the internet. The sentences used for adaptation were extracted from the tweet data based on certain rules. We extracted the data of 25.86 M words and used them for adaptation. In the recognition experiments, the baseline word error rate was 36.11%, whereas that with the acoustic and language model adaptation was 17.77%. The results demonstrated the effectiveness of the proposed method.

主題

この書誌の出所

  • openalex— W4392281383(2026-08-14取得)

引用

Tetsuo Kosaka・Kazuya Saeki・Yoshitaka Aizawa・Kato Masaharu・Takashi Nose(2024-02-29) Simultaneous Adaptation of Acoustic and Language Models for Emotional Speech Recognition Using Tweet Data 『IEICE Transactions on Information and Systems』 E107.D(3) pp. 363-373 Institute of Electronics, Information and Communication Engineers

Kosaka2024SimultaneousAdaptationAcoustic
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON