本文へ移動

論文 ·日本語 ·未確認

Acoustic Model Adaptation for Emotional Speech Recognition Using Twitter-Based Emotional Speech Corpus

Tetsuo Kosaka Yoshitaka Aizawa Masaharu Kato Takashi Nose

刊行年
2018-11-01
言語
英語
OpenAlex
W2921306346
DOI
10.23919/apsipa.2018.8659756
MAG
2921306346
URL
https://doi.org/10.23919/apsipa.2018.8659756

要旨

In recent years, Japanese Twitter-based emotional speech (JTES) was constructed as an emotional speech corpus. This corpus is based on tweets, and has features wherein an emotional label is assigned to each sentence, and sentences are selected considering the balance of both phoneme and prosody. Compared to speech recognition without emotion, emotional speech recognition is a difficult task. In this study, we aim to improve the performance of emotional speech recognition on the JTES corpus using acoustic model adaptation. For recognition, a deep neural network-based hidden Markov model (DNN-HMM) is used as the acoustic model. As a baseline, a word error rate (WER) of 38.0% was obtained when the DNN-HMM was trained by the corpus of spontaneous Japanese. This model was used as an initial model for adaptation. In this study, various types of adaptation were examined, and substantial performance improvement was achieved. Finally, a WER of 23.05% was obtained using speaker adaptation.

主題

この書誌の出所

  • openalex— W2921306346(2026-08-14取得)

引用

Tetsuo Kosaka・Yoshitaka Aizawa・Masaharu Kato・Takashi Nose(2018-11-01) Acoustic Model Adaptation for Emotional Speech Recognition Using Twitter-Based Emotional Speech Corpus pp. 1747-1751

Kosaka2018AcousticModelAdaptation
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON