論文 ·日本語 ·未確認
Prosody-aware subword embedding considering Japanese intonation systems and its application to DNN-based multi-dialect speech synthesis
Takanori Akiyama ・ Shinnosuke Takamichi ・ Hiroshi Saruwatari
- 刊行年
- 2018-11-01
- 言語
- 英語
- OpenAlex
- W2921047845
- DOI
- 10.23919/apsipa.2018.8659465
- MAG
- 2921047845
- URL
- https://doi.org/10.23919/apsipa.2018.8659465
要旨
This paper presents prosody-aware subword embedding considering Japanese intonation systems and its application to DNN (deep neural network)-based multi-dialect speech synthesis. In accordance with recent improvements of speech synthesis in rich-resourced languages, the research trend is shifting to more challenging languages such as Japanese dialects that still have undefined prosodic contexts. Conventional prosody-aware word embedding can unsupervisedly extract the contexts in a data-driven manner using words and F0 sequences. However, accurate contexts for unknown words are difficult to generate. To solve this problem, we propose prosody-aware subword embedding considering Japanese intonation systems. The unsupervised subword model, which is trained considering language and acoustic characteristics, can tokenize an unknown word into known subwords suitable for prosody-aware embedding. We also propose a modulation filtering method considering intra-subword moras to improve the embedding accuracies. We apply the methods to not only Japanese but also Japanese multi-dialect speech synthesis. In the multi-dialect case, we propose subword models shared among dialects and embedding models conditioned by dialect information. The experimental evaluation demonstrates that the proposed multi-dialect methods can improve speech quality in some Japanese dialects.
主題
この書誌の出所
- openalex— W2921047845(2026-08-14取得)
引用
Takanori Akiyama・Shinnosuke Takamichi・Hiroshi Saruwatari(2018-11-01) Prosody-aware subword embedding considering Japanese intonation systems and its application to DNN-based multi-dialect speech synthesis pp. 659-664