論文 ·日本語 ·未確認
A stochastic approach to phoneme and accent estimation
Tohru Nagano ・ Shinsuke Mori ・ Masafumi Nishimura
- 刊行年
- 2005-09-04
- 言語
- 英語
- OpenAlex
- W35654218
- DOI
- 10.21437/interspeech.2005-575
- MAG
- 35654218
- URL
- https://doi.org/10.21437/interspeech.2005-575
要旨
We present a new stochastic approach to estimate accurately phonemes and accents for Japanese TTS (Text-to-Speech) systems. Front-end process of TTS system assigns phonemes and accents to an input plain text, which is critical for creating intelligible and natural speech. Rule-based approaches that build hierarchical structures are widely used for this purpose. However, considering scalability and the ease of domain adaptation, rule-based approaches have well-known limitations. In this paper, we present a stochastic method based on an n-gram model for phonemes and accents estimation. The proposed method estimates not only phonemes and accents but word segmentation and part-of-speech (POS) simultaneously. We implemented a system for Japanese which solves tokenization, linguistic annotation, text-to-phonemes conversion, homograph disambiguation, and accents generation at the same time, and observed promising results.
主題
この書誌の出所
- openalex— W35654218(2026-08-14取得)
引用
Tohru Nagano・Shinsuke Mori・Masafumi Nishimura(2005-09-04) A stochastic approach to phoneme and accent estimation 2005(69) pp. 3293-3296