本文へ移動

論文 ·日本語 ·未確認

A stochastic approach to phoneme and accent estimation

Tohru Nagano Shinsuke Mori Masafumi Nishimura

刊行年
2005-09-04
言語
英語
OpenAlex
W35654218
DOI
10.21437/interspeech.2005-575
MAG
35654218
URL
https://doi.org/10.21437/interspeech.2005-575

要旨

We present a new stochastic approach to estimate accurately phonemes and accents for Japanese TTS (Text-to-Speech) systems. Front-end process of TTS system assigns phonemes and accents to an input plain text, which is critical for creating intelligible and natural speech. Rule-based approaches that build hierarchical structures are widely used for this purpose. However, considering scalability and the ease of domain adaptation, rule-based approaches have well-known limitations. In this paper, we present a stochastic method based on an n-gram model for phonemes and accents estimation. The proposed method estimates not only phonemes and accents but word segmentation and part-of-speech (POS) simultaneously. We implemented a system for Japanese which solves tokenization, linguistic annotation, text-to-phonemes conversion, homograph disambiguation, and accents generation at the same time, and observed promising results.

主題

この書誌の出所

  • openalex— W35654218(2026-08-14取得)

引用

Tohru Nagano・Shinsuke Mori・Masafumi Nishimura(2005-09-04) A stochastic approach to phoneme and accent estimation 2005(69) pp. 3293-3296

Nagano2005StochasticApproachPhoneme
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON