本文へ移動

論文 ·日本語 ·未確認

A pointwise approach to pronunciation estimation for a TTS front-end

Shinsuke Mori Graham Neubig

刊行年
2011-08-27
言語
英語
OpenAlex
W2401761779
DOI
10.21437/interspeech.2011-571
MAG
2401761779
URL
https://doi.org/10.21437/interspeech.2011-571

要旨

In this paper, we propose a pointwise approach to the Japanese TTS front-end. In this approach, phoneme sequence estimation of sentences is decomposed into two tasks: word segmentation of the input sentence and phoneme estimation of each word. Then these two tasks are solved by pointwise classifiers without referring to the neighboring classification results. In contrast to existing sequence-based methods, an n-gram model based on sequences of word-phoneme pairs for example, this framework enables us to use various language resources such as sentences in which only a few words are annotated, or an unsegmented list of compound words, among others. In the experiments, we compared a joint tri-gram model with the combination of a pointwise word segmenter and a pointwise phoneme sequence estimator. The results showed that our framework successfully enables a TTS front-end to refer to a partially annotated corpus and/or a word sequence list annotated with phoneme sequences to realize a far larger improvement in accuracy.

主題

この書誌の出所

  • openalex— W2401761779(2026-08-14取得)

引用

Shinsuke Mori・Graham Neubig(2011-08-27) A pointwise approach to pronunciation estimation for a TTS front-end pp. 2181-2184

MoriNeubig2011PointwiseApproachPronunciation
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON