本文へ移動

論文 ·日本語 ·未確認

Unsupervised Lexicon Acquisition from Speech and Text

Gakuto Kurata Shinsuke Mori Nobuyasu Itoh Masafumi Nishimura

刊行年
2007-04-01
言語
英語
OpenAlex
W2017380725
DOI
10.1109/icassp.2007.366939
MAG
2017380725
URL
https://doi.org/10.1109/icassp.2007.366939

要旨

When introducing a large vocabulary continuous speech recognition (LVCSR) system into a specific domain, it is preferable to add the necessary domain-specific words and their correct pronunciations selectively to the lexicon, especially in the areas where the LVCSR system should be updated frequently by adding new words. In this paper, we propose an unsupervised method of word acquisition in Japanese, where no spaces exist between words. In our method, by taking advantage of the speech of the target domain, we selected the domain-specific words among an enormous number of word candidates extracted from the raw corpora. The experiments showed that the acquired lexicon was of good quality and that it contributed to the performance of the LVCSR system for the target domain.

主題

この書誌の出所

  • openalex— W2017380725(2026-08-14取得)

引用

Gakuto Kurata・Shinsuke Mori・Nobuyasu Itoh・Masafumi Nishimura(2007-04-01) Unsupervised Lexicon Acquisition from Speech and Text pp. IV-421

Kurata2007UnsupervisedLexiconAcquisition
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON