本文へ移動

論文 ·日本語 ·未確認

Benchmark test for speech recognition using the Corpus of Spontaneous Japanese

Tatsuya Kawahara

刊行年
2003-01-01
出版
Tokyo Institute of Technology
言語
英語
OpenAlex
W87954838
MAG
87954838
URL
http://t2r2.star.titech.ac.jp/cgi-bin/publicationinfo.cgi?q_publication_content_number=CTT100485181

要旨

We present benchmark results of automatic speech recognition using the Corpus of Spontaneous Japanese (CSJ), which has been developed in the five-year national project and will be the largest spontaneous speech databases.New test-sets are designed for both academic presentation speech and extemporaneous public speech, which are the two major categories in the corpus.The testsets are selected to cover the variation of acoustic and linguistic factors in spontaneous speech: word perplexity, degree of disfluency, and the speaking rate.Baseline acoustic and language models are set up using an almost complete set (500 hours and 6.67M words) of the CSJ.Statistical modeling of pronunciation variation is also incorporated into the language model based on the alignment of large-scale transcriptions.The benchmark results verified the effects of the factors considered in the test-set design.

主題

この書誌の出所

  • openalex— W87954838(2026-08-14取得)

引用

Tatsuya Kawahara(2003-01-01) Benchmark test for speech recognition using the Corpus of Spontaneous Japanese pp. 135-138 Tokyo Institute of Technology

Kawahara2003BenchmarkTestSpeech
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON