論文 ·日本語 ·未確認
Designing text corpus using phone-error distribution for acoustic modeling
Hiroko Murakami ・ Koichi Shinoda ・ Sadaoki Furui
- 刊行年
- 2011-12-01
- 言語
- 英語
- OpenAlex
- W2008737763
- DOI
- 10.1109/asru.2011.6163929
- MAG
- 2008737763
- URL
- https://doi.org/10.1109/asru.2011.6163929
要旨
It is expensive to prepare a sufficient amount of training data for acoustic modeling for developing large vocabulary continuous speech recognition systems. This is a serious problem especially for resource-deficient languages. We propose an active learning method that effectively reduces the amount of training data without any degradation in recognition performance. It is used to design a text corpus for read speech collection. It first estimates phone-error distribution using a small amount of fully transcribed speech data. Second, it constructs a sentence set whose phone-occurrence distribution is close to the phone-error distribution and collects its speech data. It then extends this process to diphones and triphones and collects more speech data. We evaluated our method with simulation experiments using the Corpus of Spontaneous Japanese. It required only 76 h of speech data to achieve word accuracy of 74.7%, while the conventional training method required 152 h of data to achieve the same rate.
主題
この書誌の出所
- openalex— W2008737763(2026-08-14取得)
引用
Hiroko Murakami・Koichi Shinoda・Sadaoki Furui(2011-12-01) Designing text corpus using phone-error distribution for acoustic modeling 1 pp. 191-195