論文 ·対照・比較 ·未確認
An Ensemble Model of Word-based and Character-based Models for Japanese and Chinese Input Method
- 刊行年
- 2012-12-01
- 言語
- 英語
- OpenAlex
- W2251238909
- MAG
- 2251238909
- URL
- https://openalex.org/W2251238909
要旨
Since Japanese and Chinese languages have too many characters to be input directly using a standard keyboard, input methods for these languages that enable users to input the characters are required. Recently, input methods based on statistical models have become popular because of their accuracy and ease of maintenance. Most of them adopt word-based models because they utilize word-segmented corpora to train the models. However, such word-based models suffer from unknown words because they cannot convert words correctly which are not in corpora. To handle this problem, we propose a character-based model that enables input methods to convert unknown words by exploiting character-aligned corpora automatically generated by a monotonic alignment tool. In addition to the character-based model, we propose an ensemble model of both character-based and word-based models to achieve higher accuracy. The ensemble model combines these two models by linear interpolation. All of these models are based on joint source channel model to utilize rich context through higher order joint n-gram. Experiments on Japanese and Chinese datasets showed that the character-based model performs reasonably and the ensemble model outperforms the word-based baseline model. As a future work, the effectiveness of incorporating large raw data should be investigated.
主題
この書誌の出所
- openalex— W2251238909(2026-08-14取得)
引用
Yoh Okuno・Shinsuke Mori(2012-12-01) An Ensemble Model of Word-based and Character-based Models for Japanese and Chinese Input Method pp. 15-28