本文へ移動

論文 ·対照・比較 ·未確認

An Ensemble Model of Word-based and Character-based Models for Japanese and Chinese Input Method

Yoh Okuno Shinsuke Mori

刊行年
2012-12-01
言語
英語
OpenAlex
W2251238909
MAG
2251238909
URL
https://openalex.org/W2251238909

要旨

Since Japanese and Chinese languages have too many characters to be input directly using a standard keyboard, input methods for these languages that enable users to input the characters are required. Recently, input methods based on statistical models have become popular because of their accuracy and ease of maintenance. Most of them adopt word-based models because they utilize word-segmented corpora to train the models. However, such word-based models suffer from unknown words because they cannot convert words correctly which are not in corpora. To handle this problem, we propose a character-based model that enables input methods to convert unknown words by exploiting character-aligned corpora automatically generated by a monotonic alignment tool. In addition to the character-based model, we propose an ensemble model of both character-based and word-based models to achieve higher accuracy. The ensemble model combines these two models by linear interpolation. All of these models are based on joint source channel model to utilize rich context through higher order joint n-gram. Experiments on Japanese and Chinese datasets showed that the character-based model performs reasonably and the ensemble model outperforms the word-based baseline model. As a future work, the effectiveness of incorporating large raw data should be investigated.

主題

この書誌の出所

  • openalex— W2251238909(2026-08-14取得)

引用

Yoh Okuno・Shinsuke Mori(2012-12-01) An Ensemble Model of Word-based and Character-based Models for Japanese and Chinese Input Method pp. 15-28

OkunoMori2012EnsembleModelWord
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON