論文 ·日本語 ·未確認
Extended models and tools for high-performance part-of-speech tagger
Masayuki Asahara ・ Yūji Matsumoto
- 刊行年
- 2000-01-01
- 言語
- 英語
- OpenAlex
- W2000665223
- DOI
- 10.3115/990820.990824
- MAG
- 2000665223
- URL
- https://dl.acm.org/doi/pdf/10.3115/990820.990824
要旨
Statistical part-of-speech (POS) taggers achieve high accuracy and robustness when based on large scale manually tagged corpora. However, enhancements of the learning models are necessary to achieve better performance. We are developing a learning tool for a Japanese morphological analyzer called ChaSen. Currently we use a fine-grained POS tag set with about 500 tags. To apply a normal tri gram model on the tag set, we need unrealistic size of corpora. Even, for a bi-gram model, we cannot prepare a moderate size of an annotated corpus, when we take all the tags as distinct. A usual technique to cope with such fine-grained tags is to reduce the size of the tag set by grouping the set of tags into equivalence classes. We introduce the concept of position-wise grouping where the tag set is partitioned into different equivalence classes at each position in the conditional probabilities in the Markov Model. Moreover, to cope with the data sparseness problem caused by exceptional phenomena, we introduce several other techniques such as word-level statistics, smoothing of word-level and POS-level statistics and a selective tri-gram model. To help users determine probabilistic parameters, we introduce an error-driven method for the parameter selection. We then give results of experiments to see the effect of the tools applied to an existing Japanese morphological analyzer.
主題
この書誌の出所
- openalex— W2000665223(2026-08-14取得)
引用
Masayuki Asahara・Yūji Matsumoto(2000-01-01) Extended models and tools for high-performance part-of-speech tagger 1 pp. 21-27