論文 ·用例に日本語 ·未確認
Smoothing a lexicon-based POS tagger for Arabic and Hebrew
Saib Mansour ・ Khalil Sima’an ・ Yoad Winter
- 刊行年
- 2007-01-01
- 言語
- 英語
- OpenAlex
- W2033830177
- DOI
- 10.3115/1654576.1654593
- MAG
- 2033830177
- URL
- https://dl.acm.org/doi/pdf/10.5555/1654576.1654593
要旨
We propose an enhanced Part-of-Speech (POS) tagger of Semitic languages that treats Modern Standard Arabic (henceforth Arabic) and Modern Hebrew (henceforth Hebrew) using the same probabilistic model and architectural setting. We start out by porting an existing Hidden Markov Model POS tagger for Hebrew to Arabic by exchanging a morphological analyzer for Hebrew with Buckwalter's (2002) morphological analyzer for Arabic. This gives state-of-the-art accuracy (96.12%), comparable to Habash and Rambow's (2005) analyzer-based POS tagger on the same Arabic datasets. However, further improvement of such analyzer-based tagging methods is hindered by the incomplete coverage of standard morphological analyzer (Bar Haim et al., 2005). To overcome this coverage problem we supplement the output of Buckwalter's analyzer with synthetically constructed analyses that are proposed by a model which uses character information (Diab et al., 2004) in a way that is similar to Nakagawa's (2004) system for Chinese and Japanese. A version of this extended model that (unlike Nakagawa) incorporates synthetically constructed analyses also for known words achieves 96.28% accuracy on the standard Arabic test set.
主題
この書誌の出所
- openalex— W2033830177(2026-08-14取得)
引用
Saib Mansour・Khalil Sima’an・Yoad Winter(2007-01-01) Smoothing a lexicon-based POS tagger for Arabic and Hebrew pp. 97-97