論文 ·日本語 ·未確認
Handling of Out-of-vocabulary Words in Japanese-English Machine Translation by Exploiting Parallel Corpus.
- 刊行年
- 2015-01-01
- 言語
- 英語
- OpenAlex
- W2401186096
- MAG
- 2401186096
- URL
- https://openalex.org/W2401186096
要旨
A large number of loanwords and orthographic variants in Japanese pose a challenge for machine translation. In this article, we present a hybrid model for handling out-of-vocabulary words in Japanese-to-English statistical machine translation output by exploiting parallel corpus. As the Japanese writing system makes use of four different script sets (kanji, hiragana, katakana, and romaji), we treat these scripts differently. A machine transliteration model is built to transliterate out-of-vocabulary Japanese katakana words into English words. A Japanese dependency structure analyzer is employed to tackle out-of-vocabulary kanji and hiragana words. The evaluation results demonstrate that it is an effective approach for addressing out-of-vocabulary word problems and decreasing the OOVs rate in the Japanese-to-English machine translation tasks.
主題
この書誌の出所
- openalex— W2401186096(2026-08-14取得)
引用
Juan Luo・Yves Lepage(2015-01-01) Handling of Out-of-vocabulary Words in Japanese-English Machine Translation by Exploiting Parallel Corpus. 23 pp. 1-20