本文へ移動

論文 ·日本語 ·未確認

Handling of Out-of-vocabulary Words in Japanese-English Machine Translation by Exploiting Parallel Corpus.

Juan Luo Yves Lepage

刊行年
2015-01-01
言語
英語
OpenAlex
W2401186096
MAG
2401186096
URL
https://openalex.org/W2401186096

要旨

A large number of loanwords and orthographic variants in Japanese pose a challenge for machine translation. In this article, we present a hybrid model for handling out-of-vocabulary words in Japanese-to-English statistical machine translation output by exploiting parallel corpus. As the Japanese writing system makes use of four different script sets (kanji, hiragana, katakana, and romaji), we treat these scripts differently. A machine transliteration model is built to transliterate out-of-vocabulary Japanese katakana words into English words. A Japanese dependency structure analyzer is employed to tackle out-of-vocabulary kanji and hiragana words. The evaluation results demonstrate that it is an effective approach for addressing out-of-vocabulary word problems and decreasing the OOVs rate in the Japanese-to-English machine translation tasks.

主題

この書誌の出所

  • openalex— W2401186096(2026-08-14取得)

引用

Juan Luo・Yves Lepage(2015-01-01) Handling of Out-of-vocabulary Words in Japanese-English Machine Translation by Exploiting Parallel Corpus. 23 pp. 1-20

LuoLepage2015HandlingOutVocabulary
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON