論文 ·対照・比較 ·未確認

Finding ideographic representations of Japanese names written in Latin script via language identification and corpus validation

Yan Qu Gregory Grefenstette

刊行年
2004-01-01
言語
英語
openalex
W2139755556
doi
10.3115/1218955.1218979
mag
2139755556
URL
http://dl.acm.org/ft_gateway.cfm?id=1218979&type=pdf

要旨

Multilingual applications frequently involve dealing with proper names, but names are often missing in bilingual lexicons. This problem is exacerbated for applications involving translation between Latin-scripted languages and Asian languages such as Chinese, Japanese and Korean (CJK) where simple string copying is not a solution. We present a novel approach for generating the ideographic representations of a CJK name written in a Latin script. The proposed approach involves first identifying the origin of the name, and then back-transliterating the name to all possible Chinese characters using language-specific mappings. To reduce the massive number of possibilities for computation, we apply a three-tier filtering process by filtering first through a set of attested bigrams, then through a set of attested terms, and lastly through the WWW for a final validation. We illustrate the approach with English-to-Japanese back-transliteration. Against test sets of Japanese given names and surnames, we have achieved average precisions of 73% and 90%, respectively.

主題

この書誌の出所

  • openalex— W2139755556(2026-08-13取得)

引用キー: QuGrefenstette2004FindingIdeographicRepresentations

書誌 56,535件 語別索引 17,251件 資源 113件 研究者 303名 JSON