論文 ·対照・比較 ·未確認
Finding ideographic representations of Japanese names written in Latin script via language identification and corpus validation
- 刊行年
- 2004-01-01
- 言語
- 英語
- openalex
- W2139755556
- doi
- 10.3115/1218955.1218979
- mag
- 2139755556
- URL
- http://dl.acm.org/ft_gateway.cfm?id=1218979&type=pdf
要旨
Multilingual applications frequently involve dealing with proper names, but names are often missing in bilingual lexicons. This problem is exacerbated for applications involving translation between Latin-scripted languages and Asian languages such as Chinese, Japanese and Korean (CJK) where simple string copying is not a solution. We present a novel approach for generating the ideographic representations of a CJK name written in a Latin script. The proposed approach involves first identifying the origin of the name, and then back-transliterating the name to all possible Chinese characters using language-specific mappings. To reduce the massive number of possibilities for computation, we apply a three-tier filtering process by filtering first through a set of attested bigrams, then through a set of attested terms, and lastly through the WWW for a final validation. We illustrate the approach with English-to-Japanese back-transliteration. Against test sets of Japanese given names and surnames, we have achieved average precisions of 73% and 90%, respectively.
主題
この書誌の出所
- openalex— W2139755556(2026-08-13取得)
引用キー: QuGrefenstette2004FindingIdeographicRepresentations