本文へ移動

論文 ·日本語 ·未確認

Japanese Unknown Word Identification by Character-based Chunking

浅原 正幸 ・ 松本 裕治 ・ Masayuki Asahara ・ Yūji Matsumoto

刊行年
2004
収録
『Proceedings of 20th International Conference on Computational Linguistics 20』 pp. 459-465
言語
英語
CiNii CRID
1010000781802385555
OpenAlex
W2007619971
DOI
10.3115/1220355.1220421
MAG
2007619971
URL
https://cir.nii.ac.jp/crid/1010000781802385555

要旨

We introduce a character-based chunking for unknown word identification in Japanese text. A major advantage of our method is an ability to detect low frequency unknown words of unrestricted character type patterns. The method is built upon SVM-based chunking, by use of character n-gram and surrounding context of n-best word segmentation candidates from statistical morphological analysis as features. It is applied to newspapers and patent texts, achieving 95% precision and 55-70% recall for newspapers and more than 85% precision for patent texts.

主題

この書誌の出所

  • cinii— 1010000781802385555(2026-08-12取得)
  • openalex— W2007619971(2026-08-14取得)

引用

浅原 正幸・松本 裕治・Masayuki Asahara・Yūji Matsumoto(2004) Japanese Unknown Word Identification by Character-based Chunking 『Proceedings of 20th International Conference on Computational Linguistics 20』 pp. 459-465

浅原2004JapaneseUn
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON