論文 ·日本語 ·未確認
Japanese Unknown Word Identification by Character-based Chunking
浅原 正幸 ・ 松本 裕治 ・ Masayuki Asahara ・ Yūji Matsumoto
- 刊行年
- 2004
- 収録
- 『Proceedings of 20th International Conference on Computational Linguistics 20』 pp. 459-465
- 言語
- 英語
- CiNii CRID
- 1010000781802385555
- OpenAlex
- W2007619971
- DOI
- 10.3115/1220355.1220421
- MAG
- 2007619971
- URL
- https://cir.nii.ac.jp/crid/1010000781802385555
要旨
We introduce a character-based chunking for unknown word identification in Japanese text. A major advantage of our method is an ability to detect low frequency unknown words of unrestricted character type patterns. The method is built upon SVM-based chunking, by use of character n-gram and surrounding context of n-best word segmentation candidates from statistical morphological analysis as features. It is applied to newspapers and patent texts, achieving 95% precision and 55-70% recall for newspapers and more than 85% precision for patent texts.
主題
この書誌の出所
- cinii— 1010000781802385555(2026-08-12取得)
- openalex— W2007619971(2026-08-14取得)
引用
浅原 正幸・松本 裕治・Masayuki Asahara・Yūji Matsumoto(2004) Japanese Unknown Word Identification by Character-based Chunking 『Proceedings of 20th International Conference on Computational Linguistics 20』 pp. 459-465