本文へ移動

論文 ·日本語 ·未確認

Morphological Analysis of the Corpus of Spontaneous Japanese

Kiyotaka Uchimoto Kazuhiro Takaoka Chikashi Nobata Atsushi Yamada Shuichi Sekine Hitoshi Isahara

刊行年
2004-06-21
収録
『IEEE Transactions on Speech and Audio Processing』 12(4) pp. 382-390
出版
Institute of Electrical and Electronics Engineers
言語
英語
OpenAlex
W2167915960
DOI
10.1109/tsa.2004.828700
MAG
2167915960
ISSN
1063-6676
URL
https://doi.org/10.1109/tsa.2004.828700

要旨

This paper describes two methods for detecting word segments and their morphological information in a Japanese spontaneous speech corpus, and describes how to tag a large spontaneous speech corpus accurately by using the two methods. The first method is used to detect any type of word segments. The second method is used when there are several definitions for word segments and their POS categories, and when one type of word segments includes another type of word segments. In this paper, we show that by using semi-automatic analysis, we achieve a precision of better than 99% for detecting and tagging short-unit words and 97% for long-unit words; the two types of words that comprise the corpus. We also show that better accuracy is achieved by using both methods than by using only the first.

主題

この書誌の出所

  • openalex— W2167915960(2026-08-14取得)

引用

Kiyotaka Uchimoto・Kazuhiro Takaoka・Chikashi Nobata・Atsushi Yamada・Shuichi Sekine・Hitoshi Isahara(2004-06-21) Morphological Analysis of the Corpus of Spontaneous Japanese 『IEEE Transactions on Speech and Audio Processing』 12(4) pp. 382-390 Institute of Electrical and Electronics Engineers

Uchimoto2004MorphologicalAnalysisCorpus
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON