論文 ·日本語 ·未確認

武道歌の計量テキスト分析における 形態素解析辞書の選択とテキストデータの加工

Katsunori Kobayashi Koya Satoh

刊行年
2025-03-16
出版
National Institute of Informatics
言語
日本語
openalex
W7143721967
doi
10.15034/0002002413
URL
https://bunkyo.repo.nii.ac.jp/records/2002413

要旨

Martial arts poems are 31-character, fixed-form poems that describe the principles, techniques, and spirit of martial arts training. When performing text mining on the content of these poems, it is necessary to select an appropriate morphological analysis dictionary. This study aims to investigate how to make that selection. Additionally, unlike prose, these poems contain no punctuation marks and use specific rhetorical methods characteristic of tanka. Therefore, we examined whether processing the text could improve the accuracy of the analysis. For the analysis, we used the martial arts poem anthology “Tsukahara Bokuden Ikunsho,”which consists of 96 poems composed by Tsukahara Bokuden at the end of the Muromachi period. Morphological analysis was performed using the analysis software “Web Chamame,” developed by the National Institute for Japanese Language and Linguistics. The dictionaries found to be suitable for this collection of poems were the “Middle Japanese Literary UniDic,” “Middle Japanese Colloquial UniDic,” and “Early Modern Japanese Literary UniDic,” all developed by the National Institute for Japanese Language and Linguistics. The morphological analysis performed using these dictionaries demonstrated significant accuracy, approximately 97%. This indicates that quantitative text analysis of martial arts poems requires selecting a morphological analysis dictionary appropriate for the period and literary style of the poems. The most suitable method involves analyzing with software, such as “Web Chamame,” and selecting the most appropriate dictionary from multiple options.Furthermore, this research concludes that a dictionary selection guideline of 95% analysis accuracy is appropriate. Additionally, it was suggested that when extracting samples to compare analysis accuracy, a minimum of 60 poems or over 1,000 morphemes is required. However, more studies are needed to determine whether these results are generally applicable. Lastly, we found no decrease in analytical accuracy due to the combination of compound sentences and sentence-ending particles—a rhetorical device unique to tanka. Therefore, we conclude that it is not necessary to process text data, such as by adding punctuation or dividing sentences.

主題

この書誌の出所

  • openalex— W7143721967(2026-08-12取得)

引用キー: KobayashiSatoh2025

書誌 34,896件 語別索引 17,251件 資源 113件 研究者 303名 JSON