学位論文 ·対照・比較 ·未確認
万葉集の未解読歌の解読
- 刊行年
- 2025-03-01
- 出版
- National Institute of Informatics
- 言語
- 日本語
- OpenAlex
- W7146619091
- URL
- https://dspace.jaist.ac.jp/dspace/bitstream/10119/19836/5/paper.pdf
要旨
Man'yoshu, Japan's oldest anthology of poetry, encompasses 4,516 poems, though 42 of them remain undeciphered.Previous studies assumed these compositions, written in Man'yogana (a writing system of using Chinese characters to represent Japanese phonemes), should be interpreted as Japanese.However, Vovin successfully demonstrated that Man'yoshu's poem number 9, which was previously undeciphered, became intelligible when it was read as Old Korean.The goal of this study is to decode the other undeciphered poems of Man'yoshu as Old Korean.We implement Vovin's method as a computational system to automatically or semi-automatically decipher undeciphered poems.Our proposed method primarily involves two modules.The first module converts sequences of Chinese characters into sequences of phonemes using the " Chinese character phoneme dictionary."The module then performs morphological analysis on these phoneme sequences using a Korean word dictionary, and finally generates sequences of Korean words.The second module enlarges our Chinese character phoneme dictionary by adding Middle Korean (MK) pronunciations.This is achieved by estimating Middle Korean pronunciations from Chinese pronunciations.To implement the first module, we begin by creating the Chinese character phoneme dictionary which defines the phonetic symbols associated with each Chinese character.This dictionary compiles five types of pronunciations: Man'yogana, Idu (a writing system of using Chinese character to represent Korean phonemes), Middle Korean pronunciations derived from Late Han Chinese (LHC) pronunciations, Middle Korean pronunciations derived from Early Middle Chinese (EMC) pronunciations, and jeongyong pronunciations that represent Korean phonemes intended to convey the Chinese character's meaning.For a given input sequence of Chinese characters, the system converts it to multiple possible sequences of phonemes by looking up the Chinese character phoneme dictionary for each of the input Chinese characters.Next, for each phonetic sequence, morphological analysis is performed to obtain sequences of Middle Korean words.MeCab is employed for this process, since it is the only morphological analyzer for which a Middle Korean word dictionary is available.Specifically, this study utilizes the MkHanDic dictionary with 9,653 words as the MK word dictionary.Then, the most appropriate word sequences are chosen from all generated sequences using the following two-step procedures.(1) Any grammatically incorrect word sequences are removed by the grammatical check.Specifically, we eliminate word sequences that begin with a verbal ending, begin with a
主題
この書誌の出所
- openalex— W7146619091(2026-08-14取得)
引用
啓晶 佐々木(2025-03-01) 万葉集の未解読歌の解読 National Institute of Informatics