本文へ移動

論文 ·日本語 ·未確認

Construction of an idiom corpus and its application to idiom identification based on WSD incorporating idiom-specific features

Chikara Hashimoto ・ Daisuke Kawahara

刊行年
2008-01-01
言語
英語
OpenAlex
W2045705339
DOI
10.3115/1613715.1613844
MAG
2045705339
URL
https://dl.acm.org/doi/pdf/10.5555/1613715.1613844

要旨

Some phrases can be interpreted either idiomatically (figuratively) or literally in context, and the precise identification of idioms is indispensable for full-fledged natural language processing (NLP). To this end, we have constructed an idiom corpus for Japanese. This paper reports on the corpus and the results of an idiom identification experiment using the corpus. The corpus targets 146 ambiguous idioms, and consists of 102, 846 sentences, each of which is annotated with a literal/idiom label. For idiom identification, we targeted 90 out of the 146 idioms and adopted a word sense disambiguation (WSD) method using both common WSD features and idiom-specific features. The corpus and the experiment are the largest of their kind, as far as we know. As a result, we found that a standard supervised WSD method works well for the idiom identification and achieved an accuracy of 89.25% and 88.86% with/without idiom-specific features and that the most effective idiom-specific feature is the one involving the adjacency of idiom constituents.

主題

この書誌の出所

  • openalex— W2045705339(2026-08-14取得)

引用

Chikara Hashimoto・Daisuke Kawahara(2008-01-01) Construction of an idiom corpus and its application to idiom identification based on WSD incorporating idiom-specific features pp. 992-992

HashimotoKawahara2008ConstructionIdiomCorpus
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON