本文へ移動

論文 ·日本語 ·未確認

Automatically Annotating A Five-Billion-Word Corpus of Japanese Blogs for Affect and Sentiment Analysis

Michał Ptaszyński Rafał Rzepka Kenji Araki Yoshio Momouchi

刊行年
2012-07-12
言語
英語
OpenAlex
W2097891230
MAG
2097891230
URL
https://openalex.org/W2097891230

要旨

This paper presents our research on automatic annotation of a five-billion-word corpus of Japanese blogs with information on affect and sentiment. We first perform a study in emotion blog corpora to discover that there has been no large scale emotion corpus available for the Japanese language. We choose the largest blog corpus for the language and annotate it with the use of two systems for affect analysis: ML-Ask for word- and sentence-level affect analysis and CAO for detailed analysis of emoticons. The annotated information includes affective features like sentence subjectivity (emotive/non-emotive) or emotion classes (joy, sadness, etc.), useful in affect analysis. The annotations are also generalized on a 2-dimensional model of affect to obtain information on sentence valence/polarity (positive/negative) useful in sentiment analysis. The annotations are evaluated in several ways. Firstly, on a test set of a thousand sentences extracted randomly and evaluated by over forty respondents. Secondly, the statistics of annotations are compared to other existing emotion blog corpora. Finally, the corpus is applied in several tasks, such as generation of emotion object ontology or retrieval of emotional and moral consequences of actions. 1

主題

この書誌の出所

  • openalex— W2097891230(2026-08-14取得)

引用

Michał Ptaszyński・Rafał Rzepka・Kenji Araki・Yoshio Momouchi(2012-07-12) Automatically Annotating A Five-Billion-Word Corpus of Japanese Blogs for Affect and Sentiment Analysis pp. 89-98

Ptaszyński2012AutomaticallyAnnotatingFive
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON