論文 ·日本語 ·未確認

JWSAN: Japanese word similarity and association norm

Keisuke Inohara Akira Utsumi

刊行年
2021-06-18
収録
『Language Resources and Evaluation』 56(1) pp. 109-137
出版
Springer Science+Business Media
言語
英語
openalex
W3175891392
doi
10.1007/s10579-021-09543-7
mag
3175891392
issn
1574-020X
URL
https://link.springer.com/content/pdf/10.1007/s10579-021-09543-7.pdf

要旨

Abstract We present a new Japanese dataset, Japanese Word Similarity and Association Norm (JWSAN), comprising human rating scores of similarity and association for 2145 word pairs, with a clear distinction between word similarity and word association. Computational models of human semantic memory or mental lexicon, such as distributed semantic models, must predict not only association but also similarity. People can distinguish between word similarity and association. However, although the SimLex-999 dataset is publicly available for English, there is no Japanese similarity dataset with a clear distinction between the two types of word relatedness. JWSAN is the first large Japanese dataset with similarity and association ratings, containing noun, verb, and adjective word pairs. It is also characterized by data collection from a sufficient number of age- and-gender-controlled assessors, with similarity and association ratings obtained via a web-based survey conducted of 6450 native speakers of Japanese. In addition, the effects of the gender and age of the raters were also examined; these factors were only given scant consideration in the past. This dataset can act as a benchmark for improving distributed semantic models in Japanese.

主題

この書誌の出所

  • openalex— W3175891392(2026-08-12取得)

引用キー: InoharaUtsumi2021JWSAN

書誌 34,896件 語別索引 17,251件 資源 113件 研究者 303名 JSON