論文 ·日本語 ·未確認
NWJC2Vec
- 刊行年
- 2018-05-31
- 収録
- 『Terminology International Journal of Theoretical and Applied Issues in Specialized Communication』 24(1) pp. 7-22
- 出版
- John Benjamins Publishing Company
- 言語
- 英語
- openalex
- W2807444963
- doi
- 10.1075/term.00011.asa
- mag
- 2807444963
- issn
- 0929-9971
- URL
- https://doi.org/10.1075/term.00011.asa
要旨
Abstract In this paper, we present a word embedding dataset NWJC2Vec constructed using ‘NINJAL Web Japanese Corpus (NWJC)’. NWJC is a Web-crawled text corpus that contains 25.8 billion tokens. We construct two types of the word embedding dataset: one is based on the surface form, and the other is based on the complete morpheme information provided by UniDic, which is a lexicon for the Japanese morphological analyser MeCab. We perform an evaluation of the dataset by comparing it with the ‘Word List by Semantic Principles (Bunrui Goihyo)’.
主題
この書誌の出所
- openalex— W2807444963(2026-08-12取得)
引用キー: Asahara2018NWJCVec