論文 ·日本語 ·未確認

NWJC2Vec

Masayuki Asahara

刊行年
2018-05-31
収録
『Terminology International Journal of Theoretical and Applied Issues in Specialized Communication』 24(1) pp. 7-22
出版
John Benjamins Publishing Company
言語
英語
openalex
W2807444963
doi
10.1075/term.00011.asa
mag
2807444963
issn
0929-9971
URL
https://doi.org/10.1075/term.00011.asa

要旨

Abstract In this paper, we present a word embedding dataset NWJC2Vec constructed using ‘NINJAL Web Japanese Corpus (NWJC)’. NWJC is a Web-crawled text corpus that contains 25.8 billion tokens. We construct two types of the word embedding dataset: one is based on the surface form, and the other is based on the complete morpheme information provided by UniDic, which is a lexicon for the Japanese morphological analyser MeCab. We perform an evaluation of the dataset by comparing it with the ‘Word List by Semantic Principles (Bunrui Goihyo)’.

主題

この書誌の出所

  • openalex— W2807444963(2026-08-12取得)

引用キー: Asahara2018NWJCVec

書誌 34,896件 語別索引 17,251件 資源 113件 研究者 303名 JSON