本文へ移動

論文 ·日本語 ·未確認

NWJC2Vec

Masayuki Asahara

刊行年
2018-05-31
収録
『Terminology International Journal of Theoretical and Applied Issues in Specialized Communication』 24(1) pp. 7-22
出版
John Benjamins Publishing Company
言語
英語
OpenAlex
W2807444963
DOI
10.1075/term.00011.asa
MAG
2807444963
ISSN
0929-9971
URL
https://doi.org/10.1075/term.00011.asa

要旨

Abstract In this paper, we present a word embedding dataset NWJC2Vec constructed using ‘NINJAL Web Japanese Corpus (NWJC)’. NWJC is a Web-crawled text corpus that contains 25.8 billion tokens. We construct two types of the word embedding dataset: one is based on the surface form, and the other is based on the complete morpheme information provided by UniDic, which is a lexicon for the Japanese morphological analyser MeCab. We perform an evaluation of the dataset by comparing it with the ‘Word List by Semantic Principles (Bunrui Goihyo)’.

主題

この書誌の出所

  • openalex— W2807444963(2026-08-12取得)

引用

Masayuki Asahara(2018-05-31) NWJC2Vec 『Terminology International Journal of Theoretical and Applied Issues in Specialized Communication』 24(1) pp. 7-22 John Benjamins Publishing Company

Asahara2018NWJCVec
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON