論文 ·日本語 ·未確認

A Web Corpus and Word Sketches for Japanese

Irena Srdanovic Erjavec Tomaz Erjavec Adam Kilgarriff Irena Srdanovic´ Erjavec Irena Srdanović Tomaž Erjavec

刊行年
2008
収録
『自然言語処理』 15(2) pp. 137-159
言語
英語
doi
10.5715/jnlp.15.2_137
issn
1340-7619
jstage_journal
jnlp1994
openalex
W1983774632
mag
1983774632
URL
https://www.jstage.jst.go.jp/article/jnlp1994/15/2/15_2_137/_article/-char/ja/

要旨

Of all the major world languages, Japanese is lagging behind in terms of publicly accessible and searchable corpora. In this paper we describe the development of JpWaC (Japanese Web as Corpus), a large corpus of 400 million words of Japanese web text, and its encoding for the Sketch Engine. The Sketch Engine is a web-based corpus query tool that supports fast concordancing, grammatical processing, ‘word sketching’ (one-page summaries of a word's grammatical and collocational behaviour), a distributional thesaurus, and robot use. We describe the steps taken to gather and process the corpus and to establish its validity, in terms of the kinds of language it contains. We then describe the development of a shallow grammar for Japanese to enable word sketching. We believe that the Japanese web corpus as loaded into the Sketch Engine will be a useful resource for a wide number of Japanese researchers, learners, and NLP developers.

主題

この書誌の出所

  • jstage— 10.5715/jnlp.15.2_137(2026-08-13取得)
  • jstage— 10.11185/imt.3.529(2026-08-13取得)
  • openalex— W1983774632(2026-08-13取得)

引用キー: Erjavec2008WebCorpusWord

書誌 53,525件 語別索引 17,251件 資源 113件 研究者 303名 JSON