論文 ·日本語 ·未確認

Keyword Extraction from a Document using Word Co-occurrence Statistical Information

Yutaka Matsuo Mitsuru Ishizuka

刊行年
2002-01-01
収録
『Transactions of the Japanese Society for Artificial Intelligence』 17(3) pp. 217-223
出版
The Japanese Society for Artificial Intelligence
言語
英語
openalex
W140751283
doi
10.1527/tjsai.17.217
mag
140751283
issn
1346-0714
URL
https://www.jstage.jst.go.jp/article/tjsai/17/3/17_3_217/_pdf

要旨

We present a new keyword extraction algorithm that applies to a single document without using a large corpus. Frequent terms are extracted first, then a set of co-occurrence between each term and the frequent terms, i.e., occurrences in the same sentences, is generated. The distribution of co-occurrence shows the importance of a term in the document as follows. If the probability distribution of co-occurrence between term a and the frequent terms is biased to a particular subset of the frequent terms, then term a is likely to be a keyword. The degree of the biases of the distribution is measured by χ²-measure. We show our algorithm performs well for indexing technical papers.

主題

この書誌の出所

  • openalex— W140751283(2026-08-13取得)

引用キー: MatsuoIshizuka2002KeywordExtractionDocument

書誌 56,535件 語別索引 17,251件 資源 113件 研究者 303名 JSON