論文 ·用例に日本語 ·未確認

Wikipedia Titles As Noun Tag Predictors

Armin Hoenen

刊行年
2016-05-01
言語
英語
openalex
W2578679828
doi
10.63317/58x54dtvvtxq
mag
2578679828
URL
http://www.lrec-conf.org/proceedings/lrec2016/pdf/18_Paper.pdf

要旨

In this paper, we investigate a covert labeling cue, namely the probability that a title (by example of the Wikipedia titles) is a noun.If this probability is very large, any list such as or comparable to the Wikipedia titles can be used as a reliable word-class (or part-of-speech tag) predictor or noun lexicon.This may be especially useful in the case of Low Resource Languages (LRL) where labeled data is lacking and putatively for Natural Language Processing (NLP) tasks such as Word Sense Disambiguation, Sentiment Analysis and Machine Translation.Profitting from the ease of digital publication on the web as opposed to print, LRL speaker communities produce resources such as Wikipedia and Wiktionary, which can be used for an assessment.We provide statistical evidence for a strong noun bias for the Wikipedia titles from 2 corpora (English, Persian) and a dictionary (Japanese) and for a typologically balanced set of 17 languages including LRLs.Additionally, we conduct a small experiment on predicting noun tags for out-of-vocabulary items in part-of-speech tagging for English.

主題

この書誌の出所

  • openalex— W2578679828(2026-08-12取得)

引用キー: Hoenen2016WikipediaTitlesNoun

書誌 34,896件 語別索引 17,251件 資源 113件 研究者 303名 JSON