本文へ移動

論文 ·日本語 ·未確認

Automatically Extracting Variant-Normalization Pairs for Japanese Text Normalization

Itsumi Saito Kyosuke Nishida Kugatsu Sadamitsu Kuniko Saito Junji Tomita

刊行年
2017-11-01
言語
英語
OpenAlex
W2771291041
MAG
2771291041
URL
https://openalex.org/W2771291041

要旨

Social media texts, such as tweets from Twitter, contain many types of non-standard tokens, and the number of normalization approaches for handling such noisy text has been increasing. We present a method for automatically extracting pairs of a variant word and its normal form from unsegmented text on the basis of a pair-wise similarity approach. We incorporated the acquired variant-normalization pairs into Japanese morphological analysis. The experimental results show that our method can extract widely covered variants from large Twitter data and improve the recall of normalization without degrading the overall accuracy of Japanese morphological analysis.

主題

この書誌の出所

  • openalex— W2771291041(2026-08-14取得)

引用

Itsumi Saito・Kyosuke Nishida・Kugatsu Sadamitsu・Kuniko Saito・Junji Tomita(2017-11-01) Automatically Extracting Variant-Normalization Pairs for Japanese Text Normalization 1 pp. 937-946

Saito2017AutomaticallyExtractingVariant
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON