本文へ移動

論文 ·日本語 ·未確認

Improving Neural Text Normalization with Data Augmentation at Character- and Morphological Levels

Itsumi Saito Jun Suzuki Kyosuke Nishida Kugatsu Sadamitsu Satoshi Kobashikawa Ryo Masumura Yūji Matsumoto Junji Tomita

刊行年
2017-11-01
言語
英語
OpenAlex
W2773842746
MAG
2773842746
URL
https://openalex.org/W2773842746

要旨

In this study, we investigated the effectiveness of augmented data for encoder-decoder-based neural normalization models. Attention based encoder-decoder models are greatly effective in generating many natural languages. % such as machine translation or machine summarization. In general, we have to prepare for a large amount of training data to train an encoder-decoder model. Unlike machine translation, there are few training data for text-normalization tasks. In this paper, we propose two methods for generating augmented data. The experimental results with Japanese dialect normalization indicate that our methods are effective for an encoder-decoder model and achieve higher BLEU score than that of baselines. We also investigated the oracle performance and revealed that there is sufficient room for improving an encoder-decoder model.

主題

この書誌の出所

  • openalex— W2773842746(2026-08-14取得)

引用

Itsumi Saito・Jun Suzuki・Kyosuke Nishida・Kugatsu Sadamitsu・Satoshi Kobashikawa・Ryo Masumura・Yūji Matsumoto・Junji Tomita(2017-11-01) Improving Neural Text Normalization with Data Augmentation at Character- and Morphological Levels 2 pp. 257-262

Saito2017ImprovingNeuralText
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON