本文へ移動

学位論文 ·日本語 ·未確認

Japanese Text Normalization with Encoder-Decoder Model

Taishi Ikeda Hiroyuki Shindo Yuji Matsumoto

刊行年
2016-12-01
出版
National Institute of Informatics
言語
英語
OpenAlex
W2737886250
MAG
2737886250
URL
https://naist.repo.nii.ac.jp/records/8246

要旨

Text normalization is the task of transforming lexical variants to their canonical forms. We model the problem of text normalization as a character-level sequence to sequence learning problem and present a neural encoder-decoder model for solving it. To train the encoder-decoder model, many sentences pairs are generally required. However, Japanese non-standard canonical pairs are scarce in the form of parallel corpora. To address this issue, we propose a method of data augmentation to increase data size by converting existing resources into synthesized non-standard forms using handcrafted rules. We conducted an experiment to demonstrate that the synthesized corpus contributes to stably train an encoder-decoder model and improve the performance of Japanese text normalization.

主題

この書誌の出所

  • openalex— W2737886250(2026-08-14取得)

引用

Taishi Ikeda・Hiroyuki Shindo・Yuji Matsumoto(2016-12-01) Japanese Text Normalization with Encoder-Decoder Model pp. 129-137 National Institute of Informatics

Ikeda2016JapaneseTextNormalization
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON