論文 ·日本語 ·未確認
Large Context End-to-end Automatic Speech Recognition via Extension of Hierarchical Recurrent Encoder-decoder Models
Ryo Masumura ・ Tomohiro Tanaka ・ Takafumi Moriya ・ Yusuke Shinohara ・ Takanobu Oba ・ Yushi Aono
- 刊行年
- 2019-04-17
- 言語
- 英語
- OpenAlex
- W2937649809
- DOI
- 10.1109/icassp.2019.8683843
- MAG
- 2937649809
- URL
- https://doi.org/10.1109/icassp.2019.8683843
要旨
This paper describes a novel end-to-end automatic speech recognition (ASR) method that takes into consideration long-range sequential context information beyond utterance boundaries. In spontaneous ASR tasks such as those for discourses and conversations, the input speech often comprises a series of utterances. Accordingly, the relationships between the utterances should be leveraged for transcribing the individual utterances. While most previous end-to-end ASR methods only focus on utterance-level ASR that handles single utterances independently, the proposed method (which we call "large-context end-to-end ASR") can explicitly utilize relationships between a current target utterance and all preceding utterances. The method is modeled by combining an attention-based encoder-decoder model, which is one of the most representative end-to-end ASR models, with hierarchical recurrent encoder-decoder models, which are effective language models for capturing long-range sequential contexts beyond the utterance boundaries. Experiments on Japanese discourse speech tasks demonstrate the proposed method yields significant ASR performance improvements compared with the conventional utterance-level end-to-end ASR system.
主題
この書誌の出所
- openalex— W2937649809(2026-08-14取得)
引用
Ryo Masumura・Tomohiro Tanaka・Takafumi Moriya・Yusuke Shinohara・Takanobu Oba・Yushi Aono(2019-04-17) Large Context End-to-end Automatic Speech Recognition via Extension of Hierarchical Recurrent Encoder-decoder Models pp. 5661-5665