本文へ移動

論文 ·用例に日本語 ·未確認

Joint CTC/attention decoding for end-to-end speech recognition

Takaaki Hori Shinji Watanabe John R. Hershey

刊行年
2017-01-01
言語
英語
OpenAlex
W2739883972
DOI
10.18653/v1/p17-1048
MAG
2739883972
URL
https://www.aclweb.org/anthology/P17-1048.pdf

要旨

End-to-end automatic speech recognition (ASR) has become a popular alternative to conventional DNN/HMM systems because it avoids the need for linguistic resources such as pronunciation dictionary, tokenization, and contextdependency trees, leading to a greatly simplified model-building process. There are two major types of end-to-end architectures for ASR: attention-based methods use an attention mechanism to perform alignment between acoustic frames and recognized symbols, and connectionist temporal classification (CTC), uses Markov assumptions to efficiently solve sequential problems by dynamic programming. This paper proposes a joint decoding algorithm for end-to-end ASR with a hybrid CTC/attention architecture, which effectively utilizes both advantages in decoding. We have applied the proposed method to two ASR benchmarks (spontaneous Japanese and Mandarin Chinese), and showing the comparable performance to conventional state-of-the-art DNN/HMM ASR systems without linguistic resources.

主題

この書誌の出所

  • openalex— W2739883972(2026-08-14取得)

引用

Takaaki Hori・Shinji Watanabe・John R. Hershey(2017-01-01) Joint CTC/attention decoding for end-to-end speech recognition pp. 518-529

Hori2017JointCTCAttention
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON