論文 ·日本語 ·未確認
Context Adaptive Neural Network Based Acoustic Models for Rapid Adaptation
Marc Delcroix ・ Keisuke Kinoshita ・ Atsunori Ogawa ・ Christian Huemmer ・ Tomohiro Nakatani
- 刊行年
- 2018-01-26
- 収録
- 『IEEE/ACM Transactions on Audio Speech and Language Processing』 26(5) pp. 895-908
- 出版
- Institute of Electrical and Electronics Engineers
- 言語
- 英語
- OpenAlex
- W2791636785
- DOI
- 10.1109/taslp.2018.2798821
- MAG
- 2791636785
- ISSN
- 2329-9290
- URL
- https://doi.org/10.1109/taslp.2018.2798821
要旨
The adaptation of automatic speech recognition systems to a speaker or an environment is important if we are to achieve high speech recognition performance ubiquitously. Recently, deep neural network (DNN) based acoustic models have been made adaptive to speakers or environments by the addition of an auxiliary feature representing the acoustic context information such as speaker or noise characteristics to the network input. The addition of such auxiliary features to the input realizes only the adaptation of the bias term of the input layer. In this paper, we introduce “context adaptive neural networks,” which are an alternative approach for exploiting auxiliary features that can achieve adaptation of all the parameters of a layer including the linear transformation matrices and the bias terms. A context adaptive neural network is a neural network with one of its layers factorized into sublayers, each associated with an acoustic context class representing a class of speakers or noise conditions. The output of the factorized layer is obtained as a weighted sum of the contributions of all of the sublayers. The weighting coefficients, or context class weights, are derived from the auxiliary features, by transforming them through an auxiliary network. The auxiliary network and the main network can be trained jointly, which enables the context classes that optimize the training criterion to be learned automatically. We perform experiments on three tasks, i.e., two speaker adaptation experiments using DNN models with medium-sized (Wall Street Journal) and large (Continuous Spontaneous Japanese) training datasets, and one environmental adaptation of a convolutional neural network based acoustic model with CHiME3 data. These experiments confirm the potential of the proposed approach in various settings.
主題
この書誌の出所
- openalex— W2791636785(2026-08-14取得)
引用
Marc Delcroix・Keisuke Kinoshita・Atsunori Ogawa・Christian Huemmer・Tomohiro Nakatani(2018-01-26) Context Adaptive Neural Network Based Acoustic Models for Rapid Adaptation 『IEEE/ACM Transactions on Audio Speech and Language Processing』 26(5) pp. 895-908 Institute of Electrical and Electronics Engineers