論文 ·日本語 ·未確認
A Parameter-Efficient Multi-Step Fine-Tuning of Multilingual and Multi-Task Learning Model for Japanese Dialect Speech Recognition
Yuta Kamiya ・ Shogo Miwa ・ Atsuhiko Kai
- 刊行年
- 2024-10-17
- 言語
- 英語
- OpenAlex
- W4405633822
- DOI
- 10.1109/o-cocosda64382.2024.10800608
- URL
- https://doi.org/10.1109/o-cocosda64382.2024.10800608
要旨
This paper addresses the challenge of developing a unified spoken language model for Japanese dialects. Due to the limited availability of such speech resources, self-supervised learning (SSL) methods show promise. However, the diversity of Japanese dialects also requires multilingual approaches to language modeling. One of the few related works improved performance through multi-step fine-tuning, using standard Japanese speech and multitask learning that incorporates dialect identification. However, this method is computationally expensive and runs the risk of forgetting acquired knowledge. To address these challenges, we propose a large-scale multilingual SSL model-based multistage fine-tuning strategy using lightweight adapter modules per domain. Our method achieves a relative reduction in character error rate of 18% compared to simple full fine-tuning method, especially in the speaker adaptation scenario, while using only half the training parameters required for full fine-tuning.
主題
この書誌の出所
- openalex— W4405633822(2026-08-14取得)
引用
Yuta Kamiya・Shogo Miwa・Atsuhiko Kai(2024-10-17) A Parameter-Efficient Multi-Step Fine-Tuning of Multilingual and Multi-Task Learning Model for Japanese Dialect Speech Recognition pp. 1-6