本文へ移動

論文 ·日本語 ·未確認

A Parameter-Efficient Multi-Step Fine-Tuning of Multilingual and Multi-Task Learning Model for Japanese Dialect Speech Recognition

Yuta Kamiya Shogo Miwa Atsuhiko Kai

刊行年
2024-10-17
言語
英語
OpenAlex
W4405633822
DOI
10.1109/o-cocosda64382.2024.10800608
URL
https://doi.org/10.1109/o-cocosda64382.2024.10800608

要旨

This paper addresses the challenge of developing a unified spoken language model for Japanese dialects. Due to the limited availability of such speech resources, self-supervised learning (SSL) methods show promise. However, the diversity of Japanese dialects also requires multilingual approaches to language modeling. One of the few related works improved performance through multi-step fine-tuning, using standard Japanese speech and multitask learning that incorporates dialect identification. However, this method is computationally expensive and runs the risk of forgetting acquired knowledge. To address these challenges, we propose a large-scale multilingual SSL model-based multistage fine-tuning strategy using lightweight adapter modules per domain. Our method achieves a relative reduction in character error rate of 18% compared to simple full fine-tuning method, especially in the speaker adaptation scenario, while using only half the training parameters required for full fine-tuning.

主題

この書誌の出所

  • openalex— W4405633822(2026-08-14取得)

引用

Yuta Kamiya・Shogo Miwa・Atsuhiko Kai(2024-10-17) A Parameter-Efficient Multi-Step Fine-Tuning of Multilingual and Multi-Task Learning Model for Japanese Dialect Speech Recognition pp. 1-6

Kamiya2024ParameterEfficientMulti
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON