論文 ·日本語 ·未確認
CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects
Shinnosuke Takamichi ・ Hiroshi Saruwatari
- 刊行年
- 2018-05-01
- 言語
- 英語
- OpenAlex
- W2806666542
- DOI
- 10.63317/23papnq84qja
- MAG
- 2806666542
- URL
- http://www.lrec-conf.org/proceedings/lrec2018/pdf/67.pdf
要旨
Public parallel corpora of dialects can accelerate related studies such as spoken language processing.Various corpora have been collected using a well-equipped recording environment, such as voice recording in an anechoic room.However, due to geographical and expense issues, it is impossible to use such a perfect recording environment for collecting all existing dialects.To address this problem, we used web-based recording and crowdsourcing platforms to construct a crowdsourced parallel speech corpus of Japanese dialects (CPJD corpus) including parallel text and speech data of 21 Japanese dialects.We recruited native dialect speakers on the crowdsourcing platform, and the hired speakers recorded their dialect speech using their personal computer or smartphone in their homes.This paper shows the results of the data collection and analyzes the audio data in terms of the signal-to-noise ratio and mispronunciations.
主題
この書誌の出所
- openalex— W2806666542(2026-08-14取得)
引用
Shinnosuke Takamichi・Hiroshi Saruwatari(2018-05-01) CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects