本文へ移動

論文 ·日本語 ·未確認

CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects

Shinnosuke Takamichi Hiroshi Saruwatari

刊行年
2018-05-01
言語
英語
OpenAlex
W2806666542
DOI
10.63317/23papnq84qja
MAG
2806666542
URL
http://www.lrec-conf.org/proceedings/lrec2018/pdf/67.pdf

要旨

Public parallel corpora of dialects can accelerate related studies such as spoken language processing.Various corpora have been collected using a well-equipped recording environment, such as voice recording in an anechoic room.However, due to geographical and expense issues, it is impossible to use such a perfect recording environment for collecting all existing dialects.To address this problem, we used web-based recording and crowdsourcing platforms to construct a crowdsourced parallel speech corpus of Japanese dialects (CPJD corpus) including parallel text and speech data of 21 Japanese dialects.We recruited native dialect speakers on the crowdsourcing platform, and the hired speakers recorded their dialect speech using their personal computer or smartphone in their homes.This paper shows the results of the data collection and analyzes the audio data in terms of the signal-to-noise ratio and mispronunciations.

主題

この書誌の出所

  • openalex— W2806666542(2026-08-14取得)

引用

Shinnosuke Takamichi・Hiroshi Saruwatari(2018-05-01) CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects

TakamichiSaruwatari2018CPJDCorpus
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON