論文 ·日本語 ·未確認
Creating dialect sub-corpora by clustering: a case in Japanese for an adaptive method
- 刊行年
- 2018-05-01
- 言語
- 英語
- OpenAlex
- W2807616547
- DOI
- 10.63317/388yq5ns8e45
- MAG
- 2807616547
- URL
- http://www.lrec-conf.org/proceedings/lrec2018/pdf/142.pdf
要旨
We propose a pipeline through which to derive clusters of dialects, given a mixed corpus composed of different dialects, when their standard counterpart is sufficiently resourced.The test case is Japanese, where the written standard language is sufficiently equipped with adequate resources.Our method starts by detecting non-standard contents first, and then clusters what is deemed dialectal.We report the results on the clustering of mixed Twitter corpus into four dialects (Kansai, Tohoku, Chugoku and Kyushu).
主題
この書誌の出所
- openalex— W2807616547(2026-08-14取得)
引用
Yo Sato・Kevin Heffernan(2018-05-01) Creating dialect sub-corpora by clustering: a case in Japanese for an adaptive method
SatoHeffernan2018CreatingDialectSub