本文へ移動

論文 ·日本語 ·未確認

Creating dialect sub-corpora by clustering: a case in Japanese for an adaptive method

Yo Sato Kevin Heffernan

刊行年
2018-05-01
言語
英語
OpenAlex
W2807616547
DOI
10.63317/388yq5ns8e45
MAG
2807616547
URL
http://www.lrec-conf.org/proceedings/lrec2018/pdf/142.pdf

要旨

We propose a pipeline through which to derive clusters of dialects, given a mixed corpus composed of different dialects, when their standard counterpart is sufficiently resourced.The test case is Japanese, where the written standard language is sufficiently equipped with adequate resources.Our method starts by detecting non-standard contents first, and then clusters what is deemed dialectal.We report the results on the clustering of mixed Twitter corpus into four dialects (Kansai, Tohoku, Chugoku and Kyushu).

主題

この書誌の出所

  • openalex— W2807616547(2026-08-14取得)

引用

Yo Sato・Kevin Heffernan(2018-05-01) Creating dialect sub-corpora by clustering: a case in Japanese for an adaptive method

SatoHeffernan2018CreatingDialectSub
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON