論文 ·日本語 ·未確認

JBLiMP: Japanese Benchmark of Linguistic Minimal Pairs

Taiga Someya Yohei Oseki

刊行年
2023-01-01
言語
英語
openalex
W4386566493
doi
10.18653/v1/2023.findings-eacl.117
URL
https://aclanthology.org/2023.findings-eacl.117.pdf

要旨

In this paper, we introduce JBLiMP (Japanese Benchmark of Linguistic Minimal Pairs), a novel dataset for targeted syntactic evaluations of language models in Japanese. JBLiMP consists of 331 minimal pairs, which are created based on acceptability judgments extracted from journal articles in theoretical linguistics. These minimal pairs are grouped into 11 categories, each covering a different linguistic phenomenon. JBLiMP is unique in that it successfully combines two important features independently observed in existing datasets: (i) coverage of complex linguistic phenomena (cf. CoLA) and (ii) presentation of sentences as minimal pairs (cf. BLiMP). In addition, JBLiMP is the first dataset for targeted syntactic evaluations of language models in Japanese, thus allowing the comparison of syntactic knowledge of language models across different languages. We then evaluate the syntactic knowledge of several language models on JBLiMP: GPT-2, LSTM, and n-gram language models. The results demonstrated that all the architectures achieved comparable overall accuracies around 75%. Error analyses by linguistic phenomenon further revealed that these language models successfully captured local dependencies like nominal structures, but not long-distance dependencies such as verbal agreement and binding.

主題

この書誌の出所

  • openalex— W4386566493(2026-08-13取得)

引用キー: SomeyaOseki2023JBLiMP

書誌 56,535件 語別索引 17,251件 資源 113件 研究者 303名 JSON