本文へ移動

論文 ·日本語 ·未確認

Controlled and Balanced Dataset for Japanese Lexical Simplification

Tomonori Kodaira Tomoyuki Kajiwara Mamoru Komachi

刊行年
2016-01-01
言語
英語
OpenAlex
W2510721067
DOI
10.18653/v1/p16-3001
MAG
2510721067
URL
https://www.aclweb.org/anthology/P16-3001.pdf

要旨

We propose a new dataset for evaluating a Japanese lexical simplification method. Previous datasets have several deficiencies. All of them substitute only a single target word, and some of them extract sentences only from newswire corpus. In addition, most of these datasets do not allow ties and integrate simplification ranking from all the annotators without considering the quality. In contrast, our dataset has the following advantages: (1) it is the first controlled and balanced dataset for Japanese lexical simplification with high correlation with human judgment and (2) the consistency of the simplification ranking is improved by allowing candidates to have ties and by considering the reliability of annotators.

主題

この書誌の出所

  • openalex— W2510721067(2026-08-14取得)

引用

Tomonori Kodaira・Tomoyuki Kajiwara・Mamoru Komachi(2016-01-01) Controlled and Balanced Dataset for Japanese Lexical Simplification

Kodaira2016ControlledBalancedDataset
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON