本文へ移動

論文 ·日本語 ·未確認

Crowdsourced Corpus of Sentence Simplification with Core Vocabulary

Akihiro Katsuta ・ Kazuhide Yamamoto

刊行年
2018-05-01
言語
英語
OpenAlex
W2806010723
DOI
10.63317/2a5ax2twwiob
MAG
2806010723
URL
http://www.lrec-conf.org/proceedings/lrec2018/pdf/327.pdf

要旨

We present a new Japanese crowdsourced data set of simplified sentences created from more complex ones.Our simplicity standard involves all rewritable words in the simplified sentences being drawn from a core vocabulary of 2,000 words.Our simplified corpus is a collection of complex sentences from Japanese textbooks and reference books together with simplified sentences generated by humans, paired with data on how the complex sentences were paraphrased.The corpus contains a total of 15,000 sentences, in both complex and simple versions.In addition, we investigate the differences in the simplification operations used by each annotator.The aim is to understand whether a crowdsourced complex-simple parallel corpus is an appropriate data source for automated simplification by machine learning.The results, that there was a high level of agreement between the annotators building the data set.So, we believe that this corpus is a good quality data set for machine learning for simplification.We therefore plan to expand the scale of the simplified corpus in the future.

主題

この書誌の出所

  • openalex— W2806010723(2026-08-12取得)

引用

Akihiro Katsuta・Kazuhide Yamamoto(2018-05-01) Crowdsourced Corpus of Sentence Simplification with Core Vocabulary

KatsutaYamamoto2018CrowdsourcedCorpusSentence
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON