論文 ·日本語 ·未確認

Automatic Assessment of Japanese Text Readability Based on a Textbook Corpus

Satoshi Sato Suguru Matsuyoshi Y. Kondoh

刊行年
2008-05-01
言語
英語
openalex
W63748263
doi
10.63317/33amje3bxi5d
mag
63748263
URL
http://www.lrec-conf.org/proceedings/lrec2008/pdf/165_paper.pdf

要旨

This paper describes a method of readability measurement of Japanese texts based on a newly compiled textbook corpus.The textbook corpus consists of 1,478 sample passages extracted from 127 textbooks of elementary school, junior high school, high school, and university; it is divided into thirteen grade levels and the total size is about a million characters.For a given text passage, the readability measurement method determines the grade level to which the passage is the most similar by using character-unigram models, which are constructed from the textbook corpus.Because this method does not require sentence-boundary analysis and word-boundary analysis, it is applicable to texts that include incomplete sentences and non-regular text fragments.The performance of this method, which is measured by the correlation coefficient, is considerably high (R > 0.9); in case that the length of a text passage is limited in 25 characters, the correlation coefficient is still high (R = 0.83).

主題

この書誌の出所

  • openalex— W63748263(2026-08-13取得)

引用キー: Sato2008AutomaticAssessmentJapanese

書誌 56,535件 語別索引 17,251件 資源 113件 研究者 303名 JSON