論文 ·日本語 ·未確認
Mining Revision Log of Language Learning SNS for Automated Japanese Error Correction
Tomoya Mizumoto ・ Mamoru Komachi ・ Masaaki Nagata ・ Yūji Matsumoto
- 刊行年
- 2013-01-01
- 収録
- 『Transactions of the Japanese Society for Artificial Intelligence』 28(5) pp. 420-432
- 出版
- The Japanese Society for Artificial Intelligence
- 言語
- 英語
- openalex
- W2334823724
- doi
- 10.1527/tjsai.28.420
- mag
- 2334823724
- issn
- 1346-0714
- URL
- https://www.jstage.jst.go.jp/article/tjsai/28/5/28_B-C76/_pdf
要旨
Recently, natural language processing research has begun to pay attention to second language learning. However, it is not easy to acquire a large-scale learners' corpus, which is important for a research for second language learning by natural language processing. We present an attempt to extract a large-scale Japanese learners' corpus from the revision log of a language learning social network service.This corpus is easy to obtain in large-scale, covers a wide variety of topics and styles, and can be a great source of knowledge for both language learners and instructors. We also demonstrate that the extracted learners' corpus of Japanese as a second language can be used as training data for learners' error correction using a statistical machine translation approach.We evaluate different granularities of tokenization to alleviate the problem of word segmentation errors caused by erroneous input from language learners.We propose a character-based SMT approach to alleviate the problem of erroneous input from language learners.Experimental results show that the character-based model outperforms the word-based model when corpus size is small and test data is written by the learners whose L1 is English.
主題
この書誌の出所
- openalex— W2334823724(2026-08-13取得)
引用キー: Mizumoto2013MiningRevisionLog