本文へ移動

論文 ·日本語 ·未確認

Building a Japanese Typo Dataset from Wikipedia’s Revision History

Yu Tanaka Yugo Murawaki Daisuke Kawahara Sadao Kurohashi

刊行年
2020-01-01
言語
英語
OpenAlex
W3037128534
DOI
10.18653/v1/2020.acl-srw.31
MAG
3037128534
URL
https://doi.org/10.18653/v1/2020.acl-srw.31

要旨

User generated texts contain many typos for which correction is necessary for NLP systems to work. Although a large number of typo–correction pairs are needed to develop a data-driven typo correction system, no such dataset is available for Japanese. In this paper, we extract over half a million Japanese typo–correction pairs from Wikipedia’s revision history. Unlike other languages, Japanese poses unique challenges: (1) Japanese texts are unsegmented so that we cannot simply apply a spelling checker, and (2) the way people inputting kanji logographs results in typos with drastically different surface forms from correct ones. We address them by combining character-based extraction rules, morphological analyzers to guess readings, and various filtering methods. We evaluate the dataset using crowdsourcing and run a baseline seq2seq model for typo correction.

主題

この書誌の出所

  • openalex— W3037128534(2026-08-14取得)

引用

Yu Tanaka・Yugo Murawaki・Daisuke Kawahara・Sadao Kurohashi(2020-01-01) Building a Japanese Typo Dataset from Wikipedia’s Revision History pp. 230-236

Tanaka2020BuildingJapaneseTypo
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON