論文 ·日本語 ·未確認
Building a Japanese Typo Dataset from Wikipedia’s Revision History
Yu Tanaka ・ Yugo Murawaki ・ Daisuke Kawahara ・ Sadao Kurohashi
- 刊行年
- 2020-01-01
- 言語
- 英語
- OpenAlex
- W3037128534
- DOI
- 10.18653/v1/2020.acl-srw.31
- MAG
- 3037128534
- URL
- https://doi.org/10.18653/v1/2020.acl-srw.31
要旨
User generated texts contain many typos for which correction is necessary for NLP systems to work. Although a large number of typo–correction pairs are needed to develop a data-driven typo correction system, no such dataset is available for Japanese. In this paper, we extract over half a million Japanese typo–correction pairs from Wikipedia’s revision history. Unlike other languages, Japanese poses unique challenges: (1) Japanese texts are unsegmented so that we cannot simply apply a spelling checker, and (2) the way people inputting kanji logographs results in typos with drastically different surface forms from correct ones. We address them by combining character-based extraction rules, morphological analyzers to guess readings, and various filtering methods. We evaluate the dataset using crowdsourcing and run a baseline seq2seq model for typo correction.
主題
この書誌の出所
- openalex— W3037128534(2026-08-14取得)
引用
Yu Tanaka・Yugo Murawaki・Daisuke Kawahara・Sadao Kurohashi(2020-01-01) Building a Japanese Typo Dataset from Wikipedia’s Revision History pp. 230-236