本文へ移動

プレプリント ·用例に日本語 ·未確認

Character Feature Engineering for Japanese Word Segmentation

Mike Tian-Jian Jiang

刊行年
2019-10-03
収録
『arXiv (Cornell University)』
出版
Cornell University
言語
英語
OpenAlex
W2977894884
DOI
10.48550/arxiv.1910.01761
MAG
2977894884
ISSN
2331-8422
URL
https://arxiv.org/pdf/1910.01761

要旨

On word segmentation problems, machine learning architecture engineering often draws attention. The problem representation itself, however, has remained almost static as either word lattice ranking or character sequence tagging, for at least two decades. The latter of-ten shows stronger predictive power than the former for out-of-vocabulary (OOV) issue. When the issue escalating to rapid adaptation, which is a common scenario for industrial applications, active learning of partial annotations or re-training with additional lexical re-sources is usually applied, however, from a somewhat word-based perspective. Not only it is uneasy for end-users to comply with linguistically consistent word boundary decisions, but also the risk/cost of forking models permanently with estimated weights is seldom affordable. To overcome the obstacle, this work provides an alternative, which uses linguistic intuition about character compositions, such that a sophisticated feature set and its derived scheme can enable dynamic lexicon expansion with the model remaining intact. Experiment results suggest that the proposed solution, with or without external lexemes, performs competitively in terms of F1 score and OOV recall across various datasets.

主題

この書誌の出所

  • openalex— W2977894884(2026-08-14取得)

引用

Mike Tian-Jian Jiang(2019-10-03) Character Feature Engineering for Japanese Word Segmentation 『arXiv (Cornell University)』 Cornell University

Jiang2019CharacterFeatureEngineering
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON