本文へ移動

論文 ·日本語 ·未確認

Robust segmentation of Japanese text into a lattice for parsing

Gary Kacmarcik Chris Brockett Hisami Suzuki

刊行年
2000-01-01
言語
英語
OpenAlex
W1999741130
DOI
10.3115/990820.990877
MAG
1999741130
URL
https://dl.acm.org/doi/pdf/10.3115/990820.990877

要旨

We describe a segmentation component that utilizes minimal syntactic knowledge to produce a lattice of word candidates for a broad coverage Japanese NL parser. The segmenter is a finite state morphological analyzer and text normalizer designed to handle the orthographic variations characteristic of written Japanese, including alternate spellings, script variation, vowel extensions and word-internal parenthetical material. This architecture differs from conventional Japanese wordbreakers in that it does not attempt to simultaneously attack the problems of identifying segmentation candidates and choosing the most probable analysis. To minimize duplication of effort between components and to give the segmenter greater freedom to address orthography issues, the task of choosing the best analysis is handled by the parser, which has access to a much richer set of linguistic information. By maximizing recall in the segmenter and allowing a precision of 34.7%, our parser currently achieves a breaking accuracy of ~97% over a wide variety of corpora.

主題

この書誌の出所

  • openalex— W1999741130(2026-08-14取得)

引用

Gary Kacmarcik・Chris Brockett・Hisami Suzuki(2000-01-01) Robust segmentation of Japanese text into a lattice for parsing 1 pp. 390-396

Kacmarcik2000RobustSegmentationJapanese
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON