論文 ·日本語 ·未確認

Machine Learning Approaches to the Exhaustive Auto-Identification of Japanese Clauses

Yong Zhong

刊行年
2024-12-13
言語
英語
openalex
W4409403731
doi
10.1145/3711542.3711585
URL
https://dl.acm.org/doi/pdf/10.1145/3711542.3711585

要旨

Automatic clause identification is important for both Natural Language Processing and Second Language Writing Research. In order to explore the optimal model for exhaustive auto-identification of Japanese clauses, this paper trained six distinct machine learning models using the morphological information from the BCCWJ corpus and its clause boundary annotation dataset, BCCWJ-CBL, and evaluated and compared the prediction performance of each model. Results indicated that both the XGBoost and Random Forest models achieved high F1 scores of 0.982 for the lenient metric and 0.976 for the stringent metric in the evaluation, and the two can be considered as the optimal tools for exhaustive auto-identification of Japanese clauses currently available. Additionally, this paper also found that the five types of morphological information (active form, lemma pronunciation, macro-categorization of parts of speech (POS), original form, and meso-categorization of POS) of current words are the most important features influencing the prediction performance of these models.

主題

この書誌の出所

  • openalex— W4409403731(2026-08-12取得)

引用キー: Zhong2024MachineLearningApproaches

書誌 34,896件 語別索引 17,251件 資源 113件 研究者 303名 JSON