論文 ·日本語 ·未確認
Machine Learning Approaches to the Exhaustive Auto-Identification of Japanese Clauses
- 刊行年
- 2024-12-13
- 言語
- 英語
- openalex
- W4409403731
- doi
- 10.1145/3711542.3711585
- URL
- https://dl.acm.org/doi/pdf/10.1145/3711542.3711585
要旨
Automatic clause identification is important for both Natural Language Processing and Second Language Writing Research. In order to explore the optimal model for exhaustive auto-identification of Japanese clauses, this paper trained six distinct machine learning models using the morphological information from the BCCWJ corpus and its clause boundary annotation dataset, BCCWJ-CBL, and evaluated and compared the prediction performance of each model. Results indicated that both the XGBoost and Random Forest models achieved high F1 scores of 0.982 for the lenient metric and 0.976 for the stringent metric in the evaluation, and the two can be considered as the optimal tools for exhaustive auto-identification of Japanese clauses currently available. Additionally, this paper also found that the five types of morphological information (active form, lemma pronunciation, macro-categorization of parts of speech (POS), original form, and meso-categorization of POS) of current words are the most important features influencing the prediction performance of these models.
主題
この書誌の出所
- openalex— W4409403731(2026-08-12取得)
引用キー: Zhong2024MachineLearningApproaches