本文へ移動

論文 ·日本語 ·未確認

An Empirical Study of Contextual Data Augmentation for Japanese Zero Anaphora Resolution

Ryuto Konno ・ Yuichiroh Matsubayashi ・ Shun Kiyono ・ Hiroki Ouchi ・ Ryo Takahashi ・ Kentaro Inui

刊行年
2020-01-01
収録
『arXiv (Cornell University)』 pp. 4956-4968
出版
Cornell University
言語
英語
OpenAlex
W3113524397
DOI
10.18653/v1/2020.coling-main.435
MAG
3113524397
ISSN
2331-8422
URL
https://www.aclweb.org/anthology/2020.coling-main.435.pdf

要旨

One critical issue of zero anaphora resolution (ZAR) is the scarcity of labeled data. This study explores how effectively this problem can be alleviated by data augmentation. We adopt a state-ofthe-art data augmentation method, called the contextual data augmentation (CDA), that generates labeled training instances using a pretrained language model. The CDA has been reported to work well for several other natural language processing tasks, including text classification and machine translation This study addresses two underexplored issues on CDA, that is, how to reduce the computational cost of data augmentation and how to ensure the quality of the generated data. We also propose two methods to adapt CDA to ZAR: [MASK]-based augmentation and linguistically-controlled masking. Consequently, the experimental results on Japanese ZAR show that our methods contribute to both the accuracy gain and the computation cost reduction. Our closer analysis reveals that the proposed method can improve the quality of the augmented training data when compared to the conventional CDA.

主題

この書誌の出所

  • openalex— W3113524397(2026-08-14取得)
  • openalex— W3095514862(2026-08-14取得)

引用

Ryuto Konno・Yuichiroh Matsubayashi・Shun Kiyono・Hiroki Ouchi・Ryo Takahashi・Kentaro Inui(2020-01-01) An Empirical Study of Contextual Data Augmentation for Japanese Zero Anaphora Resolution 『arXiv (Cornell University)』 pp. 4956-4968 Cornell University

Konno2020EmpiricalContextualData
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON