プレプリント ·日本語 ·未確認

JaQuAD: Japanese Question Answering Dataset for Machine Reading Comprehension

ByungHoon So Kyuhong Byun Kyung‐Won Kang Seong Jin Cho

刊行年
2022-02-03
収録
『arXiv (Cornell University)』
出版
Cornell University
言語
英語
openalex
W4221152489
doi
10.48550/arxiv.2202.01764
issn
2331-8422
URL
https://arxiv.org/pdf/2202.01764

要旨

Question Answering (QA) is a task in which a machine understands a given document and a question to find an answer. Despite impressive progress in the NLP area, QA is still a challenging problem, especially for non-English languages due to the lack of annotated datasets. In this paper, we present the Japanese Question Answering Dataset, JaQuAD, which is annotated by humans. JaQuAD consists of 39,696 extractive question-answer pairs on Japanese Wikipedia articles. We finetuned a baseline model which achieves 78.92% for F1 score and 63.38% for EM on test set. The dataset and our experiments are available at https://github.com/SkelterLabsInc/JaQuAD.

主題

この書誌の出所

  • openalex— W4221152489(2026-08-12取得)

引用キー: So2022JaQuAD

書誌 34,896件 語別索引 17,251件 資源 113件 研究者 303名 JSON