本文へ移動

論文 ·対照・比較 ·未確認

WCC-JC: A Web-Crawled Corpus for Japanese-Chinese Neural Machine Translation

Jinyi Zhang Ye Tian Jiannan Mao Mei Han Tadahiro Matsumoto

刊行年
2022-06-13
収録
『Applied Sciences』 12(12) pp. 6002-6002
出版
Multidisciplinary Digital Publishing Institute
言語
英語
OpenAlex
W4282828036
DOI
10.3390/app12126002
ISSN
2076-3417
URL
https://www.mdpi.com/2076-3417/12/12/6002/pdf?version=1655455139

要旨

Currently, there are only a limited number of Japanese-Chinese bilingual corpora of a sufficient amount that can be used as training data for neural machine translation (NMT). In particular, there are few corpora that include spoken language such as daily conversation. In this research, we attempt to construct a Japanese-Chinese bilingual corpus of a certain scale by crawling the subtitle data of movies and TV series from the websites. We calculated the BLEU scores of the constructed WCC-JC (Web Crawled Corpus—Japanese and Chinese) and the other compared corpora. We also manually evaluated the translation results using the translation model trained on the WCC-JC to confirm the quality and effectiveness.

主題

この書誌の出所

  • openalex— W4282828036(2026-08-14取得)

引用

Jinyi Zhang・Ye Tian・Jiannan Mao・Mei Han・Tadahiro Matsumoto(2022-06-13) WCC-JC: A Web-Crawled Corpus for Japanese-Chinese Neural Machine Translation 『Applied Sciences』 12(12) pp. 6002-6002 Multidisciplinary Digital Publishing Institute

Zhang2022WCCJC
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON