本文へ移動

プレプリント ·用例に日本語 ·未確認

Coursera Corpus Mining and Multistage Fine-Tuning for Improving Lectures Translation

Haiyue Song Raj Dabre Atsushi Fujita Sadao Kurohashi

刊行年
2019-12-26
収録
『arXiv (Cornell University)』
出版
Cornell University
言語
英語
OpenAlex
W2998300596
DOI
10.48550/arxiv.1912.11739
MAG
2998300596
ISSN
2331-8422
URL
https://arxiv.org/pdf/1912.11739

要旨

Lectures translation is a case of spoken language translation and there is a lack of publicly available parallel corpora for this purpose. To address this, we examine a language independent framework for parallel corpus mining which is a quick and effective way to mine a parallel corpus from publicly available lectures at Coursera. Our approach determines sentence alignments, relying on machine translation and cosine similarity over continuous-space sentence representations. We also show how to use the resulting corpora in a multistage fine-tuning based domain adaptation for high-quality lectures translation. For Japanese--English lectures translation, we extracted parallel data of approximately 40,000 lines and created development and test sets through manual filtering for benchmarking translation performance. We demonstrate that the mined corpus greatly enhances the quality of translation when used in conjunction with out-of-domain parallel corpora via multistage training. This paper also suggests some guidelines to gather and clean corpora, mine parallel sentences, address noise in the mined data, and create high-quality evaluation splits. For the sake of reproducibility, we will release our code for parallel data creation.

主題

この書誌の出所

  • openalex— W2998300596(2026-08-14取得)
  • openalex— W3032491448(2026-08-14取得)

引用

Haiyue Song・Raj Dabre・Atsushi Fujita・Sadao Kurohashi(2019-12-26) Coursera Corpus Mining and Multistage Fine-Tuning for Improving Lectures Translation 『arXiv (Cornell University)』 Cornell University

Song2019CourseraCorpusMining
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON