本文へ移動

論文 ·用例に日本語 ·未確認

Towards automated creation of high quality domain-specific machine translation resources

Bartholomäus Wloka

刊行年
2015-12-11
言語
英語
OpenAlex
W2325535771
DOI
10.1145/2837185.2837273
MAG
2325535771
URL
https://doi.org/10.1145/2837185.2837273

要旨

In this paper we present a workflow for the automated creation of parallel domain-specific corpora, i.e. multilingual translated text collections of a certain domain, in which the text pieces are aligned at sentence level. The source for the text extraction are Wikipedia articles. This workflow will be adaptable to any language pair, though the first implementation is targeting English/Japanese. The workflow consists of intelligent text acquisition, text alignment including novel techniques, and large scale quality evaluation by human experts. This will enable us to create an adaptable, fine-tuned system as well as high quality corpora, which will be compiled during the implementation.

主題

この書誌の出所

  • openalex— W2325535771(2026-08-14取得)

引用

Bartholomäus Wloka(2015-12-11) Towards automated creation of high quality domain-specific machine translation resources pp. 1-4

Wloka2015AutomatedCreationHigh
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON