論文 ·用例に日本語 ·未確認
Towards automated creation of high quality domain-specific machine translation resources
- 刊行年
- 2015-12-11
- 言語
- 英語
- OpenAlex
- W2325535771
- DOI
- 10.1145/2837185.2837273
- MAG
- 2325535771
- URL
- https://doi.org/10.1145/2837185.2837273
要旨
In this paper we present a workflow for the automated creation of parallel domain-specific corpora, i.e. multilingual translated text collections of a certain domain, in which the text pieces are aligned at sentence level. The source for the text extraction are Wikipedia articles. This workflow will be adaptable to any language pair, though the first implementation is targeting English/Japanese. The workflow consists of intelligent text acquisition, text alignment including novel techniques, and large scale quality evaluation by human experts. This will enable us to create an adaptable, fine-tuned system as well as high quality corpora, which will be compiled during the implementation.
主題
この書誌の出所
- openalex— W2325535771(2026-08-14取得)
引用
Bartholomäus Wloka(2015-12-11) Towards automated creation of high quality domain-specific machine translation resources pp. 1-4