本文へ移動

論文 ·用例に日本語 ·未確認

Bilingual Spoken Monologue Corpus for Simultaneous Machine Interpretation Research

Shigeki Matsubara Akira Takagi Nobuo Kawaguchi Yasuyoshi Inagaki

刊行年
2002-05-01
言語
英語
OpenAlex
W365755509
DOI
10.63317/2airahzn8k72
MAG
365755509
URL
http://www.lrec-conf.org/proceedings/lrec2002/pdf/273.pdf

要旨

This paper describes a large-scale bilingual corpus of spoken monologues and their simultaneous interpretation, which has been constructed at CIAIR.The corpus has the following characteristics: (1) English and Japanese speeches are recorded in parallel, (2) the data contains monologue speeches such as lecture and self-introduction, and (3) the exact beginning and ending times are provided for each utterance.We have collected a total of about 70 hours of speech data and transcribed them into ASCII text files.The corpus will be made publicly available in the near future.This paper also provides an analysis of the professional interpreter's speeches using the bilingual corpus.The following points have been investigated: (1) the interpreting unit of simultaneous interpretation, (2) the difference between the beginning time of the lecturer's utterance and that of the interpreter's utterance, and (3) the interpreter's speaking speed.The characteristic features about the timing at which simultaneous interpreters start to speak is presented.The analysis will be available for the development of a simultaneous machine interpreting system.

主題

この書誌の出所

  • openalex— W365755509(2026-08-14取得)

引用

Shigeki Matsubara・Akira Takagi・Nobuo Kawaguchi・Yasuyoshi Inagaki(2002-05-01) Bilingual Spoken Monologue Corpus for Simultaneous Machine Interpretation Research

Matsubara2002BilingualSpokenMonologue
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON