本文へ移動

論文 ·対照・比較 ·未確認

Compilation of specialized comparable corpora in French and Japanese

Lorraine Goeuriot Emmanuel Morin Béatrice Daille

刊行年
2009-01-01
言語
英語
OpenAlex
W2157668051
DOI
10.3115/1690339.1690353
MAG
2157668051
URL
https://dl.acm.org/doi/pdf/10.5555/1690339.1690353

要旨

We present in this paper the development of a specialized comparable corpora compilation tool, for which quality would be close to a manually compiled corpus. The comparability is based on three levels: domain, topic and type of discourse. Domain and topic can be filtered with the keywords used through web search. But the detection of the type of discourse needs a wide linguistic analysis. The first step of our work is to automate the detection of the type of discourse that can be found in a scientific domain (science and popular science) in French and Japanese languages. First, a contrastive stylistic analysis of the two types of discourse is done on both languages. This analysis leads to the creation of a reusable, generic and robust typology. Machine learning algorithms are then applied to the typology, using shallow parsing. We obtain good results, with an average precision of 80% and an average recall of 70% that demonstrate the efficiency of this typology. This classification tool is then inserted in a corpus compilation tool which is a text collection treatment chain realized through IBM UIMA system. Starting from two specialized web documents collection in French and Japanese, this tool creates the corresponding corpus.

主題

この書誌の出所

  • openalex— W2157668051(2026-08-14取得)

引用

Lorraine Goeuriot・Emmanuel Morin・Béatrice Daille(2009-01-01) Compilation of specialized comparable corpora in French and Japanese pp. 55-55

Goeuriot2009CompilationSpecializedComparable
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON