本文へ移動

論文 ·用例に日本語 ·未確認

DueT: Image-Text Contrastive Transfer Learning with Dual-adapter Tuning

Taku Hasegawa Kyosuke Nishida Koki Maeda Kuniko Saito

刊行年
2023-01-01
言語
英語
OpenAlex
W4389524425
DOI
10.18653/v1/2023.emnlp-main.839
URL
https://aclanthology.org/2023.emnlp-main.839.pdf

要旨

This paper presents DueT, a novel transfer learning method for vision and language models built by contrastive learning. In DueT, adapters are inserted into the image and text encoders, which have been initialized using models pre-trained on uni-modal corpora and then frozen. By training only these adapters, DueT enables efficient learning with a reduced number of trainable parameters. Moreover, unlike traditional adapters, those in DueT are equipped with a gating mechanism, enabling effective transfer and connection of knowledge acquired from pre-trained uni-modal encoders while preventing catastrophic forgetting. We report that DueT outperformed simple fine-tuning, the conventional method fixing only the image encoder and training only the text encoder, and the LoRA-based adapter method in accuracy and parameter efficiency for 0-shot image and text retrieval in both English and Japanese domains.

主題

この書誌の出所

  • openalex— W4389524425(2026-08-14取得)

引用

Taku Hasegawa・Kyosuke Nishida・Koki Maeda・Kuniko Saito(2023-01-01) DueT: Image-Text Contrastive Transfer Learning with Dual-adapter Tuning pp. 13607-13624

Hasegawa2023DueT
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON