Align and tell: Boosting text-video retrieval with local alignment and fine-grained supervision
Authors:
Xiaohan Wang,
Linchao Zhu, Zhedong Zheng, Mingliang Xu,
Yi Yang
Published in IEEE Transactions on Multimedia (TMM), 2022
Recommended citation: Xiaohan Wang, Linchao Zhu, Zhedong Zheng, Mingliang Xu, Yi Yang, "Align and tell: Boosting text-video retrieval with local alignment and fine-grained supervision." IEEE Transactions on Multimedia, 2022. DOI: 10.1109/TMM.2022.3204444
Download PDF: https://zdzheng.xyz/files/2022/TMM22-Xiaohan.pdf
Abstract: Text-video retrieval is one of the basic tasks for multimodal research and has been widely harnessed in many real- world systems. Most existing approaches directly compare the global representation between videos and text descriptions and utilize the global contrastive loss to train the model. These designs overlook the local alignment and the word-level supervision signal. In this paper, we propose a new framework, called Align and Tell, for text-video retrieval. Compared to the previous work, our framework contains additional modules, i.e., two transformer decoders for local alignment and one captioning head to enhance the representation learning. First, we introduce a set of learnable queries to interact with both textual representations and video representations and project them to a fixed number of local features. After that, local contrastive learning is performed to complement the global comparison. Moreover, we design a video captioning head to provide additional supervision signals during training. This word-level supervision can enhance the visual presentation and alleviate the cross-modal gap. The captioning head can be removed during inference and does not introduce extra computational costs. Extensive empirical results demon- strate that our Align and Tell model can achieve state-of-the- art performance on four text-video retrieval datasets, including MSR-VTT, MSVD, LSMDC, and ActivityNet-Captions.
@article{wang2022align,
author = "Wang, Xiaohan and Zhu, Linchao and Zheng, Zhedong and Xu, Mingliang and Yang, Yi",
title = "Align and tell: Boosting text-video retrieval with local alignment and fine-grained supervision",
abstract = "Text-video retrieval is one of the basic tasks for multimodal research and has been widely harnessed in many real- world systems. Most existing approaches directly compare the global representation between videos and text descriptions and utilize the global contrastive loss to train the model. These designs overlook the local alignment and the word-level supervision signal. In this paper, we propose a new framework, called Align and Tell, for text-video retrieval. Compared to the previous work, our framework contains additional modules, i.e., two transformer decoders for local alignment and one captioning head to enhance the representation learning. First, we introduce a set of learnable queries to interact with both textual representations and video representations and project them to a fixed number of local features. After that, local contrastive learning is performed to complement the global comparison. Moreover, we design a video captioning head to provide additional supervision signals during training. This word-level supervision can enhance the visual presentation and alleviate the cross-modal gap. The captioning head can be removed during inference and does not introduce extra computational costs. Extensive empirical results demon- strate that our Align and Tell model can achieve state-of-the- art performance on four text-video retrieval datasets, including MSR-VTT, MSVD, LSMDC, and ActivityNet-Captions.",
journal = "IEEE Transactions on Multimedia",
url = "https://zdzheng.xyz/files/2022/TMM22-Xiaohan.pdf",
doi = "10.1109/TMM.2022.3204444",
year = "2022",
publisher = "IEEE" }