Joint Attribute Graph Reasoning and Aggregation for Composed Image Retrieval
Authors: Xiaotong Chen, Na Cai, Zhedong Zheng, Huaxin Pang, Ying Qin, Shikui Wei
Published in IEEE Transactions on Multimedia (TMM), 2025
Recommended citation: Xiaotong Chen, Na Cai, Zhedong Zheng, Huaxin Pang, Ying Qin, Shikui Wei, "Joint Attribute Graph Reasoning and Aggregation for Composed Image Retrieval." TMM, 2025. DOI: 10.1109/TMM.2026.3657153
Abstract: Given a multimodal query consisting of a reference image and a modification text pair, composed image retrieval (CIR) aims to locate a target image of interest in a large corpus. Recent CIR methods usually leverage vision and language pre-trained (VLP) models to enhance retrieval performance by generating semantically aligned multimodal representations. However, many of these methods often neglect the intricate semantic relationships between the elements of the reference image, the modification text and the target image, leading to suboptimal performance. Inspired by the intuitive interaction between nodes in the graph structure, we propose a novel attribute-based CIR architecture, which contains two key components: a Multi-grained Discriminative Attribute Extractor (MDAE) and a Latent Relation-guided Attribute-oriented Aggregator (LRAA). The MDAE is designed to produce independent attribute representations for each element of the query triplet. The LRAA, on the other hand, aims to capture more comprehensive semantic correlations between visual and textual inputs by reasoning over joint attribute relation graphs and adaptively aggregating the updated node features into a unified multimodal query embedding. Extensive experiments on three widely-used datasets show that our model consistently outperforms other competitive methods, yielding +5.94\% Avg-Recall@10, +5.31\% Avg-Recall@50 on FashionIQ, +2.03\% Recall@1 on CIRR and +1.41\% Recall@1 on Shoes respectively.
@article{chen2025joint,
author = "Chen, Xiaotong and Cai, Na and Zheng, Zhedong and Pang, Huaxin and Qin, Ying and Wei, Shikui",
title = "Joint Attribute Graph Reasoning and Aggregation for Composed Image Retrieval",
abstract = "Given a multimodal query consisting of a reference image and a modification text pair, composed image retrieval (CIR) aims to locate a target image of interest in a large corpus. Recent CIR methods usually leverage vision and language pre-trained (VLP) models to enhance retrieval performance by generating semantically aligned multimodal representations. However, many of these methods often neglect the intricate semantic relationships between the elements of the reference image, the modification text and the target image, leading to suboptimal performance. Inspired by the intuitive interaction between nodes in the graph structure, we propose a novel attribute-based CIR architecture, which contains two key components: a Multi-grained Discriminative Attribute Extractor (MDAE) and a Latent Relation-guided Attribute-oriented Aggregator (LRAA). The MDAE is designed to produce independent attribute representations for each element of the query triplet. The LRAA, on the other hand, aims to capture more comprehensive semantic correlations between visual and textual inputs by reasoning over joint attribute relation graphs and adaptively aggregating the updated node features into a unified multimodal query embedding. Extensive experiments on three widely-used datasets show that our model consistently outperforms other competitive methods, yielding +5.94\\% Avg-Recall@10, +5.31\\% Avg-Recall@50 on FashionIQ, +2.03\\% Recall@1 on CIRR and +1.41\\% Recall@1 on Shoes respectively.",
journal = "TMM",
doi = "10.1109/TMM.2026.3657153",
year = "2025" }