Minimizing the pretraining gap: Domain-aligned text-based person retrieval
Authors:
Shuyu Yang,
Yaxiong Wang, Yongrui Li, Li Zhu, Zhedong Zheng
Published in Pattern Recognition (PR), 2026
Recommended citation: Shuyu Yang, Yaxiong Wang, Yongrui Li, Li Zhu, Zhedong Zheng, "Minimizing the pretraining gap: Domain-aligned text-based person retrieval." PR, 2026. DOI: 10.1016/j.patcog.2026.113511
Download PDF: https://zdzheng.xyz/files/2026/PR_SDA_Yang.pdf
中文解读: https://zhuanlan.zhihu.com/p/2052778421856506671
Code is available at: https://github.com/Shuyu-XJTU/MRA
Abstract: In this work, we focus on text-based person retrieval, which identifies individuals based on textual descriptions. Despite advancements enabled by synthetic data for pretraining, a significant domain gap, due to variations in lighting, color, and viewpoint, limits the effectiveness of the pretrain-finetune paradigm. To overcome this issue, we propose a unified pipeline incorporating domain adaptation at both image and region levels. Our method features two key components: Domain-aware Diffusion (DaD) for image-level adaptation, which aligns image distributions between synthetic and real-world domains, e.g., CUHK-PEDES, and Multi-granularity Relation Alignment (MRA) for region-level adaptation, which aligns visual regions with descriptive sentences, thereby addressing disparities at a finer granularity. This dual-level strategy effectively bridges the domain gap, achieving state-of-the-art performance on CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets. The dataset, model, and code are available at https://github.com/Shuyu-XJTU/MRA.
@article{yang2026minimizing,
author = "Yang, Shuyu and Wang, Yaxiong and Li, Yongrui and Zhu, Li and Zheng, Zhedong",
title = "Minimizing the pretraining gap: Domain-aligned text-based person retrieval",
abstract = "In this work, we focus on text-based person retrieval, which identifies individuals based on textual descriptions. Despite advancements enabled by synthetic data for pretraining, a significant domain gap, due to variations in lighting, color, and viewpoint, limits the effectiveness of the pretrain-finetune paradigm. To overcome this issue, we propose a unified pipeline incorporating domain adaptation at both image and region levels. Our method features two key components: Domain-aware Diffusion (DaD) for image-level adaptation, which aligns image distributions between synthetic and real-world domains, e.g., CUHK-PEDES, and Multi-granularity Relation Alignment (MRA) for region-level adaptation, which aligns visual regions with descriptive sentences, thereby addressing disparities at a finer granularity. This dual-level strategy effectively bridges the domain gap, achieving state-of-the-art performance on CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets. The dataset, model, and code are available at https://github.com/Shuyu-XJTU/MRA.",
url = "https://zdzheng.xyz/files/2026/PR\_SDA\_Yang.pdf",
code = "https://github.com/Shuyu-XJTU/MRA",
blog = "https://zhuanlan.zhihu.com/p/2052778421856506671",
doi = "10.1016/j.patcog.2026.113511",
funding = "2025A1515012281, MYRG-GRG2024-00077-FST-UMDF, University of Macau Advanced Research Institute in Hengqin, FDCT/0043/2025/RIA1",
journal = "PR",
year = "2026" }