Camstyle: A novel data augmentation method for person re-identification

Authors: zhun-zhongZhun Zhong, liang-zhengLiang Zheng, Zhedong Zheng, shaozi-liShaozi Li, yi-yangYi Yang

Published in IEEE Transactions on Image Processing (TIP), 2019

Recommended citation: Zhun Zhong, Liang Zheng, Zhedong Zheng, Shaozi Li, Yi Yang, "Camstyle: A novel data augmentation method for person re-identification." IEEE Transactions on Image Processing (TIP), 2019. DOI: 10.1109/TIP.2018.2874313
Download PDF: https://zdzheng.xyz/files/2019/TIP-08485427.pdf

Code is available at: https://github.com/zhunzhong07/CamStyle

Abstract: Person re-identification (re-ID) is a cross-camera retrieval task that suffers from image style variations caused by different cameras. The art implicitly addresses this problem by learning a camera-invariant descriptor subspace. In this paper, we explicitly consider this challenge by introducing camera style (CamStyle). CamStyle can serve as a data augmentation approach that reduces the risk of deep network overfitting and that smooths the CamStyle disparities. Specifically, with a style transfer model, labeled training images can be style transferred to each camera, and along with the original training samples, form the augmented training set. This method, while increasing data diversity against overfitting, also incurs a considerable level of noise. In the effort to alleviate the impact of noise, the label smooth regularization (LSR) is adopted. The vanilla version of our method (without LSR) performs reasonably well on few camera systems in which overfitting often occurs. With LSR, we demonstrate consistent improvement in all systems regardless of the extent of overfitting. We also report competitive accuracy compared with the state of the art on Market-1501 and DukeMTMC-re-ID. Importantly, CamStyle can be employed to the challenging problems of one view learning and unsupervised domain adaptation (UDA) in person re-identification (re-ID), both of which have critical research and application significance. The former only has labeled data in one camera view and the latter only has labeled data in the source domain. Experimental results show that CamStyle significantly improves the performance of the baseline in the two problems. Specially, for UDA, CamStyle achieves state-of-the-art accuracy based on a baseline deep re-ID model on Market-1501 and DukeMTMC-reID. Our code is available at: https://github.com/zhunzhon g07/CamStyle. Manuscript received April 22, 2018; revised August 26, 2018 and September 28, 2018; accepted September 28, 2018. Date of publication October 8, 2018; date of current version November 2, 2018. This work was supported in part by the National Nature Science Foundation of China under Grants 61572409, 61876159, 61806172, U1705286, and 61571188, in part by the Fujian Province 2011 Collaborative Innovation Center of TCM Health Management, in part by the Collaborative Innovation Center of Chinese Oolong Tea Industry-Collaborative Innovation Center (2011) of Fujian Province, in part by the Fund for Integration of Cloud Computing and Big Data, in part by the Innovation of Science and Education, in part by the Data to Decisions CRC (D2D CRC), and in part by the Cooperative Research Centre Programme. The asso ciate editor coordinating the review of this manuscript and approving it for publication was Prof. Dong Xu. (Corresponding author: Shaozi Li.) Z. Zhong is with the Cognitive Science D epartment, Xiamen University, Xiamen 361005, China, and also with the C entre for Artificial Intelligence, University of Technology Sydney, Ultimo, NSW 2007, Australia (e-mail: [email protected]). L. Zheng is with the Research School of Computer Science, The Australian National University, C anberra, ACT 0200, Australia (e-mail: [email protected]). Z. Zheng and Y . Yang are with the Centre for Artificial Intelligence, University of Technology Sydney, Ultimo, NSW 2007, Australia (e-mail: [email protected]; [email protected]). S. Li is with the Cognitive Science Department, Xiamen University, Xiamen 361005, China (e-mail: [email protected]). Color versions of one or more of the figures in this paper are available online at http://ieeexplore.ieee.org. Digital Object Identifier 10.1109/TIP.2018.2874313

@article{zhong2019camstyle,
author = "Zhong, Zhun and Zheng, Liang and Zheng, Zhedong and Li, Shaozi and Yang, Yi",
doi = "10.1109/TIP.2018.2874313",
title = "Camstyle: A novel data augmentation method for person re-identification",
abstract = "Person re-identification (re-ID) is a cross-camera retrieval task that suffers from image style variations caused by different cameras. The art implicitly addresses this problem by learning a camera-invariant descriptor subspace. In this paper, we explicitly consider this challenge by introducing camera style (CamStyle). CamStyle can serve as a data augmentation approach that reduces the risk of deep network overfitting and that smooths the CamStyle disparities. Specifically, with a style transfer model, labeled training images can be style transferred to each camera, and along with the original training samples, form the augmented training set. This method, while increasing data diversity against overfitting, also incurs a considerable level of noise. In the effort to alleviate the impact of noise, the label smooth regularization (LSR) is adopted. The vanilla version of our method (without LSR) performs reasonably well on few camera systems in which overfitting often occurs. With LSR, we demonstrate consistent improvement in all systems regardless of the extent of overfitting. We also report competitive accuracy compared with the state of the art on Market-1501 and DukeMTMC-re-ID. Importantly, CamStyle can be employed to the challenging problems of one view learning and unsupervised domain adaptation (UDA) in person re-identification (re-ID), both of which have critical research and application significance. The former only has labeled data in one camera view and the latter only has labeled data in the source domain. Experimental results show that CamStyle significantly improves the performance of the baseline in the two problems. Specially, for UDA, CamStyle achieves state-of-the-art accuracy based on a baseline deep re-ID model on Market-1501 and DukeMTMC-reID. Our code is available at: https://github.com/zhunzhon g07/CamStyle. Manuscript received April 22, 2018; revised August 26, 2018 and September 28, 2018; accepted September 28, 2018. Date of publication October 8, 2018; date of current version November 2, 2018. This work was supported in part by the National Nature Science Foundation of China under Grants 61572409, 61876159, 61806172, U1705286, and 61571188, in part by the Fujian Province 2011 Collaborative Innovation Center of TCM Health Management, in part by the Collaborative Innovation Center of Chinese Oolong Tea Industry-Collaborative Innovation Center (2011) of Fujian Province, in part by the Fund for Integration of Cloud Computing and Big Data, in part by the Innovation of Science and Education, in part by the Data to Decisions CRC (D2D CRC), and in part by the Cooperative Research Centre Programme. The asso ciate editor coordinating the review of this manuscript and approving it for publication was Prof. Dong Xu. (Corresponding author: Shaozi Li.) Z. Zhong is with the Cognitive Science D epartment, Xiamen University, Xiamen 361005, China, and also with the C entre for Artificial Intelligence, University of Technology Sydney, Ultimo, NSW 2007, Australia (e-mail: [email protected]). L. Zheng is with the Research School of Computer Science, The Australian National University, C anberra, ACT 0200, Australia (e-mail: [email protected]). Z. Zheng and Y . Yang are with the Centre for Artificial Intelligence, University of Technology Sydney, Ultimo, NSW 2007, Australia (e-mail: [email protected]; [email protected]). S. Li is with the Cognitive Science Department, Xiamen University, Xiamen 361005, China (e-mail: [email protected]). Color versions of one or more of the figures in this paper are available online at http://ieeexplore.ieee.org. Digital Object Identifier 10.1109/TIP.2018.2874313",
journal = "IEEE Transactions on Image Processing (TIP)",
volume = "28",
number = "3",
pages = "1176--1190",
year = "2019",
url = "https://zdzheng.xyz/files/2019/TIP-08485427.pdf",
code = "https://github.com/zhunzhong07/CamStyle",
publisher = "IEEE" }