Fetching the paper…
Reading the bibliography…
TIReID aims to retrieve the image corresponding to the given text query from a pool of candidate images.
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in IEEE Computer Society Conference on Computer Vision and Pattern Recognition, CVPR , 2009
2009
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition, CVPR , 2016
2016
Earlier work this paper cites.
L. Zhang, X. Wang, D. V. Kalashnikov, S. Mehrotra, and D. Ramanan, “Query-driven approach to face clustering and tagging,” IEEE Transactions on Image Processing , vol. 25, no. 10, pp. 4504–4513, 2016
2016
Earlier work this paper cites.
S. Li, T. Xiao, H. Li, B. Zhou, D. Yue, and X. Wang, “Person search with natural language description,” in IEEE Conference on Computer Vision and Pattern Recognition, CVPR , 2017
2017
Earlier work this paper cites.
Y. Sun, L. Zheng, Y. Yang, Q. Tian, and S. Wang, “Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline),” in European Conference on Computer Vision, ECCV , 2018
2018
Earlier work this paper cites.
K. Lee, X. Chen, G. Hua, H. Hu, and X. He, “Stacked cross attention for image-text matching,” in European Conference on Computer Vision, ECCV , 2018
2018
Earlier work this paper cites.
Y. Zhang and H. Lu, “Deep cross-modal projection learning for image-text matching,” in European Conference on Computer Vision, ECCV , 2018
2018
Earlier work this paper cites.
L. Wei, S. Zhang, W. Gao, and Q. Tian, “Person transfer GAN to bridge domain gap for person re-identification,” in IEEE Conference on Computer Vision and Pattern Recognition, CVPR , 2018
2018
Earlier work this paper cites.
D. Chen, H. Li, X. Liu, Y. Shen, J. Shao, Z. Yuan, and X. Wang, “Improving deep visual representation for person re-identification by global and local image-language association,” in European Conference on Computer Vision, ECCV , 2018
2018
Earlier work this paper cites.
H. Li, S. Yan, Z. Yu, and D. Tao, “Attribute-identity embedding and self-supervised learning for scalable person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 30, no. 10, pp. 3472–3485, 2019
2019
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT , 2019
2019
Earlier work this paper cites.
T. Qiao, J. Zhang, D. Xu, and D. Tao, “Mirrorgan: Learning text-to-image generation by redescription,” in IEEE Conference on Computer Vision and Pattern Recognition, CVPR , 2019
2019
Earlier work this paper cites.
Y. Wang, C. Bo, D. Wang, S. Wang, Y. Qi, and H. Lu, “Language person search with mutually connected classification loss,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP , 2019
2019
Earlier work this paper cites.
J. Liu, Z. Zha, R. Hong, M. Wang, and Y. Zhang, “Deep adversarial graph attention convolution network for text-based person search,” in 27th ACM International Conference on Multimedia, MM , 2019
2019
Earlier work this paper cites.
N. Sarafianos, X. Xu, and I. A. Kakadiaris, “Adversarial representation learning for text-to-image matching,” in IEEE/CVF International Conference on Computer Vision, ICCV , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Z. Wang, Z. Fang, J. Wang, and Y. Yang, “Vitaa: Visual-textual attributes alignment in person search by natural language,” in European Conference on Computer Vision, ECCV , 2020
2020
Earlier work this paper cites.
S. Aggarwal, R. V. Babu, and A. Chakraborty, “Text-based person search via attribute-aided matching,” in IEEE Winter Conference on Applications of Computer Vision, WACV , 2020
2020
Earlier work this paper cites.
K. Zheng, W. Liu, J. Liu, Z. Zha, and T. Mei, “Hierarchical gumbel attention network for text-based person search,” in 28th ACM International Conference on Multimedia, MM , 2020
2020
Earlier work this paper cites.
K. Niu, Y. Huang, W. Ouyang, and L. Wang, “Improving description-based person re-identification by multi-granularity image-text alignments,” IEEE Transactions on Image Processing , vol. 29, pp. 5542–5556, 2020
2020
Earlier work this paper cites.
Y. Jing, C. Si, J. Wang, W. Wang, L. Wang, and T. Tan, “Pose-guided multi-granularity attention network for text-based person search,” in Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI , 2020
2020
Earlier work this paper cites.
Z. Zheng, L. Zheng, M. Garrett, Y. Yang, M. Xu, and Y. Shen, “Dual-path convolutional image-text embeddings with instance loss,” ACM Transactions on Multimedia Computing, Communications, and Applications , vol. 16, no. 2, pp. 51:1–51:23, 2020
2020
Earlier work this paper cites.
Y. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “UNITER: universal image-text representation learning,” in European Conference on Computer Vision, ECCV , 2020
2020
Cited alongside, same era.
H. Tang, Z. Li, Z. Peng, and J. Tang, “Blockmix: Meta regularization and self-calibrated inference for metric-based meta-learning,” in 28th ACM International Conference on Multimedia, MM , 2020
2020
Cited alongside, same era.
K. Niu, Y. Huang, and L. Wang, “Textual dependency embedding for person search by language,” in 28th ACM International Conference on Multimedia, MM , 2020
2020
Cited alongside, same era.
Z. Wang, A. Zhu, Z. Zheng, J. Jin, Z. Xue, and G. Hua, “Img-net: inner-cross-modal attentional multigranular network for description-based person re-identification,” Journal of Electronic Imaging , vol. 29, no. 4, p. 043028, 2020
2020
Cited alongside, same era.
H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, and T. Li, “Clip4clip: An empirical study of CLIP for end to end video clip retrieval,” Neurocomputing , vol. 508, pp. 293–304, 2022
2022
Closest in time.
Z. Wang, Y. Lu, Q. Li, X. Tao, Y. Guo, M. Gong, and T. Liu, “CRIS: clip-driven referring image segmentation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR , 2022
2022
Closest in time.
Y. Rao, W. Zhao, G. Chen, Y. Tang, Z. Zhu, G. Huang, J. Zhou, and J. Lu, “Denseclip: Language-guided dense prediction with context-aware prompting,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR , 2022
2022
Closest in time.
B. Ni, H. Peng, M. Chen, S. Zhang, G. Meng, J. Fu, S. Xiang, and H. Ling, “Expanding language-image pretrained models for general video recognition,” in European Conference on Computer Vision, ECCV , 2022
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Zhang, G. Du, F. Liu, H. Tu, and X. Shu, “Global-local multiple granularity learning for cross-modality visible-infrared person reidentification,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–11, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
A. Zhu, Z. Wang, Y. Li, X. Wan, J. Jin, T. Wang, F. Hu, and G. Hua, “DSSL: deep surroundings-person separation learning for text-based person retrieval,” in 29th ACM International Conference on Multimedia, MM , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in 9th International Conference on Learning Representations, ICLR , 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in 38th International Conference on Machine Learning, ICML , 2021
2021
Cited alongside, same era.
C. Wang, Z. Luo, Y. Lin, and S. Li, “Text-based person search via multi-granularity embedding learning,” in Thirtieth International Joint Conference on Artificial Intelligence, IJCAI , 2021
2021
Cited alongside, same era.
Y. Chen, R. Huang, H. Chang, C. Tan, T. Xue, and B. Ma, “Cross-modal knowledge adaptation for language-based person search,” IEEE Transactions on Image Processing , vol. 30, pp. 4057–4069, 2021
2021
Cited alongside, same era.
S. Li, M. Cao, and M. Zhang, “Learning semantic-aligned feature representation for text-based person search,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP , 2022
2022
Closest in time.
Z. Shao, X. Zhang, M. Fang, Z. Lin, J. Wang, and C. Ding, “Learning granularity-unified representations for text-to-image person re-identification,” in 30th ACM International Conference on Multimedia, MM , 2022
2022
Closest in time.
L. Yao, R. Huang, L. Hou, G. Lu, M. Niu, H. Xu, X. Liang, Z. Li, X. Jiang, and C. Xu, “FILIP: fine-grained interactive language-image pre-training,” in Tenth International Conference on Learning Representations, ICLR , 2022
2022
Closest in time.
M. Cao, T. Yang, J. Weng, C. Zhang, J. Wang, and Y. Zou, “Locvtp: Video-text pre-training for temporal localization,” in European conference on computer vision, ECCV , 2022
2022
Closest in time.
Y. Ma, G. Xu, X. Sun, M. Yan, J. Zhang, and R. Ji, “X-CLIP: end-to-end multi-grained contrastive learning for video-text retrieval,” in 30th ACM International Conference on Multimedia, MM , 2022
2022
Closest in time.
S. Zhao, L. Zhu, X. Wang, and Y. Yang, “Centerclip: Token clustering for efficient text-video retrieval,” in 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR , 2022
2022
Closest in time.
R. Zhang, Z. Guo, W. Zhang, K. Li, X. Miao, B. Cui, Y. Qiao, P. Gao, and H. Li, “Pointclip: Point cloud understanding by CLIP,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR , 2022
2022
Closest in time.
H. Tang, C. Yuan, Z. Li, and J. Tang, “Learning attention-guided pyramidal features for few-shot fine-grained recognition,” Pattern Recognition , vol. 130, p. 108792, 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
H. Zhu, W. Ke, D. Li, J. Liu, L. Tian, and Y. Shan, “Dual cross-attention learning for fine-grained visual categorization and object re-identification,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR , 2022
2022
Closest in time.
Y. Liu, P. Xiong, L. Xu, S. Cao, and Q. Jin, “Ts2-net: Token shift and selection transformer for text-video retrieval,” in European Conference on Computer Vision, ECCV , 2022
2022
Closest in time.
Z. Wang, A. Zhu, J. Xue, D. Jiang, C. Liu, Y. Li, and F. Hu, “SUM: serialized updating and matching for text-based person retrieval,” Knowledge-Based Systems , vol. 248, p. 108891, 2022
2022
Closest in time.
X. Shu, W. Wen, H. Wu, K. Chen, Y. Song, R. Qiao, B. Ren, and X. Wang, “See finer, see more: Implicit modality alignment for text-based person retrieval,” in European Conference on Computer Vision Workshop on Real-World Surveillance, ECCVW , 2022
2022
Closest in time.
Z. Wang, A. Zhu, J. Xue, X. Wan, C. Liu, T. Wang, and Y. Li, “Look before you leap: Improving text-based person retrieval by learning a consistent cross-modal common manifold,” in 30th ACM International Conference on Multimedia, MM , 2022
2022
Closest in time.
Z. Wang, A. Zhu, J. Xue, X. Wan, C. Liu, T. Wang, and Y. Li, “Caibc: Capturing all-round information beyond color for text-based person retrieval,” in 30th ACM International Conference on Multimedia, MM , 2022
2022
Closest in time.
A. Farooq, M. Awais, J. Kittler, and S. S. Khalid, “Axm-net: Implicit cross-modal feature alignment for person re-identification,” in Thirty-Sixth AAAI Conference on Artificial Intelligence, AAAI , 2022
2022
Closest in time.