Fetching the paper…
Reading the bibliography…
Pretrained vision-language models (VLMs) such as CLIP have shown impressive generalization capability in downstream vision tasks with appropriate text prompts.
D. H. Wolpert, “The lack of a priori distinctions between learning algorithms,” Neural computation , vol. 8, no. 7, pp. 1341–1390, 1996
1996
Earlier work this paper cites.
L. Fei-Fei, R. Fergus, and P. Perona, “Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,” in 2004 conference on computer vision and pattern recognition workshop . IEEE, 2004, pp. 178–178
2004
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing . IEEE, 2008, pp. 722–729
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba, “Sun database: Large-scale scene recognition from abbey to zoo,” in 2010 IEEE computer society conference on computer vision and pattern recognition . IEEE, 2010, pp. 3485–3492
2010
Earlier work this paper cites.
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar, “Cats and dogs,” in 2012 IEEE conference on computer vision and pattern recognition . IEEE, 2012, pp. 3498–3505
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” in Proceedings of the IEEE international conference on computer vision workshops , 2013, pp. 554–561
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
L. Bossard, M. Guillaumin, and L. V. Gool, “Food-101–mining discriminative components with random forests,” in European conference on computer vision . Springer, 2014, pp. 446–461
2014
Earlier work this paper cites.
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi, “Describing textures in the wild,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2014, pp. 3606–3613
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
F. Santambrogio, “ { \{ Euclidean, metric, and Wasserstein } \} gradient flows: an overview,” Bulletin of Mathematical Sciences , vol. 7, no. 1, pp. 87–154, 2017
2017
Earlier work this paper cites.
Q. Liu, “Stein variational gradient descent as gradient flow,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in 2017 IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
H. Caesar, J. Uijlings, and V. Ferrari, “Coco-stuff: Thing and stuff classes in context,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 1209–1218
2018
Earlier work this paper cites.
J. Yu, J. Li, Z. Yu, and Q. Huang, “Multimodal transformer with multi-view visual representation for image captioning,” IEEE transactions on circuits and systems for video technology , vol. 30, no. 12, pp. 4467–4480, 2019
2019
Earlier work this paper cites.
M. Arbel, A. Korba, A. Salim, and A. Gretton, “Maximum mean discrepancy gradient flow,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
A. Liutkus, U. Simsekli, S. Majewski, A. Durmus, and F.-R. Stöter, “Sliced-wasserstein flows: Nonparametric generative modeling via optimal transport and diffusions,” in International Conference on Machine Learning . PMLR, 2019, pp. 4104–4113
2019
Earlier work this paper cites.
J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington, “Wide neural networks of any depth evolve as linear models under gradient descent,” in Advances in Neural Information Processing Systems , vol. 32, 2019, pp. 8572–8583
2019
Earlier work this paper cites.
P. Helber, B. Bischke, A. Dengel, and D. Borth, “Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 12, no. 7, pp. 2217–2226, 2019
2019
Earlier work this paper cites.
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar, “Do imagenet classifiers generalize to imagenet?” in International Conference on Machine Learning . PMLR, 2019, pp. 5389–5400
2019
Earlier work this paper cites.
H. Wang, S. Ge, Z. Lipton, and E. P. Xing, “Learning robust global representations by penalizing local predictive power,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
A. Gupta, P. Dollar, and R. Girshick, “Lvis: A dataset for large vocabulary instance segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 5356–5364
2019
Cited alongside, same era.
W. Zhang, C. Ma, Q. Wu, and X. Yang, “Language-guided navigation via cross-modal grounding and alternate adversarial learning,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 31, no. 9, pp. 3469–3481, 2020
2020
Cited alongside, same era.
Z. Yang, T. Kumar, T. Chen, J. Su, and J. Luo, “Grounding-tracking-integration,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 31, no. 9, pp. 3433–3443, 2020
2020
Cited alongside, same era.
T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn, “Gradient surgery for multi-task learning,” Advances in Neural Information Processing Systems , vol. 33, pp. 5824–5836, 2020
2020
Cited alongside, same era.
T. Mei, J. J. Corso, G. Kim, J. Luo, C. Shen, and H. Zhang, “Guest editorial introduction to the special section on video and language,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 32, no. 1, pp. 1–4, 2022
2022
Closest in time.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision-language models,” International Journal of Computer Vision , vol. 130, no. 9, pp. 2337–2348, 2022
2022
Closest in time.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu., “Conditional prompt learning for vision-language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 816–16 825
2022
Closest in time.
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Geng, S.-j. Huang, and S. Chen, “Recent advances in open set recognition: A survey,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 10, pp. 3614–3631, 2020
2020
Cited alongside, same era.
S. Yang, L. Liu, and M. Xu, “Free lunch for few-shot learning: Distribution calibration,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
W. Jiang, K. Huang, J. Geng, and X. Deng, “Multi-scale metric learning for few-shot learning,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 31, no. 3, pp. 1091–1102, 2020
2020
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 8748–8763
2021
Cited alongside, same era.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 4904–4916
2021
Cited alongside, same era.
2021
Cited alongside, same era.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 3045–3059
2021
Cited alongside, same era.
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 4582–4597
2021
Cited alongside, same era.
2022
Closest in time.
D. Cheng, J. Zhou, N. Wang, and X. Gao, “Hybrid dynamic contrast and probability distillation for unsupervised person re-id,” IEEE Transactions on Image Processing , vol. 31, pp. 3334–3346, 2022
2022
Closest in time.
X. Zhai, X. Wang, B. Mustafa, A. Steiner, D. Keysers, A. Kolesnikov, and L. Beyer, “Lit: Zero-shot transfer with locked-image text tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 123–18 133
2022
Closest in time.
C. Liang, W. Wang, T. Zhou, and Y. Yang, “Visual abductive reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 565–15 575
2022
Closest in time.
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang, “Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” in International Conference on Machine Learning . PMLR, 2022, pp. 23 318–23 340
2022
Closest in time.
2022
Closest in time.
H. Wang, W. Liang, J. Shen, L. Van Gool, and W. Wang, “Counterfactual cycle-consistent learning for instruction following and generation in vision-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 471–15 481
2022
Closest in time.
Y. Lu, J. Liu, Y. Zhang, Y. Liu, and X. Tian, “Prompt distribution learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 5206–5215
2022
Closest in time.
2022
Closest in time.
Y. Yao, Q. Chen, A. Zhang, W. Ji, Z. Liu, T.-S. Chua, and M. Sun, “Pevl: Position-enhanced pre-training and prompt tuning for vision-language models,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022
2022
Closest in time.
2022
Closest in time.
Y. Zhang, K. Zhou, and Z. Liu, “Neural prompt search,” arXiv preprint arXiv:2206.04673 , 2022
2022
Closest in time.
J. Wang, C. Lan, C. Liu, Y. Ouyang, T. Qin, W. Lu, Y. Chen, W. Zeng, and P. Yu, “Generalizing to unseen domains: A survey on domain generalization,” IEEE Transactions on Knowledge and Data Engineering , 2022
2022
Closest in time.
K. Zhou, Z. Liu, Y. Qiao, T. Xiang, and C. C. Loy, “Domain generalization: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2022
2022
Closest in time.
Y. Wang, L. Qi, Y. Shi, and Y. Gao, “Feature-based style randomization for domain generalization,” IEEE Transactions on Circuits and Systems for Video Technology , 2022
2022
Closest in time.
T. Li, L. Tan, Z. Huang, Q. Tao, Y. Liu, and X. Huang, “Low dimensional trajectory hypothesis is true: Dnns can be trained in tiny subspaces,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2022
2022
Closest in time.
B. W. Larsen, S. Fort, N. Becker, and S. Ganguli, “How many degrees of freedom do we need to train deep networks: a loss landscape perspective,” in International Conference on Learning Representations , 2022
2022
Closest in time.