Fetching the paper…
Reading the bibliography…
Contrastive Language-Image Pre-training (CLIP) has been shown to learn visual representations with great transferability, which achieves promising accuracy for zero-shot classification.
Multi-view 3d object detection network for autonomous driving
Chen, X.; Ma, H.; Wan, J.; Li, B.; and Xia, T. 2017 · 1915
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Wu, Z.; Song, S.; Khosla, A.; Yu, F.; Zhang, L.; Tang, X.; and Xiao, J. 2015 · 1920
Earlier work this paper cites.
Learning Generative Visual Models from Few Training Examples: An Incremental Bayesian Approach Tested on 101 Object Categories
Li, F.; Fergus, R.; and Perona, P. 2004 · 2004
Earlier work this paper cites.
Automated Flower Classification over a Large Number of Classes
Nilsback, M. E.; and Zisserman, A. 2008 · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia, D.; Wei, D.; Socher, R.; Li, L. J.; Kai, L.; and Li, F. F. 2009 · 2009
Earlier work this paper cites.
SUN database: Large-scale scene recognition from abbey to zoo
Xiao, J.; Hays, J.; Ehinger, K. A.; Oliva, A.; and Torralba, A. 2010 · 2010
Earlier work this paper cites.
End-to-end object detection with adaptive clustering transformer
Zheng, M.; Gao, P.; Wang, X.; Li, H.; and Dong, H. 2020 · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Soomro, K.; Zamir, A. R.; and Shah, M. 2012 · 2012
Earlier work this paper cites.
Cats and dogs
Vedaldi, A. 2012 · 2012
Earlier work this paper cites.
Describing Textures in the Wild
Cimpoi, M.; Maji, S.; Kokkinos, I.; Mohamed, S.; and Vedaldi, A. 2013 · 2013
Earlier work this paper cites.
Fine-Grained Visual Classification of Aircraft
Maji, S.; Rahtu, E.; Kannala, J.; Blaschko, M.; and Vedaldi, A. 2013 · 2013
Earlier work this paper cites.
Food-101 – Mining Discriminative Components with Random Forests
Bossard, L.; Guillaumin, M.; and Gool, L. V. 2014 · 2014
Earlier work this paper cites.
3D Object Representations for Fine-Grained Categorization
Krause, J.; Stark, M.; Deng, J.; and Li, F. F. 2014 · 2014
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Cited alongside, same era.
Gutman, D.; Codella, N. C.; Celebi, E.; Helba, B.; Marchetti, M.; Mishra, N.; and Halpern, A. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification
Helber, P.; Bischke, B.; Dengel, A.; and Borth, D. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Pct: Point cloud transformer
Guo, M.-H.; Cai, J.-X.; Liu, Z.-N.; Mu, T.-J.; Martin, R. R.; and Hu, S.-M. 2021 · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L.; and Liang, P. 2021 · 2021
Later among the works it cites.
Dual-stream network for visual recognition
Mao, M.; Zhang, R.; Zheng, H.; Gao, P.; Ma, T.; Peng, Y.; Ding, E.; Zhang, B.; and Han, S. 2021 · 2021
Later among the works it cites.
Trackformer: Multi-object tracking with transformers
Meinhardt, T.; Kirillov, A.; Leal-Taixe, L.; and Feichtenhofer, C. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Image transformer
Parmar, N.; Vaswani, A.; Uszkoreit, J.; Kaiser, L.; Shazeer, N.; Ku, A.; and Tran, D. 2018 · 2018
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
Recht, B.; Roelofs, R.; Schmidt, L.; and Shankar, V. 2019 · 2019
Cited alongside, same era.
Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data
Uy, M. A.; Pham, Q.-H.; Hua, B.-S.; Nguyen, T.; and Yeung, S.-K. 2019 · 2019
Cited alongside, same era.
Learning robust global representations by penalizing local predictive power
Wang, H.; Ge, S.; Lipton, Z.; and Xing, E. P. 2019 · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
Rao, Y.; Zhao, W.; Chen, G.; Tang, Y.; Zhu, Z.; Huang, G.; Zhou, J.; and Lu, J. 2021 · 2021
Later among the works it cites.
SegFormer: Simple and efficient design for semantic segmentation with transformers
Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J. M.; and Luo, P. 2021 · 2021
Later among the works it cites.
Learning to prompt for vision-language models
Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2021 · 2021
Later among the works it cites.
Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model
Du, Y.; Wei, F.; Zhang, Z.; Shi, M.; Gao, Y.; and Li, G. 2022 · 2022
Closest in time.
A Multi-level Mesh Mutual Attention Model for Visual Question Answering
Lei, Z.; Zhang, G.; Wu, L.; Zhang, K.; and Liang, R. 2022 · 2022
Closest in time.
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
Li, K.; Wang, Y.; Zhang, J.; Gao, P.; Song, G.; Liu, Y.; Li, H.; and Qiao, Y. 2022 · 2022
Closest in time.
Frozen clip models are efficient video learners
Lin, Z.; Geng, S.; Zhang, R.; Gao, P.; de Melo, G.; Wang, X.; Dai, J.; Qiao, Y.; and Li, H. 2022 · 2022
Closest in time.
PointCLIP V2: Adapting CLIP for Powerful 3D Open-world Learning
Zhu, X.; Zhang, R.; He, B.; Zeng, Z.; Zhang, S.; and Gao, P. 2022 · 2022
Closest in time.