Fetching the paper…
Reading the bibliography…
Recent developments in vision language models (VLM) have shown great potential for diverse applications related to image understanding.
Hoiem, D., S. K. Divvala, and J. H. Hays, Pascal VOC 2008 challenge. World Literature Today , Vol. 24, No. 1, 2009, pp. 1–4
2009
Earlier work this paper cites.
Lin, T.-Y., M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 , Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
Chakraborty, P., Y. O. Adu-Gyamfi, S. Poddar, V. Ahsani, A. Sharma, and S. Sarkar, Traffic congestion detection from camera images using deep convolution neural networks. Transportation Research Record , Vol. 2672, No. 45, 2018, pp. 222–231
2018
Earlier work this paper cites.
Dorafshan, S., R. J. Thomas, and M. Maguire, SDNET2018: An annotated image dataset for non-contact concrete crack detection using deep convolutional neural networks. Data in brief , Vol. 21, 2018, pp. 1664–1668
2018
Earlier work this paper cites.
Li, J., D. Li, C. Xiong, and S. Hoi, Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning , PMLR, 2022, pp. 12888–12900
2022
Earlier work this paper cites.
Minderer, M., A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen, et al., Simple open-vocabulary object detection. In European Conference on Computer Vision , Springer, 2022, pp. 728–755
2022
Earlier work this paper cites.
2023
Cited alongside, same era.
Hu, Y., D. Ou, X. Wang, and R. Yu, Enabling Vision-and-Language Navigation for Intelligent Connected Vehicles Using Large Pre-Trained Models. In 2023 IEEE International Conferences on Internet of Things (iThings) and IEEE Green Computing & Communications (GreenCom) and IEEE Cyber, Physical & Social Computing (CPSCom) and IEEE Smart Data (SmartData) and IEEE Congress on Cybermatics (Cybermatics) , IEEE, 2023, pp. 390–396
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Pan, C., B. Yaman, T. Nesti, A. Mallik, A. G. Allievi, S. Velipasalar, and L. Ren, VLP: Vision Language Planning for Autonomous Driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 14760–14769
2024
Closest in time.
2024
Closest in time.
Wang, S., D. C. Anastasiu, Z. Tang, M.-C. Chang, Y. Yao, L. Zheng, M. S. Rahman, M. S. Arya, A. Sharma, P. Chakraborty, et al., The 8th AI City Challenge. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 7261–7272
2024
Closest in time.
Hasan, M. Z., J. Chen, J. Wang, M. S. Rahman, A. Joshi, S. Velipasalar, C. Hegde, A. Sharma, and S. Sarkar, Vision-language models can identify distracted driver behavior from naturalistic videos. IEEE Transactions on Intelligent Transportation Systems , 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Terven, J., D.-M. Córdova-Esparza, and J.-A. Romero-González, A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas. Machine Learning and Knowledge Extraction , Vol. 5, No. 4, 2023, pp. 1680–1716
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Radford, A., J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transferable visual models from natural language supervision. In International conference on machine learning , PMLR, 2021a, pp. 8748–8763
Cited in the paper.
Radford, A., J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transferable visual models from natural language supervision. In International conference on machine learning , PMLR, 2021b, pp. 8748–8763
Cited in the paper.
2024
Closest in time.
2024
Closest in time.
Liu, H., C. Li, Q. Wu, and Y. J. Lee, Visual instruction tuning. Advances in neural information processing systems , Vol. 36, 2024
2024
Closest in time.