Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) models, which fall under the category of vision-language models, conventionally execute multiple downsampling processes on image inputs to strike a balance between computational efficiency and model performance.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Language models for image captioning: The quirks and what works, 2015
Jacob Devlin, Hao Cheng, Hao Fang, Saurabh Gupta, Li Deng, Xiaodong He, Geoffrey Zweig, and Margaret Mitchell · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Visual question answering: A survey of methods and datasets, 2016
Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton van den Hengel · 2016
Earlier work this paper cites.
Visual dialog, 2017
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M. F. Moura, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language, 2019
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks, 2019
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations, 2020
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2020
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision, 2021
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts, 2022
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, and Furu Wei · 2022
Cited alongside, same era.
An end-to-end fast no-reference video quality predictor with spatiotemporal feature fusion
Anish Kumar Vishwakarma and Kishor M. Bhurchandi · 2023
Later among the works it cites.
Drivegpt4: Interpretable end-to-end autonomous driving via large language model
Zhenhua Xu, Yujia Zhang, Enze Xie, Zhen Zhao, Yong Guo, Kenneth KY Wong, Zhenguo Li, and Hengshuang Zhao · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Hang Zhang, Xin Li, Lidong Bing, and at al · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Explainable artificial intelligence for autonomous driving: A comprehensive overview and field guide for future research directions, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi · 2022
Cited alongside, same era.
Parameter-efficient image-to-video transfer learning for action recognition
Junting Pan, Ziyi Lin, Xiatian Zhu, Jing Shao, and Hongsheng Li ST-Adapter · 2022
Cited alongside, same era.
Explaining autonomous driving actions with visual question answering, 2023
Shahin Atakishiyev, Mohammad Salameh, Housam Babiker, and Randy Goebel · 2023
Cited alongside, same era.
Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving
Vikrant Dewangan, Tushar Choudhary, Shivam Chandhok, Shubham Priyadarshan, Anushka Jain, Arun K Singh, Siddharth Srivastava, Krishna Murthy Jatavallabhula, and K Madhava Krishna · 2023
Cited alongside, same era.
Xinpeng Ding, Jianhua Han, Hang Xu, Wei Zhang, and Xiaomeng Li · 2023
Cited alongside, same era.
Content-adaptive downsampling in convolutional neural networks, 2023
Robin Hesse, Simone Schaub-Meyer, and Stefan Roth · 2023
Cited alongside, same era.
Drama: Joint risk localization and captioning in driving
Srikanth Malla, Chiho Choi, Isht Dwivedi, Joon Hee Choi, and Jiachen Li · 2023
Cited alongside, same era.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi
Cited in the paper.
Shahin Atakishiyev, Mohammad Salameh, Hengshuai Yao, and Randy Goebel · 2024
Later among the works it cites.
Akshay Gopalkrishnan, Ross Greer, and Mohan Trivedi · 2024
Later among the works it cites.
Yolov8: A novel object detection algorithm with enhanced performance and robustness
Rejin Varghese and Sambath M · 2024
Later among the works it cites.
Holistic autonomous driving understanding by bird’s-eye-view injected multi-modal large models
Ding Xinpeng, Han Jinahua, Xu Hang, Laing Xiaodan, Hang Xu, Zhang Wei, and Li Xiaomeng · 2024
Later among the works it cites.
Image understanding through visual question answering: A review from past research
Nagamani Yanda, J. Tagore Babu, K. Aswin Kumar, M. Taraka Rama Rao, K. V. Ranjith Varma, and N. Rahul Babu · 2024
Later among the works it cites.
Dynrefer: Delving into region-level multi-modality tasks via dynamic resolution, 2024
Yuzhong Zhao, Feng Liu, Yue Liu, Mingxiang Liao, Chen Gong, Qixiang Ye, and Fang Wan · 2024
Later among the works it cites.
Controlcap: Controllable region-level captioning
Yuzhong Zhao, Yue Liu, Zonghao Guo, Weijia Wu, Chen Gong, Qixiang Ye, and Fang Wan · 2025
Closest in time.