Fetching the paper…
Reading the bibliography…
In recent years, the integration of vision and language understanding has led to significant advancements in artificial intelligence, particularly through Vision-Language Models (VLMs).
Eye movements and vision
Alfred L. Yarbus · 1967
Earlier work this paper cites.
Shifts in selective visual attention: towards the underlying neural circuitry
Christof Koch and Shimon Ullman · 1985
Earlier work this paper cites.
Control of selective visual attention: Modeling the where pathway
Ernst Niebur and Christof Koch · 1995
Earlier work this paper cites.
Interacting with eye movements in virtual environments
Vildan Tanriverdi and Robert JK Jacob · 2000
Earlier work this paper cites.
What can a mouse cursor tell us more?: correlation of eye/mouse movements on web browsing
Mon-Chu Chen, John R. Anderson, and Myeong-Ho Sohn · 2001
Earlier work this paper cites.
A nonparametric approach to bottom-up visual saliency
Wolfgang Kienzle, Felix Wichmann, Bernhard Scholkopf, and Matthias O. Franz · 2006
Earlier work this paper cites.
Eye movements and the control of actions in everyday life
Michael F Land · 2006
Earlier work this paper cites.
Learning to predict where humans look
Tilke Judd, Krista A. Ehinger, Frédo Durand, and Antonio Torralba · 2009
Earlier work this paper cites.
Towards predicting web searcher gaze position from mouse movements
Qi Guo and Eugene Agichtein · 2010
Earlier work this paper cites.
No clicks, no problem: using cursor movements to understand and improve search
Jeff Huang, Ryen W. White, and Susan T. Dumais · 2011
Earlier work this paper cites.
User see, user point: gaze and cursor alignment in web search
Jeff Huang, Ryen W. White, and Georg Buscher · 2012
Earlier work this paper cites.
Salicon: Saliency in context
Ming Jiang, Shengsheng Huang, Juanyong Duan, and Qi Zhao · 2015
Earlier work this paper cites.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
Abhishek Das, Harsh Agrawal, C. Lawrence Zitnick, Devi Parikh, and Dhruv Batra · 2016
Earlier work this paper cites.
Deepgaze ii: Reading fixations from deep features trained on object recognition
Matthias Kümmerer, Thomas S. A. Wallis, and Matthias Bethge · 2016
Earlier work this paper cites.
Shallow and deep convolutional networks for saliency prediction
Junting Pan, Elisa Sayrol, Xavier Giro i Nieto, Kevin McGuinness, and Noel E. O’Connor · 2016
Earlier work this paper cites.
Seeing with humans: Gaze-assisted neural image captioning
Yusuke Sugano and Andreas Bulling · 2016
Earlier work this paper cites.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Bubbleview: an interface for crowdsourcing image importance maps and tracking visual attention
Nam Wook Kim, Zoya Bylinskii, Michelle A Borkin, Krzysztof Z Gajos, Aude Oliva, Fredo Durand, and Hanspeter Pfister · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Referring expression generation and comprehension via attributes
Jingyu Liu, Liang Wang, and Ming-Hsuan Yang · 2017
Cited alongside, same era.
Saliency revisited: Analysis of mouse movements versus fixations
Hamed Rezazadegan Tavakoli, Fawad Ahmed, Ali Borji, and Jorma T. Laaksonen · 2017
Cited alongside, same era.
Object referring in videos with language and human gaze
Arun Balajee Vasudevan, Dengxin Dai, and Luc Van Gool · 2018
Cited alongside, same era.
Interactive image segmentation with first click attention
Zheng Lin, Zhao Zhang, Lin-Zhuo Chen, Ming-Ming Cheng, and Shao-Ping Lu · 2020
Cited alongside, same era.
Point and ask: Incorporating pointing into visual question answering
Arjun Mani, Nobline Yoo, Will Hinthorn, and Olga Russakovsky · 2020
Cited alongside, same era.
Connecting vision and language with localized narratives
https://huggingface.co/OpenAssistant/reward-model-deberta-v3-large-v2 , 2023
Openassistant/reward-model-deberta-v3-large-v2 · 2023
Closest in time.
Openflamingo: An open-source framework for training large autoregressive vision-language models
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al · 2023
Closest in time.
Disclip: Open-vocabulary referring expression generation
Lior Bracha, Eitan Shaar, Aviv Shamsian, Ethan Fetaya, and Gal Chechik · 2023
Closest in time.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao · 2023
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut, and Vittorio Ferrari · 2020
Cited alongside, same era.
A high-level description and performance evaluation of pupil invisible
Marc Tonsen, Chris Kay Baumann, and Kai Dierkes · 2020
Cited alongside, same era.
Human gaze assisted artificial intelligence: A review
Ruohan Zhang, Akanksha Saran, Bo Liu, Yifeng Zhu, Sihang Guo, Scott Niekum, Dana H. Ballard, and Mary M. Hayhoe · 2020
Cited alongside, same era.
Unifying vision-and-language tasks via text generation
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Looking for info: Evaluation of gaze based information retrieval in augmented reality
Robin Piening, Robin Piening, Ken Pfeuffer, Augusto Esteves, Tim Mittermeier, Sarah Prange, Philippe Schröder, and Florian Alt · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi · 2023
Closest in time.
Multimodal-gpt: A vision and language model for dialogue with humans
Tao Gong, Chengqi Lyu, Shilong Zhang, Yudong Wang, Miao Zheng, Qian Zhao, Kuikun Liu, Wenwei Zhang, Ping Luo, and Kai Chen · 2023
Closest in time.
Grill: Grounded vision-language pre-training via aligning text and image regions
Woojeong Jin, Subhabrata Mukherjee, Yu Cheng, Yelong Shen, Weizhu Chen, Ahmed Hassan Awadallah, Damien Jose, and Xiang Ren · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Closest in time.
Eye-gaze-guided vision transformer for rectifying shortcut learning
Chong Ma, Lin Zhao, Yuzhong Chen, Sheng Wang, Lei Guo, Tuo Zhang, Dinggang Shen, Xi Jiang, and Tianming Liu · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Kosmos-2: Grounding multimodal large language models to the world
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei · 2023
Closest in time.
Gvgnet: Gaze-directed visual grounding for learning under-specified object referring intention
Kun Qian, Zhuoyang Zhang, Wei Song, and Jianfeng Liao · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
MosaicML NLP Team · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Connecting vision and language with video localized narratives
Paul Voigtlaender, Soravit Changpinyo, Jordi Pont-Tuset, Radu Soricut, and Vittorio Ferrari · 2023
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al · 2023
Closest in time.
Regionblip: A unified multi-modal pre-training framework for holistic and regional comprehension
Qiang Zhou, Chaohui Yu, Shaofeng Zhang, Sitong Wu, Zhibing Wang, and Fan Wang · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Closest in time.