Fetching the paper…
Reading the bibliography…
Vision-language models (VLMs) have recently shown promising results in traditional downstream tasks.
Objects, attributes, and visual attention: Which, what, and where
Nancy Kanwisher and Jon Driver · 1992
Earlier work this paper cites.
Allocentric and egocentric spatial representations: Definitions, distinctions, and interconnections
Roberta L Klatzky · 1998
Earlier work this paper cites.
rob@ work: Robot assistant in industrial environments
Evert Helms, Rolf Dieter Schraft, and M Hagele · 2002
Earlier work this paper cites.
Visual landmarks detection and recognition for mobile robot navigation
Jean-Bernard Hayet, Frédéric Lerasle, and Michel Devy · 2003
Earlier work this paper cites.
Allocentric and egocentric updating of spatial memories
Weimin Mou, Timothy P McNamara, Christine M Valiquette, and Björn Rump · 2004
Earlier work this paper cites.
FastSLAM: A scalable method for the simultaneous localization and mapping problem in robotics
Michael Montemerlo and Sebastian Thrun · 2007
Earlier work this paper cites.
Learning object affordances: from sensory–motor coordination to imitation
Luis Montesano, Manuel Lopes, Alexandre Bernardino, and José Santos-Victor · 2008
Earlier work this paper cites.
Describing objects by their attributes
Ali Farhadi, Ian Endres, Derek Hoiem, and David Forsyth · 2009
Earlier work this paper cites.
Understanding egocentric activities
Alireza Fathi, Ali Farhadi, and James M Rehg · 2011
Earlier work this paper cites.
Affordance prediction via learned object attributes
Tucker Hermans, James M Rehg, and Aaron Bobick · 2011
Earlier work this paper cites.
Interacting with a robot: a guide robot understanding natural language instructions
Loreto Susperregi, Izaskun Fernandez, Ane Fernandez, Santiago Fernandez, Iñaki Maurtua, and Irene Lopez de Vallejo · 2012
Earlier work this paper cites.
A review on video-based human activity recognition
Shian-Ru Ke, Hoang Le Uyen Thuc, Yong-Jin Lee, Jenq-Neng Hwang, Jang-Hee Yoo, and Kyoung-Ho Choi · 2013
Earlier work this paper cites.
From allo-to egocentric spatial ability in early alzheimer’s disease: a study with virtual reality spatial tasks
Francesca Morganti, Stefano Stefanini, and Giuseppe Riva · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Path planning navigation of mobile robot with obstacles avoidance using fuzzy logic controller
Anish Pandey, Rakesh Kumar Sonkar, Krishna Kant Pandey, and DR Parhi · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions
Sven Bambach, Stefan Lee, David J Crandall, and Chen Yu · 2015
Earlier work this paper cites.
Evidence for the embodiment of space perception: concurrent hand but not arm action moderates reachability and egocentric distance perception
Stéphane Grade, Mauro Pesenti, and Martin G Edwards · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
A review of human activity recognition methods
Michalis Vrigkas, Christophoros Nikou, and Ioannis A Kakadiaris · 2015
Earlier work this paper cites.
A pointing gesture based egocentric interaction system: Dataset, approach and application
Yichao Huang, Xiaorui Liu, Xin Zhang, and Lianwen Jin · 2016
Earlier work this paper cites.
Recognition of activities of daily living with egocentric vision: A review
Thi-Hoa-Cuc Nguyen, Jean-Christophe Nebel, and Francisco Florez-Revuelta · 2016
Earlier work this paper cites.
Multiple-robot simultaneous localization and mapping: A review
Sajad Saeedi, Michael Trentini, Mae Seto, and Howard Li · 2016
Earlier work this paper cites.
Image captioning with semantic attention
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo · 2016
Cited alongside, same era.
Next-active-object prediction from egocentric videos
Antonino Furnari, Sebastiano Battiato, Kristen Grauman, and Giovanni Maria Farinella · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Egogesture: A new dataset and benchmark for egocentric hand gesture recognition
Yifan Zhang, Congqi Cao, Jian Cheng, and Hanqing Lu · 2018
Cited alongside, same era.
Egovqa-an egocentric video question answering benchmark dataset
Chenyou Fan · 2019
Cited alongside, same era.
A comprehensive survey of deep learning for image captioning
MD Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga · 2019
A-okvqa: A benchmark for visual question answering using world knowledge
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi · 2022
Later among the works it cites.
Vl-checklist: Evaluating pre-trained vision-language models with objects, attributes and relations
Tiancheng Zhao, Tianqi Zhang, Mingwei Zhu, Haozhan Shen, Kyusong Lee, Xiaopeng Lu, and Jianwei Yin · 2022
Later among the works it cites.
Vlue: A multi-task benchmark for evaluating vision-language models
Wangchunshu Zhou, Yan Zeng, Shizhe Diao, and Xinsong Zhang · 2022
Later among the works it cites.
Towards language models that can see: Computer vision through the lens of natural language
William Berrios, Gautam Mittal, Tristan Thrush, Douwe Kiela, and Amanpreet Singh · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Cited alongside, same era.
Human activity recognition: A survey
Charmi Jobanputra, Jatna Bavishi, and Nishant Doshi · 2019
Cited alongside, same era.
A review: On path planning strategies for navigation of mobile robot
BK Patle, Anish Pandey, DRK Parhi, AJDT Jagadeesh, et al · 2019
Cited alongside, same era.
Object detection with deep learning: A review
Zhong-Qiu Zhao, Peng Zheng, Shou-tao Xu, and Xindong Wu · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Forecasting human-object interaction: Joint prediction of motor attention and egocentric activity
Miao Liu, Siyu Tang, Yin Li, and James M Rehg · 2020
Cited alongside, same era.
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi · 2023
Closest in time.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al · 2023
Closest in time.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, et al · 2023
Closest in time.
Voxposer: Composable 3d value maps for robotic manipulation with language models
Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei · 2023
Closest in time.
Video-llava: Learning united visual representation by alignment before projection
Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Paco: Parts and attributes of common objects
Vignesh Ramanathan, Anmol Kalia, Vladan Petrovic, Yi Wen, Baixue Zheng, Baishan Guo, Rui Wang, Aaron Marquez, Rama Kovvuri, Abhishek Kadian, et al · 2023
Closest in time.
Tiny lvlm-ehub: Early multimodal experiments with bard
Wenqi Shao, Yutao Hu, Peng Gao, Meng Lei, Kaipeng Zhang, Fanqing Meng, Peng Xu, Siyuan Huang, Hongsheng Li, Yu Qiao, et al · 2023
Closest in time.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M Sadler, Wei-Lun Chao, and Yu Su · 2023
Closest in time.
Pandagpt: One model to instruction-follow them all
Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai · 2023
Closest in time.
Distilling internet-scale vision-language models into embodied agents
Theodore Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus, and Ishita Dasgupta · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Openchat: Advancing open-source language models with mixed-quality data
Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu · 2023
Closest in time.
Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models
Peng Xu, Wenqi Shao, Kaipeng Zhang, Peng Gao, Shuo Liu, Meng Lei, Fanqing Meng, Siyuan Huang, Yu Qiao, and Ping Luo · 2023
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality, 2023
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chaoya Jiang, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Closest in time.
Object detection in 20 years: A survey
Zhengxia Zou, Keyan Chen, Zhenwei Shi, Yuhong Guo, and Jieping Ye · 2023
Closest in time.