Zero shot detection
Pengkai Zhu, Hanxiao Wang, and Venkatesh Saligrama · 2020
Later among the works it cites.
Exploiting a joint embedding space for generalized zero-shot semantic segmentation
Donghyeon Baek, Youngmin Oh, and Bumsub Ham · 2021
Later among the works it cites.
Robustnav: Towards benchmarking robustness in embodied navigation
Prithvijit Chattopadhyay, Judy Hoffman, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2021
Later among the works it cites.
Transformer interpretability beyond attention visualization
Hila Chefer, Shir Gur, and Lior Wolf · 2021
Later among the works it cites.
Sign: Spatial-information incorporated generative network for generalized zero-shot semantic segmentation
Jiaxin Cheng, Soumyaroop Nandi, Prem Natarajan, and Wael Abd-Almageed · 2021
Later among the works it cites.
Zero-shot detection via vision and language knowledge distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui · 2021
Later among the works it cites.
No rl, no simulation: Learning to navigate without navigating
Meera Hahn, Devendra Singh Chaplot, and Shubham Tulsiani · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
Mdetr - modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion · 2021
Later among the works it cites.
Simple but effective: Clip embeddings for embodied ai
Apoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2021
Later among the works it cites.
Sscnav: Confidence-aware semantic scene completion for visual semantic navigation
Yiqing Liang, Boyuan Chen, and Shuran Song · 2021
Later among the works it cites.
Memory-augmented reinforcement learning for image-goal navigation
Lina Mezghani, Sainbayar Sukhbaatar, Thibaut Lavril, Oleksandr Maksymets, Dhruv Batra, Piotr Bojanowski, and Alahari Karteek · 2021
Later among the works it cites.
Interesting object, curious agent: Learning task-agnostic exploration
Simone Parisi, Victoria Dean, Deepak Pathak, and Abhinav Kumar Gupta · 2021
Later among the works it cites.
Combined scaling for zero-shot transfer learning
Hieu Pham, Zihang Dai, Golnaz Ghiasi, Hanxiao Liu, Adams Wei Yu, Minh-Thang Luong, Mingxing Tan, and Quoc V. Le · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Visual room rearrangement
Luca Weihs, Matt Deitke, Aniruddha Kembhavi, and Roozbeh Mottaghi · 2021
Later among the works it cites.
Denseclip: Extract free dense labels from clip
Chong Zhou, Chen Change Loy, and Bo Dai · 2021
Later among the works it cites.
Zero experience required: Plug & play modular transfer learning for semantic visual navigation
Ziad Al-Halah, Santhosh K. Ramakrishnan, and Kristen Grauman · 2022
Closest in time.
Procthor: Large-scale embodied ai using procedural generation
Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Jordi Salvador, Kiana Ehsani, Winson Han, Eric Kolve, Ali Farhadi, Aniruddha Kembhavi, and Roozbeh Mottaghi · 2022
Closest in time.
Continuous scene representations for embodied ai
Samir Yitzhak Gadre, Kiana Ehsani, Shuran Song, and Roozbeh Mottaghi · 2022
Closest in time.
Visual language maps for robot navigation
Original
Chenguang Huang, Oier Mees, Andy Zeng, and Wolfram Burgard · 2022
Closest in time.
Zson: Zero-shot object-goal navigation using multimodal goal embeddings
Original
Arjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman, and Dhruv Batra · 2022
Closest in time.
Simple open-vocabulary object detection with vision transformers
Matthias Minderer, Alexey A. Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, Xiao Wang, Xiaohua Zhai, Thomas Kipf, and Neil Houlsby · 2022
Closest in time.
The introspective agent: Interdependence of strategy, physiology, and sensing for embodied agents
Sarah Pratt, Luca Weihs, and Ali Farhadi · 2022
Closest in time.
Poni: Potential functions for objectgoal navigation with interaction-free learning
Santhosh K. Ramakrishnan, Devendra Singh Chaplot, Ziad Al-Halah, Jitendra Malik, and Kristen Grauman · 2022
Closest in time.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Closest in time.
Lit: Zero-shot transfer with locked-image text tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer · 2022
Closest in time.