Fetching the paper…
Reading the bibliography…
The world is filled with a wide variety of objects.
A. Kadian, J. Truong, A. Gokaslan, A. Clegg, E. Wijmans, S. Lee, M. Savva, S. Chernova, and D. Batra · 1912
Earlier work this paper cites.
A robot exploration and mapping strategy based on a semantic hierarchy of spatial representations
B. Kuipers and Y.-T. Byun · 1991
Earlier work this paper cites.
Monte carlo localization for mobile robots
F. Dellaert, D. Fox, W. Burgard, and S. Thrun · 1999
Earlier work this paper cites.
Objectnav revisited: On evaluation of embodied agents navigating to objects, 2020
D. Batra, A. Gokaslan, A. Kembhavi, O. Maksymets, R. Mottaghi, M. Savva, A. Toshev, and E. Wijmans · 2006
Earlier work this paper cites.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
H. Mei, M. Bansal, and M. R. Walter · 2015
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Semi-parametric topological memory for navigation
N. Savinov, A. Dosovitskiy, and V. Koltun · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Earlier work this paper cites.
Language-conditioned imitation learning for robot manipulation tasks, 2020
S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. B. Amor · 2020
Earlier work this paper cites.
Objectnav revisited: On evaluation of embodied agents navigating to objects, 2020
D. Batra, A. Gokaslan, A. Kembhavi, O. Maksymets, R. Mottaghi, M. Savva, A. Toshev, and E. Wijmans · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Structformer: Learning spatial structure for language-guided semantic rearrangement of novel objects
W. Liu, C. Paxton, T. Hermans, and D. Fox · 2021
Earlier work this paper cites.
Rapid Exploration for Open-World Navigation with Latent Goal Models
D. Shah, B. Eysenbach, N. Rhinehart, and S. Levine · 2021
Earlier work this paper cites.
Learning generalizable robotic reward functions from ”in-the-wild” human videos, 2021
A. S. Chen, S. Nair, and C. Finn · 2021
Earlier work this paper cites.
Do as i can and not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, M. Yan, and A. Zeng · 2022
Earlier work this paper cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Earlier work this paper cites.
Clip-nav: Using clip for zero-shot vision-and-language navigation, 2022
V. S. Dorbala, G. Sigurdsson, R. Piramuthu, J. Thomason, and G. S. Sukhatme · 2022
Earlier work this paper cites.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action, 2022
D. Shah, B. Osinski, B. Ichter, and S. Levine · 2022
Earlier work this paper cites.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2022
Earlier work this paper cites.
Progprompt: Generating situated robot task plans using large language models, 2022
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2022
Earlier work this paper cites.
Depth360: Self-supervised learning for monocular depth estimation using learnable camera distortion model
N. Hirose and K. Tahara · 2022
Earlier work this paper cites.
Bc-z: Zero-shot task generalization with robotic imitation learning, 2022
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Earlier work this paper cites.
What matters in language conditioned robotic imitation learning over unstructured data
O. Mees, L. Hermann, and W. Burgard · 2022
Cited alongside, same era.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard · 2022
Cited alongside, same era.
Navigating to objects in the real world, 2022
T. Gervet, S. Chintala, D. Batra, J. Malik, and D. S. Chaplot · 2022
Cited alongside, same era.
Simple but effective: Clip embeddings for embodied ai, 2022
A. Khandelwal, L. Weihs, R. Mottaghi, and A. Kembhavi · 2022
Cited alongside, same era.
Last-mile embodied visual navigation, 2022
J. Wasserman, K. Yadav, G. Chowdhary, A. Gupta, and U. Jain · 2022
Cited alongside, same era.
Vip: Towards universal visual reward and representation via value-implicit pre-training, 2023
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2023
Later among the works it cites.
Affordances from human videos as a versatile representation for robotics
S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y. Lu, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. Ryoo, G. Salazar, P. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. Vuong, A. Wahid, S. Welker, P. Wohlhart, J. Wu, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2023
Later among the works it cites.
Vlfm: Vision-language frontier maps for zero-shot semantic navigation, 2023
N. Yokoyama, S. Ha, D. Batra, J. Wang, and B. Bucher · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Cited alongside, same era.
Learning to imitate object interactions from internet videos
A. Patel, A. Wang, I. Radosavovic, and J. Malik · 2022
Cited alongside, same era.
Simple open-vocabulary object detection
M. Minderer, A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen, et al · 2022
Cited alongside, same era.
Exaug: Robot-conditioned navigation policies via geometric experience augmentation, 2022
N. Hirose, D. Shah, A. Sridhar, and S. Levine · 2022
Cited alongside, same era.
Bridgedata v2: A dataset for robot learning at scale
H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V. Myers, K. Fang, C. Finn, and S. Levine · 2023
Cited alongside, same era.
Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot, 2023
H.-S. Fang, H. Fang, Z. Tang, J. Liu, C. Wang, J. Wang, H. Zhu, and C. Lu · 2023
Cited alongside, same era.
Scaling up and distilling down: Language-guided robot skill acquisition
H. Ha, P. Florence, and S. Song · 2023
Cited alongside, same era.
Later among the works it cites.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models
W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, X. Song, et al · 2023
Later among the works it cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Later among the works it cites.
ViNT: A foundation model for visual navigation
D. Shah, A. Sridhar, N. Dashora, K. Stachowicz, K. Black, N. Hirose, and S. Levine · 2023
Later among the works it cites.
Sacson: Scalable autonomous control for social navigation
N. Hirose, D. Shah, A. Sridhar, and S. Levine · 2023
Later among the works it cites.
Scaling open-vocabulary object detection
M. Minderer, A. Gritsenko, and N. Houlsby · 2023
Later among the works it cites.
Zoedepth: Zero-shot transfer by combining relative and metric depth
S. F. Bhat, R. Birkl, D. Wofk, P. Wonka, and M. Müller · 2023
Later among the works it cites.
Zoedepth: Zero-shot transfer by combining relative and metric depth, 2023
S. F. Bhat, R. Birkl, D. Wofk, P. Wonka, and M. Müller · 2023
Later among the works it cites.
Droid: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Closest in time.
Multi-task robot data for dual-arm fine manipulation
H. Kim, Y. Ohmura, and Y. Kuniyoshi · 2024
Closest in time.
Dinov2: Learning robust visual features without supervision, 2024
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski · 2024
Closest in time.
Pivot: Iterative visual prompting elicits actionable knowledge for vlms
S. Nasiriany, F. Xia, W. Yu, T. Xiao, J. Liang, I. Dasgupta, A. Xie, D. Driess, A. Wahid, Z. Xu, et al · 2024
Closest in time.
Gensim: Generating robotic simulation tasks via large language models, 2024
L. Wang, Y. Ling, Z. Yuan, M. Shridhar, C. Bao, Y. Qin, B. Wang, H. Xu, and X. Wang · 2024
Closest in time.
https://www.youtube.com/ , Accessed: 2024-06-06
youtube · 2024
Closest in time.
Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots
C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song · 2024
Closest in time.
Openfmnav: Towards open-set zero-shot object navigation via vision-language foundation models
Y. Kuang, H. Lin, and M. Jiang · 2024
Closest in time.
Vunet: Dynamic scene view synthesis for traversability estimation using an rgb camera
N. Hirose, A. Sadeghian, F. Xia, R. Martín-Martín, and S. Savarese · 2069
Closest in time.