Fetching the paper…
Reading the bibliography…
Open-Vocabulary Mobile Manipulation (OVMM) is a crucial capability for autonomous robots, especially when faced with the challenges posed by unknown and dynamic environments.
A formal basis for the heuristic determination of minimum cost paths
P. E. Hart, N. J. Nilsson, and B. Raphael · 1968
Earlier work this paper cites.
Integration of sonar and stereo range data using a grid-based representation
L. Matthies and A. Elfes · 1988
Earlier work this paper cites.
Sensor fusion in certainty grids for mobile robots
H. P. Moravec · 1988
Earlier work this paper cites.
A frontier-based approach for autonomous exploration
B. Yamauchi · 1997
Earlier work this paper cites.
The dynamic window approach to collision avoidance
D. Fox, W. Burgard, and S. Thrun · 1997
Earlier work this paper cites.
Rrt-connect: An efficient approach to single-query path planning
J. Kuffner and S. LaValle · 2000
Earlier work this paper cites.
Planning and control in unstructured terrain
B. P. Gerkey and K. Konolige · 2008
Earlier work this paper cites.
The Open Motion Planning Library
I. A. Şucan, M. Moll, and L. E. Kavraki · 2012
Earlier work this paper cites.
Robot placement based on reachability inversion
N. Vahrenkamp, T. Asfour, and R. Dillmann · 2013
Earlier work this paper cites.
Orb-slam: a versatile and accurate monocular slam system
R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos · 2015
Earlier work this paper cites.
Orb-slam: A versatile and accurate monocular slam system
R. Mur-Artal, J. M. M. Montiel, and J. D. Tardós · 2015
Earlier work this paper cites.
Structure-from-motion revisited
J. L. Schönberger and J.-M. Frahm · 2016
Earlier work this paper cites.
Pixelwise view selection for unstructured multi-view stereo
J. L. Schönberger, E. Zheng, M. Pollefeys, and J.-M. Frahm · 2016
Earlier work this paper cites.
Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras
R. Mur-Artal and J. D. Tardós · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel · 2018
Earlier work this paper cites.
Online temporal calibration for monocular visual-inertial systems
T. Qin and S. Shen · 2018
Earlier work this paper cites.
Vins-mono: A robust and versatile monocular visual-inertial state estimator
T. Qin, P. Li, and S. Shen · 2018
Earlier work this paper cites.
Reuleaux: Robot base placement by reachability analysis
A. Makhal and A. K. Goins · 2018
Earlier work this paper cites.
On evaluation of embodied navigation agents
P. Anderson, A. Chang, D. S. Chaplot, A. Dosovitskiy, S. Gupta, V. Koltun, J. Kosecka, J. Malik, R. Mottaghi, M. Savva, et al · 2018
Earlier work this paper cites.
3d scene graph: A structure for unified semantics, 3d space, and camera
I. Armeni, Z.-Y. He, J. Gwak, A. R. Zamir, M. Fischer, J. Malik, and S. Savarese · 2019
Earlier work this paper cites.
Object goal navigation using goal-oriented semantic exploration
D. S. Chaplot, D. Gandhi, A. Gupta, and R. Salakhutdinov · 2020
Earlier work this paper cites.
Frontier detection and reachability analysis for efficient 2d graph-slam based active exploration
Z. Sun, B. Wu, C.-Z. Xu, S. E. Sarma, J. Yang, and H. Kong · 2020
Earlier work this paper cites.
Inertial-only optimization for visual-inertial initialization
C. Campos, J. M. Montiel, and J. D. Tardós · 2020
Earlier work this paper cites.
Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam
C. Campos, R. Elvira, J. J. G. Rodríguez, J. M. Montiel, and J. D. Tardós · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2021
Earlier work this paper cites.
Nerf in the wild: Neural radiance fields for unconstrained photo collections
R. Martin-Brualla, N. Radwan, M. S. Sajjadi, J. T. Barron, A. Dosovitskiy, and D. Duckworth · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Zson: Zero-shot object-goal navigation using multimodal goal embeddings
A. Majumdar, G. Aggarwal, B. Devnani, J. Hoffman, and D. Batra · 2022
Earlier work this paper cites.
Zero experience required: Plug & play modular transfer learning for semantic visual navigation
Z. Al-Halah, S. K. Ramakrishnan, and K. Grauman · 2022
Cited alongside, same era.
Autonomous exploration development environment and the planning algorithms
C. Cao, H. Zhu, F. Yang, Y. Xia, H. Choset, J. Oh, and J. Zhang · 2022
Cited alongside, same era.
Language-driven semantic segmentation
B. Li, K. Q. Weinberger, S. Belongie, V. Koltun, and R. Ranftl · 2022
Cited alongside, same era.
Detecting twenty-thousand classes using image-level supervision
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra · 2022
Cited alongside, same era.
Grounded language-image pre-training
L. H. Li*, P. Zhang*, H. Zhang*, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, K.-W. Chang, and J. Gao · 2022
Cited alongside, same era.
Glipv2: Unifying localization and vision-language understanding
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
C. Zhou, Q. Li, C. Li, J. Yu, Y. Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He, et al · 2023
Later among the works it cites.
A brief overview of chatgpt: The history, status quo and potential future development
T. Wu, S. He, J. Liu, S. Sun, K. Liu, Q.-L. Han, and Y. Tang · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Later among the works it cites.
A comprehensive capability analysis of gpt-3 and gpt-3.5 series models
J. Ye, X. Chen, N. Xu, C. Zu, Z. Shao, S. Liu, Y. Cui, Z. Zhou, C. Gong, Y. Shen, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Zhang, P. Zhang, X. Hu, Y.-C. Chen, L. H. Li, X. Dai, L. Wang, L. Yuan, J.-N. Hwang, and J. Gao · 2022
Cited alongside, same era.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, M. Attarian, K. M. Choromanski, A. Wong, S. Welker, F. Tombari, A. Purohit, M. S. Ryoo, V. Sindhwani, J. Lee, et al · 2022
Cited alongside, same era.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Cited alongside, same era.
Pddl planning with pretrained large language models
T. Silver, V. Hariprasad, R. S. Shuttleworth, N. Kumar, T. Lozano-Pérez, and L. P. Kaelbling · 2022
Cited alongside, same era.
Lens: Localization enhanced by nerf synthesis
A. Moreau, N. Piasco, D. Tsishkou, B. Stanciulescu, and A. de La Fortelle · 2022
Cited alongside, same era.
Nice-slam: Neural implicit scalable encoding for slam
Z. Zhu, S. Peng, V. Larsson, W. Xu, H. Bao, Z. Cui, M. R. Oswald, and M. Pollefeys · 2022
Cited alongside, same era.
Vox-fusion: Dense tracking and mapping with voxel-based neural implicit representation
X. Yang, H. Li, H. Zhai, Y. Ming, Y. Liu, and G. Zhang · 2022
Cited alongside, same era.
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Visual instruction tuning, 2023
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Later among the works it cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick · 2023
Later among the works it cites.
Y. Cheng, L. Li, Y. Xu, X. Li, Z. Yang, W. Wang, and Y. Yang · 2023
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al · 2023
Later among the works it cites.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
D. Shah, B. Osiński, S. Levine, et al · 2023
Later among the works it cites.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
C. H. Song, J. Wu, C. Washington, B. M. Sadler, W.-L. Chao, and Y. Su · 2023
Later among the works it cites.
Llm+ p: Empowering large language models with optimal planning proficiency
B. Liu, Y. Jiang, X. Zhang, Q. Liu, S. Zhang, J. Biswas, and P. Stone · 2023
Later among the works it cites.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Later among the works it cites.
3d gaussian splatting for real-time radiance field rendering
B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis · 2023
Later among the works it cites.
Nerf-pose: A first-reconstruct-then-regress approach for weakly-supervised 6d object pose estimation
F. Li, S. R. Vutukur, H. Yu, I. Shugurov, B. Busam, S. Yang, and S. Ilic · 2023
Later among the works it cites.
Graspnerf: Multiview-based 6-dof grasp detection for transparent and specular objects using generalizable nerf
Q. Dai, Y. Zhu, Y. Geng, C. Ruan, J. Zhang, and H. Wang · 2023
Later among the works it cites.
Open-vocabulary queryable scene representations for real world planning
B. Chen, F. Xia, B. Ichter, K. Rao, K. Gopalakrishnan, M. S. Ryoo, A. Stone, and D. Kappler · 2023
Later among the works it cites.
Ns3d: Neuro-symbolic grounding of 3d objects and relations
J. Hsu, J. Mao, and J. Wu · 2023
Later among the works it cites.
Representation granularity enables time-efficient autonomous exploration in large, complex worlds
C. Cao, H. Zhu, Z. Ren, H. Choset, and J. Zhang · 2023
Later among the works it cites.
Grounded sam: Assembling open-world models for diverse visual tasks, 2024
T. Ren, S. Liu, A. Zeng, J. Lin, K. Li, H. Cao, J. Chen, X. Huang, Y. Chen, F. Yan, Z. Zeng, H. Zhang, F. Li, J. Yang, H. Li, Q. Jiang, and L. Zhang · 2024
Closest in time.
Multimodal foundation models: From specialists to general-purpose assistants
C. Li, Z. Gan, Z. Yang, J. Yang, L. Li, L. Wang, J. Gao, et al · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
Vision-language models for vision tasks: A survey
J. Zhang, J. Huang, S. Jin, and S. Lu · 2024
Closest in time.
Depth anything: Unleashing the power of large-scale unlabeled data
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao · 2024
Closest in time.
Gaussian Splatting SLAM
H. Matsuki, R. Murai, P. H. J. Kelly, and A. J. Davison · 2024
Closest in time.
Manigaussian: Dynamic gaussian splatting for multi-task robotic manipulation
G. Lu, S. Zhang, Z. Wang, C. Liu, J. Lu, and Y. Tang · 2024
Closest in time.
Object-aware gaussian splatting for robotic manipulation
Y. Li and D. Pathak · 2024
Closest in time.
Lp-ovod: Open-vocabulary object detection by linear probing
C. Pham, T. Vu, and K. Nguyen · 2024
Closest in time.