Fetching the paper…
Reading the bibliography…
Embodied visual tracking is a fundamental skill in Embodied AI, enabling an agent to follow a specific target in dynamic environments using only egocentric vision.
Field studies of pedestrian walking speed and start-up time
R. L. Knoblauch, M. T. Pietrucha, and M. Nitzburg · 1996
Earlier work this paper cites.
k-means++: The advantages of careful seeding
D. Arthur and S. Vassilvitskii · 2006
Earlier work this paper cites.
Reciprocal n-body collision avoidance
J. Van Den Berg, S. J. Guy, M. Lin, and D. Manocha · 2011
Earlier work this paper cites.
A novel vision-based tracking algorithm for a human-following mobile robot
M. Gupta, S. Kumar, L. Behera, and V. K. Subramanian · 2016
Earlier work this paper cites.
Cross-stitch networks for multi-task learning
I. Misra, A. Shrivastava, A. Gupta, and M. Hebert · 2016
Earlier work this paper cites.
Unrealcv: Virtual worlds for computer vision
W. Qiu, F. Zhong, Y. Zhang, S. Qiao, Z. Xiao, T. S. Kim, and Y. Wang · 2017
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niebner, M. Savva, S. Song, A. Zeng, and Y. Zhang · 2017
Earlier work this paper cites.
Coarse-to-fine uav target tracking with deep reinforcement learning
W. Zhang, K. Song, X. Rong, and Y. Li · 2018
Earlier work this paper cites.
End-to-end active object tracking via reinforcement learning
W. Luo, P. Sun, F. Zhong, W. Liu, T. Zhang, and Y. Wang · 2018
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao · 2019
Earlier work this paper cites.
End-to-end active object tracking and its real-world deployment via reinforcement learning
W. Luo, P. Sun, F. Zhong, W. Liu, T. Zhang, and Y. Wang · 2019
Earlier work this paper cites.
Learning discriminative model prediction for tracking
G. Bhat, M. Danelljan, L. V. Gool, and R. Timofte · 2019
Earlier work this paper cites.
Pose-assisted multi-camera collaboration for active object tracking
J. Li, J. Xu, F. Zhong, X. Kong, Y. Qiao, and Y. Wang · 2020
Earlier work this paper cites.
Object goal navigation using goal-oriented semantic exploration
D. S. Chaplot, D. P. Gandhi, A. Gupta, and R. R. Salakhutdinov · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Earlier work this paper cites.
Enhancing continuous control of mobile robots for end-to-end visual active tracking
A. Devo, A. Dionigi, and G. Costante · 2021
Earlier work this paper cites.
Towards distraction-robust active visual tracking
F. Zhong, P. Sun, W. Luo, T. Yan, and Y. Wang · 2021
Earlier work this paper cites.
Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai
S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. M. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang, et al · 2021
Earlier work this paper cites.
A survey of embodied ai: From simulators to research tasks
J. Duan, S. Yu, H. L. Tan, H. Zhu, and C. Tan · 2022
Earlier work this paper cites.
Offline reinforcement learning for visual navigation
D. Shah, A. Bhorkar, H. Leen, I. Kostrikov, N. Rhinehart, and S. Levine · 2022
Earlier work this paper cites.
Rspt: reconstruct surroundings and predict trajectory for generalizable active object tracking
F. Zhong, X. Bi, Y. Zhang, W. Zhang, and Y. Wang · 2023
Earlier work this paper cites.
Core challenges of social robot navigation: A survey
C. Mavrogiannis, F. Baldini, A. Wang, D. Zhao, P. Trautman, A. Steinfeld, and J. Oh · 2023
Earlier work this paper cites.
Habitat 3.0: A co-habitat for humans, avatars and robots
X. Puig, E. Undersander, A. Szot, M. D. Cote, T.-Y. Yang, R. Partsey, R. Desai, A. W. Clegg, M. Hlavac, S. Y. Min, et al · 2023
Earlier work this paper cites.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Earlier work this paper cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, et al · 2023
Earlier work this paper cites.
Gnm: A general navigation model to drive any robot
D. Shah, A. Sridhar, A. Bhorkar, N. Hirose, and S. Levine · 2023
Earlier work this paper cites.
3d-aware object goal navigation via simultaneous exploration and identification
J. Zhang, L. Dai, F. Meng, Q. Fan, X. Chen, K. Xu, and H. Wang · 2023
Cited alongside, same era.
Eqa-mx: Embodied question answering using multimodal expression
M. M. Islam, A. Gladstone, R. Islam, and T. Iqbal · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, et al · 2023
Cited alongside, same era.
Llama-vid: An image is worth 2 tokens in large language models
Y. Li, C. Wang, and J. Jia · 2023
Cited alongside, same era.
Instructnav: Zero-shot system for generic instruction navigation in unexplored environment
Y. Long, W. Cai, H. Wang, G. Zhan, and H. Dong · 2024
Later among the works it cites.
Navid: Video-based vlm plans the next step for vision-and-language navigation
J. Zhang, K. Wang, R. Xu, G. Zhou, Y. Hong, X. Fang, Q. Wu, Z. Zhang, and H. Wang · 2024
Later among the works it cites.
Openfmnav: Towards open-set zero-shot object navigation via vision-language foundation models
Y. Kuang, H. Lin, and M. Jiang · 2024
Later among the works it cites.
Cognav: Cognitive process modeling for object goal navigation with llms
Y. Cao, J. Zhang, Z. Yu, S. Liu, Z. Qin, Q. Zou, B. Du, and K. Xu · 2024
Later among the works it cites.
Openeqa: Embodied question answering in the era of foundation models
A. Majumdar, A. Ajay, X. Zhang, P. Putta, S. Yenamandra, M. Henaff, S. Silwal, P. Mcvay, O. Maksymets, S. Arnaud, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eva-clip: Improved training techniques for clip at scale
Q. Sun, Y. Fang, L. Wu, X. Wang, and Y. Cao · 2023
Cited alongside, same era.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Cited alongside, same era.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao · 2023
Cited alongside, same era.
Lisa++: An improved baseline for reasoning segmentation with large language model
S. Yang, T. Qu, X. Lai, Z. Tian, B. Peng, S. Liu, and J. Jia · 2023
Cited alongside, same era.
Later among the works it cites.
Longvu: Spatiotemporal adaptive compression for long video-language understanding
X. Shen, Y. Xiong, C. Zhao, L. Wu, J. Chen, C. Zhu, Z. Liu, F. Xiao, B. Varadarajan, F. Bordes, et al · 2024
Later among the works it cites.
Paligemma 2: A family of versatile vlms for transfer
A. Steiner, A. S. Pinto, M. Tschannen, D. Keysers, X. Wang, Y. Bitton, A. Gritsenko, M. Minderer, A. Sherbondy, S. Long, et al · 2024
Later among the works it cites.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Later among the works it cites.
Q. Li, Y. Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei, Y. Deng, S. Xu, Y. Zhang, et al · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al · 2024
Later among the works it cites.
Open6dor: Benchmarking open-instruction 6-dof object rearrangement and a vlm-based approach
Y. Ding, H. Geng, C. Xu, X. Fang, J. Zhang, S. Wei, Q. Dai, Z. Zhang, and H. Wang · 2024
Later among the works it cites.
Navila: Legged robot vision-language-action model for navigation
A.-C. Cheng, Y. Ji, Z. Yang, Z. Gongye, X. Zou, J. Kautz, E. Bıyık, H. Yin, S. Liu, and X. Wang · 2024
Later among the works it cites.
Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving
B. Liao, S. Chen, H. Yin, B. Jiang, C. Wang, S. Yan, X. Zhang, X. Li, Y. Zhang, Q. Zhang, et al · 2024
Later among the works it cites.
Texdreamer: Towards zero-shot high-fidelity 3d human texture generation
Y. Liu, J. Zhu, J. Tang, S. Zhang, J. Zhang, W. Cao, C. Wang, Y. Wu, and D. Huang · 2024
Later among the works it cites.
Towards learning a generalist model for embodied navigation
D. Zheng, S. Huang, L. Zhao, Y. Zhong, and L. Wang · 2024
Later among the works it cites.
Viplanner: Visual semantic imperative learning for local navigation
P. Roth, J. Nubert, F. Yang, M. Mittal, and M. Hutter · 2024
Later among the works it cites.
Rpf-search: Field-based search for robot person following in unknown dynamic environments
H. Ye, K. Cai, Y. Zhan, B. Xia, A. Ajoudani, and H. Zhang · 2025
Closest in time.
Principles and guidelines for evaluating social robot navigation algorithms
A. Francis, C. Pérez-d’Arpino, C. Li, F. Xia, A. Alahi, R. Alami, A. Bera, A. Biswas, J. Biswas, R. Chandra, et al · 2025
Closest in time.
Navgpt-2: Unleashing navigational reasoning capability for large vision-language models
G. Zhou, Y. Hong, Z. Wang, X. E. Wang, and Q. Wu · 2025
Closest in time.
π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al · 2025
Closest in time.
Spatialvla: Exploring spatial representations for visual-language-action model
D. Qu, H. Song, Q. Chen, Y. Yao, X. Ye, Y. Ding, Z. Wang, J. Gu, B. Zhao, D. Wang, et al · 2025
Closest in time.
Dexgraspvla: A vision-language-action framework towards general dexterous grasping
Y. Zhong, X. Huang, R. Li, C. Zhang, Y. Liang, Y. Yang, and Y. Chen · 2025
Closest in time.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al · 2025
Closest in time.
Introducing 4o image generation
OpenAI · 2025
Closest in time.
Q. Jiang, L. Wu, Z. Zeng, T. Ren, Y. Xiong, Y. Chen, Q. Liu, and L. Zhang · 2025
Closest in time.
Multi-view spatial context and state constraints for object-goal navigation
C. Lu, M. Liu, Z. Luan, Y. He, and B. Chen · 2025
Closest in time.