Fetching the paper…
Reading the bibliography…
Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information.
Active vision
Aloimonos, J., Weiss, I., and Bandyopadhyay, A · 1988
Earlier work this paper cites.
Animate vision
Ballard, D. H · 1991
Earlier work this paper cites.
Autonomous exploration: Driven by uncertainty
Whaite, P. and Ferrie, F. P · 1997
Earlier work this paper cites.
Active vision for dexterous grasping of novel objects
Arruda, E., Wyatt, J., and Kopicki, M · 2016
Earlier work this paper cites.
Neural modular control for embodied question answering
Das, A., Gkioxari, G., Lee, S., Parikh, D., and Batra, D · 2018
Earlier work this paper cites.
Learning to look around: Intelligently exploring unseen environments for unknown tasks
Jayaraman, D. and Grauman, K · 2018
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
Gupta, A., Dollar, P., and Girshick, R · 2019
Earlier work this paper cites.
Learning to explore using active neural slam
Chaplot, D. S., Gandhi, D., Gupta, S., Gupta, A., and Salakhutdinov, R · 2020
Earlier work this paper cites.
Deep interactive thin object selection
Liew, J. H., Cohen, S., Price, B., Mai, L., and Feng, J · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Towards large-scale small object detection: Survey and benchmarks
Cheng, G., Yuan, X., Yao, X., Yan, K., Zeng, Q., Xie, X., and Han, J · 2023
Earlier work this paper cites.
pi0: A vision-language-action flow model for general robot control
Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., et al · 2024
Earlier work this paper cites.
Visual chain-of-thought prompting for knowledge-based visual reasoning
Chen, Z., Zhou, Q., Shen, Y., Hong, Y., Sun, Z., Gutfreund, D., and Gan, C · 2024
Earlier work this paper cites.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al · 2024
Cited alongside, same era.
Openvla: An open-source vision-language-action model
Kim, M. J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., et al · 2024
Cited alongside, same era.
Sam 2: Segment anything in images and videos
Ravi, N., Gabeur, V., Hu, Y.-T., Hu, R., Ryali, C., Ma, T., Khedr, H., Rädle, R., Rolland, C., Gustafson, L., et al · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al · 2024
Cited alongside, same era.
Visual-rft: Visual reinforcement fine-tuning
Liu, Z., Sun, Z., Zang, Y., Dong, X., Cao, Y., Duan, H., Lin, D., and Wang, J · 2025
Closest in time.
Argus: Vision-centric reasoning with grounded chain-of-thought
Man, Y., Huang, D.-A., Liu, G., Sheng, S., Liu, S., Gui, L.-Y., Kautz, J., Wang, Y.-X., and Yu, Z · 2025
Closest in time.
Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation
Qi, Z., Zhang, W., Ding, Y., Dong, R., Yu, X., Li, J., Xu, L., Li, B., He, X., Fan, G., et al · 2025
Closest in time.
Zoomeye: Enhancing multimodal llms with human-like zooming capabilities through tree-based image exploration
Shen, H., Zhao, K., Zhao, T., Xu, R., Zhang, Z., Zhu, M., and Yin, J · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Contrastive region guidance: Improving grounding in vision-language models without training
Wan, D., Cho, J., Stengel-Eskin, E., and Bansal, M · 2024
Cited alongside, same era.
V*: Guided visual search as a core mechanism in multimodal llms
Wu, P. and Xie, S · 2024
Cited alongside, same era.
A survey on multimodal large language models
Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., and Chen, E · 2024
Cited alongside, same era.
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., et al · 2025
Cited alongside, same era.
Grit: Teaching mllms to think with images
Fan, Y., He, X., Yang, D., Zheng, K., Kuo, C.-C., Zheng, Y., Narayanaraju, S. J., Guan, X., and Wang, X. E · 2025
Cited alongside, same era.
Video-r1: Reinforcing video reasoning in mllms, 2025
Feng, K., Gong, K., Li, B., Guo, Z., Wang, Y., Peng, T., Wu, J., Zhang, X., Wang, B., and Yue, X · 2025
Cited alongside, same era.
Refocus: Visual editing as a chain of thought for structured image understanding
Fu, X., Liu, M., Yang, Z., Corring, J., Lu, Y., Yang, J., Roth, D., Florencio, D., and Zhang, C · 2025
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Cited alongside, same era.
Team, G. R., Abeyruwan, S., Ainslie, J., Alayrac, J.-B., Arenas, M. G., Armstrong, T., Balakrishna, A., Baruch, R., Bauza, M., Blokzijl, M., et al · 2025
Closest in time.
Grok-1.5 vision preview, 2024
xAI · 2025
Closest in time.
Magma: A foundation model for multimodal ai agents
Yang, J., Tan, R., Wu, Q., Zheng, R., Peng, B., Liang, Y., Gu, Y., Cai, M., Ye, S., Jang, J., et al · 2025
Closest in time.
Thinking in 360 ∘ : Humanoid visual search in the wild
Yu, H., Han, Y., Zhang, X., Yin, B., Chang, B., Han, X., Liu, X., Zhang, J., Pavone, M., Feng, C., et al · 2025
Closest in time.
R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning, 2025
Zhao, J., Wei, X., and Bo, L · 2025
Closest in time.
Deepeyes: Incentivizing” thinking with images” via reinforcement learning
Zheng, Z., Yang, M., Hong, J., Zhao, C., Xu, G., Yang, L., Shen, C., and Yu, X · 2025
Closest in time.
Zhu, M., Tian, Y., Chen, H., Zhou, C., Guo, Q., Liu, Y., Yang, M., and Shen, C · 2025
Closest in time.
Where to look: Can foundation models reach a target viewpoint through active exploration?
Li, L., Zhu, M., Zhao, Z., Zhao, H., Liu, K., Zhong, L., Chen, H., and Shen, C · 2026
Closest in time.
Exploring spatial intelligence from a generative perspective
Zhu, M., Jiang, S., Zheng, H., Luo, Z., Zhong, H., Li, A., Wang, K., Rong, J., Liu, Y., Chen, H., et al · 2026
Closest in time.