Fetching the paper…
Reading the bibliography…
Robotic manipulation in complex scenes demands precise perception of task-relevant details, yet fixed or suboptimal viewpoints often impair fine-grained perception and induce occlusions, constraining imitation-learned policies.
S. Ross, G. Gordon, and D. Bagnell, “A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,” in Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (AISTATS) , vol. 15, pp. 627–635, 2011
2011
Earlier work this paper cites.
G. Casiez, N. Roussel, and D. Vogel, “1€ Filter: A Simple Speed-based Low-pass Filter for Noisy Input in Interactive Systems,” in Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI) , pp. 2527–2530, 2012
2012
Earlier work this paper cites.
J. Ho and S. Ermon, “Generative Adversarial Imitation Learning,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 29, 2016
2016
Earlier work this paper cites.
J. Fu, H. Zheng, and T. Mei, “Look Closer to See Better: Recurrent Attention Convolutional Neural Network for Fine-Grained Image Recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 4438–4446, 2017
2017
Earlier work this paper cites.
Z. Wang, Y. Yin, J. Shi, et al. , “Zoom-in-Net: Deep Mining Lesions for Diabetic Retinopathy Detection,” in Proceedings of International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI) , pp. 267-275, 2017
2017
Earlier work this paper cites.
T. Osa, J. Pajarinen, G. Neumann, et al. , “An Algorithmic Perspective on Imitation Learning,” Foundations and Trends® in Robotics , vol. 7, pp. 1–179, 2018
2018
Earlier work this paper cites.
Z. Yang, T. Luo, D. Wang, et al. , “Learning to Navigate for Fine-grained Classification,” in Proceedings of the European Conference on Computer Vision (ECCV) , pp. 420-435, 2018
2018
Earlier work this paper cites.
D. Rakita, B. Mutlu, and M. Gleicher, “Remote Telemanipulation with Adapting Viewpoints in Visually Complex Environments,” Robotics: Science and Systems XV , 2019
2019
Earlier work this paper cites.
D. Morrison, P. Corke, and J. Leitner, “Multi-View Picking: Next-best-view Reaching for Improved Grasping in Clutter,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) , pp. 8762-8768, 2019
2019
Earlier work this paper cites.
R. Zeng, Y. Wen, W. Zhao, et al. , “View planning in robot active vision: A survey of systems, algorithms, and applications,” Computational Visual Media , vol. 6, pp. 225–245, 2020
2020
Earlier work this paper cites.
Nakanishi, Jun, Itadera, Shunki, Aoyama, Tadayoshi, et al. , “Towards the development of an intuitive teleoperation system for human support robot using a VR device.” Advanced Robotics , vol. 34, pp. 1239-1253, 2020
2020
Earlier work this paper cites.
M. Schwarz and S. Behnke, “Low-Latency Immersive 6D Televisualization with Spherical Rendering,” in Proceedings of the IEEE-RAS International Conference on Humanoid Robots (Humanoids) , pp. 320-325, 2020
2020
Cited alongside, same era.
X. Wang, L. Xie, C. Dong, et al. , “Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops , pp. 1905-1914, 2021
2021
Cited alongside, same era.
C. Chi, Z. Xu, S. Feng, et al. , “Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,” The International Journal of Robotics Research (IJRR) , pp. 02783649241273668, 2023
2023
Cited alongside, same era.
M. Mittal, C. Yu, Q. Yu, et al. , “Orbit: A Unified Simulation Framework for Interactive Robot Learning Environments,” IEEE Robotics and Automation Letters , vol. 8, pp. 3740–3747, 2023
2023
Cited alongside, same era.
X. Cheng, J. Li, S. Yang, et al. , “Open-TeleVision: Teleoperation with Immersive Active Visual Feedback,” arXiv, 2024
2024
Later among the works it cites.
A. Khazatsky, K. Pertsch, S. Nair, et al. , “DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset,” arXiv, 2024
2024
Later among the works it cites.
T.-C. Lin, A. U. Krishnan, and Z. Li, “erception and Action Augmentation for Teleoperation Assistance in Freeform Telemanipulation,” ACM Transactions on Human-Robot Interaction , vol. 13, pp. 1–40, 2024
2024
Later among the works it cites.
I. Chuang, A. Lee, D. Gao, et al. , “Active Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic Manipulation,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) , pp. 7952-7959, 2025
2025
Closest in time.
H. Xiong, X. Xu, J. Wu, et al. , “Vision in Action: Learning Active Perception from Human Demonstrations,” arXiv, 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Wang, Z. Xian, F. Chen, et al. , “RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation,” arXiv, 2023
2023
Cited alongside, same era.
M. Shridhar, L. Manuelli, and D. Fox, “Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation,” in Conference on Robot Learning (CoRL) , pp. 785–799, 2023
2023
Cited alongside, same era.
J. Shang and M. S. Ryoo, “Active Vision Reinforcement Learning under Limited Visual Observability,” in Advances in Neural Information Processing Systems (NeurIPS) , vol.29, pp.10316-10338, 2023
2023
Cited alongside, same era.
Q. Feng, X. Xu, and Z. Wang, “Deep learning-based small object detection: A survey,” Mathematical Biosciences and Engineering , vol. 20, pp. 6551–6590, 2023
2023
Cited alongside, same era.
M. Oquab, T. Darcet, T. Moutakanni, et al. , “DINOv2: Learning Robust Visual Features without Supervision,” arXiv, 2023
2023
Cited alongside, same era.
Y. Mu, T. Chen, S. Peng, et al. , “RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (Early Version),” arXiv, 2024
2024
Cited alongside, same era.
2025
Closest in time.
Q. Bu, J. Cai, L. Chen, et al. , “Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems,” arXiv, 2025
2025
Closest in time.
T. Chen, Z. Chen, B. Chen, et al. , “RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation,” arXiv, 2025
2025
Closest in time.
G. Wang, H. Li, S. Zhang, et al. , “Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation,” IEEE Robotics and Automation Letters , vol. 10, pp. 3422-3429, 2025
2025
Closest in time.
I. Chuang, A. Lee, D. Gao, et al. , “Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers,” arXiv, 2025
2025
Closest in time.
“Galaxea A1.” https://github.com/userguide-galaxea/A1_SDK . Accessed: Sep. 14, 2025
2025
Closest in time.