Fetching the paper…
Reading the bibliography…
Existing research studies on vision and language grounding for robot navigation focus on improving model-free deep reinforcement learning (DRL) models in synthetic environments.
Borenstein, J., Koren, Y.: Real-time obstacle avoidance for fast mobile robots. IEEE Transactions on systems, Man, and Cybernetics 19
1989
Earlier work this paper cites.
Sutton, R.S.: Integrated architectures for learning, planning, and reacting based on approximating dynamic programming. In: Machine Learning Proceedings 1990, pp. 216–224. Elsevier (1990)
1990
Earlier work this paper cites.
Borenstein, J., Koren, Y.: The vector field histogram-fast obstacle avoidance for mobile robots. IEEE transactions on robotics and automation 7
1991
Earlier work this paper cites.
Oriolo, G., Vendittelli, M., Ulivi, G.: On-line map building and navigation for autonomous mobile robots. In: Robotics and Automation, 1995. Proceedings., 1995 IEEE International Conference on. vol. 3, pp. 2900–2906. IEEE (1995)
1995
Earlier work this paper cites.
Kim, D., Nevatia, R.: Symbolic navigation with a generic map. Autonomous Robots 6
1999
Earlier work this paper cites.
Yao, H., Bhatnagar, S., Diao, D., Sutton, R.S., Szepesvári, C.: Multi-step dyna planning for policy evaluation and control. In: Advances in Neural Information Processing Systems. pp. 2187–2195 (2009)
2009
Earlier work this paper cites.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., Parikh, D.: Vqa: Visual question answering. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 2425–2433 (2015)
2015
Earlier work this paper cites.
Chen, X., Lawrence Zitnick, C.: Mind’s eye: A recurrent visual representation for image caption generation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 2422–2431 (2015)
2015
Earlier work this paper cites.
Karpathy, A., Fei-Fei, L.: Deep visual-semantic alignments for generating image descriptions. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3128–3137 (2015)
2015
Earlier work this paper cites.
Lenz, I., Knepper, R.A., Saxena, A.: Deepmpc: Learning deep latent features for model predictive control. In: Robotics: Science and Systems (2015)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Talvitie, E.: Agnostic system identification for monte carlo planning. In: AAAI. pp. 2986–2992 (2015)
2015
Earlier work this paper cites.
Vinyals, O., Toshev, A., Bengio, S., Erhan, D.: Show and tell: A neural image caption generator. In: Computer Vision and Pattern Recognition (CVPR), 2015 IEEE Conference on. pp. 3156–3164. IEEE (2015)
2015
Earlier work this paper cites.
Watter, M., Springenberg, J., Boedecker, J., Riedmiller, M.: Embed to control: A locally linear latent dynamics model for control from raw images. In: Advances in neural information processing systems. pp. 2746–2754 (2015)
2015
Cited alongside, same era.
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., Bengio, Y.: Show, attend and tell: Neural image caption generation with visual attention. In: International Conference on Machine Learning. pp. 2048–2057 (2015)
2015
Cited alongside, same era.
2016
Cited alongside, same era.
Gu, S., Lillicrap, T., Sutskever, I., Levine, S.: Continuous deep q-learning with model-based acceleration. In: International Conference on Machine Learning. pp. 2829–2838 (2016)
2016
Cited alongside, same era.
Finn, C., Levine, S.: Deep visual foresight for planning robot motion. In: Robotics and Automation (ICRA), 2017 IEEE International Conference on. pp. 2786–2793. IEEE (2017)
2017
Later among the works it cites.
2017
Later among the works it cites.
Oh, J., Singh, S., Lee, H.: Value prediction network. In: Advances in Neural Information Processing Systems. pp. 6120–6130 (2017)
2017
Later among the works it cites.
Pathak, D., Agrawal, P., Efros, A.A., Darrell, T.: Curiosity-driven exploration by self-supervised prediction. In: International Conference on Machine Learning (ICML). vol. 2017 (2017)
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770–778 (2016)
2016
Cited alongside, same era.
Huang, T.H.K., Ferraro, F., Mostafazadeh, N., Misra, I., Agrawal, A., Devlin, J., Girshick, R., He, X., Kohli, P., Batra, D., et al.: Visual storytelling. In: Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 1233–1239 (2016)
2016
Cited alongside, same era.
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., Jaśkowski, W.: Vizdoom: A doom-based ai research platform for visual reinforcement learning. In: Computational Intelligence and Games (CIG), 2016 IEEE Conference on. pp. 1–8. IEEE (2016)
2016
Cited alongside, same era.
Mei, H., Bansal, M., Walter, M.R.: Listen, attend, and walk: Neural mapping of navigational instructions to action sequences. In: AAAI. vol. 1, p. 2 (2016)
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Tamar, A., Wu, Y., Thomas, G., Levine, S., Abbeel, P.: Value iteration networks. In: Advances in Neural Information Processing Systems. pp. 2154–2162 (2016)
2016
Cited alongside, same era.
Yu, H., Wang, J., Huang, Z., Yang, Y., Xu, W.: Video paragraph captioning using hierarchical recurrent neural networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 4584–4593 (2016)
2016
Cited alongside, same era.
Alomari, M., Duckworth, P., Hawasly, M., Hogg, D.C., Cohn, A.G.: Natural language grounding and grammar induction for robotic manipulation commands. In: Proceedings of the First Workshop on Language Grounding for Robotics. pp. 35–43 (2017)
2017
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
Zhu, Y., Mottaghi, R., Kolve, E., Lim, J.J., Gupta, A., Fei-Fei, L., Farhadi, A.: Target-driven visual navigation in indoor scenes using deep reinforcement learning. In: Robotics and Automation (ICRA), 2017 IEEE International Conference on. pp. 3357–3364. IEEE (2017)
2017
Later among the works it cites.
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., van den Hengel, A.: Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). vol. 2 (2018)
2018
Closest in time.
Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., Batra, D.: Embodied Question Answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018)
2018
Closest in time.
Wang, X., Chen, W., Wang, Y.F., Wang, W.Y.: No metrics are perfect: Adversarial reward learning for visual storytelling. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 899–909. Association for Computational Linguistics (2018)
2018
Closest in time.
Wang, X., Chen, W., Wu, J., Wang, Y.F., Wang, W.Y.: Video captioning via hierarchical reinforcement learning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 4213–4222 (2018)
2018
Closest in time.
Wang, X., Wang, Y.F., Wang, W.Y.: Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning. In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers). pp. 795–801. Association for Computational Linguistics (2018)
2018
Closest in time.