Fetching the paper…
Reading the bibliography…
Grounding language to the visual observations of a navigating agent can be performed using off-the-shelf visual-language models pretrained on Internet-scale data (e.g., image captions).
T. P. McNamara, J. K. Hardy, and S. C. Hirtle, “Subjective hierarchies in spatial memory.” Journal of Experimental Psychology: Learning, Memory, and Cognition , vol. 15, no. 2, p. 211, 1989
1989
Earlier work this paper cites.
M. M. Chun and Y. Jiang, “Contextual cueing: Implicit learning and memory of visual context guides spatial attention,” Cognitive psychology , vol. 36, no. 1, pp. 28–71, 1998
1998
Earlier work this paper cites.
S. Thrun, W. Burgard, and D. Fox, “A probabilistic approach to concurrent mapping and localization for mobile robots,” Autonomous Robots , vol. 5, no. 3, pp. 253–271, 1998
1998
Earlier work this paper cites.
M. MacMahon, B. Stankiewicz, and B. Kuipers, “Walk the talk: Connecting language, knowledge, and action in route instructions,” Def , vol. 2, no. 6, p. 4, 2006
2006
Earlier work this paper cites.
E. L. Newman, J. B. Caplan, M. P. Kirschen, I. O. Korolev, R. Sekuler, and M. J. Kahana, “Learning your way around town: How virtual taxicab drivers learn to use both layout and landmark information,” Cognition , vol. 104, no. 2, pp. 231–253, 2007
2007
Earlier work this paper cites.
S. Tellex, T. Kollar, S. Dickerson, M. Walter, A. Banerjee, S. Teller, and N. Roy, “Understanding natural language commands for robotic navigation and mobile manipulation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 25, no. 1, 2011, pp. 1507–1514
2011
Earlier work this paper cites.
F. Endres, J. Hess, N. Engelhard, J. Sturm, D. Cremers, and W. Burgard, “An evaluation of the rgb-d slam system,” in 2012 IEEE international conference on robotics and automation . IEEE, 2012, pp. 1691–1696
2012
Earlier work this paper cites.
R. F. Salas-Moreno, R. A. Newcombe, H. Strasdat, P. H. Kelly, and A. J. Davison, “Slam++: Simultaneous localisation and mapping at the level of objects,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2013, pp. 1352–1359
2013
Earlier work this paper cites.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3431–3440
2015
Earlier work this paper cites.
J. McCormac, A. Handa, A. Davison, and S. Leutenegger, “Semanticfusion: Dense 3d semantic mapping with convolutional neural networks,” in 2017 IEEE International Conference on Robotics and automation (ICRA) . IEEE, 2017, pp. 4628–4635
2017
Earlier work this paper cites.
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang, “Matterport3D: Learning from RGB-D data in indoor environments,” International Conference on 3D Vision (3DV) , 2017
2017
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3674–3683
2018
Earlier work this paper cites.
M. Runz, M. Buffier, and L. Agapito, “Maskfusion: Real-time recognition, tracking and reconstruction of multiple moving objects,” in 2018 IEEE International Symposium on Mixed and Augmented Reality (ISMAR) . IEEE, 2018, pp. 10–20
2018
Earlier work this paper cites.
J. McCormac, R. Clark, M. Bloesch, A. Davison, and S. Leutenegger, “Fusion++: Volumetric object-level slam,” in 2018 international conference on 3D vision (3DV) . IEEE, 2018, pp. 32–41
2018
Earlier work this paper cites.
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell, “Speaker-follower models for vision-and-language navigation,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
B. Xu, W. Li, D. Tzoumanikas, M. Bloesch, A. Davison, and S. Leutenegger, “Mid-fusion: Octree-based object-level multi-instance dynamic slam,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 5231–5237
2019
Earlier work this paper cites.
A. Zeng, “Learning visual affordances for robotic manipulation,” Ph.D. dissertation, Princeton University, 2019
2019
Cited alongside, same era.
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, D. Parikh, and D. Batra, “Habitat: A Platform for Embodied AI Research,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2019
2019
Cited alongside, same era.
M. Labbé and F. Michaud, “Rtab-map as an open-source lidar and visual simultaneous localization and mapping library for large-scale and long-term online operation,” Journal of Field Robotics , vol. 36, no. 2, pp. 416–446, 2019
2019
Cited alongside, same era.
J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee, “Beyond the nav-graph: Vision-and-language navigation in continuous environments,” in European Conference on Computer Vision . Springer, 2020, pp. 104–120
2020
Cited alongside, same era.
2022
Closest in time.
2022
Closest in time.
N. Hughes, Y. Chang, and L. Carlone, “Hydra: a real-time spatial perception system for 3d scene graph construction and optimization,” Proceedings of Robotics: Science and Systems. New York City, NY, USA, http://dx. doi. org/10.15607/RSS , 2022
2022
Closest in time.
Y. Hong, Z. Wang, Q. Wu, and S. Gould, “Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 439–15 449
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” International Journal of Computer Vision , vol. 128, no. 2, pp. 336–359, 2020
2020
Cited alongside, same era.
P. Anderson, A. Shrivastava, J. Truong, A. Majumdar, D. Parikh, D. Batra, and S. Lee, “Sim-to-real transfer for vision-and-language navigation,” in Conference on Robot Learning . PMLR, 2021, pp. 671–681
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 8748–8763
2021
Cited alongside, same era.
B. Li, K. Q. Weinberger, S. Belongie, V. Koltun, and R. Ranftl, “Language-driven semantic segmentation,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
S.-C. Wu, J. Wald, K. Tateno, N. Navab, and F. Tombari, “Scenegraphfusion: Incremental 3d scene graph prediction from rgb-d sequences,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 7515–7525
2021
Cited alongside, same era.
P.-L. Guhur, M. Tapaswi, S. Chen, I. Laptev, and C. Schmid, “Airbert: In-domain pretraining for vision-and-language navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1634–1643
2021
Cited alongside, same era.
J. Krantz, A. Gokaslan, D. Batra, S. Lee, and O. Maksymets, “Waypoint models for instruction-guided navigation in continuous environments,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 15 162–15 171
2021
Cited alongside, same era.
2022
Closest in time.
M. Shridhar, L. Manuelli, and D. Fox, “Cliport: What and where pathways for robotic manipulation,” in Conference on Robot Learning . PMLR, 2022, pp. 894–906
2022
Closest in time.
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” IEEE Robotics and Automation Letters (RA-L) , vol. 7, no. 3, pp. 7327–7334, 2022
2022
Closest in time.
O. Mees, L. Hermann, and W. Burgard, “What matters in language conditioned robotic imitation learning over unstructured data,” IEEE Robotics and Automation Letters (RA-L) , vol. 7, no. 4, pp. 11 205–11 212, 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
K. Zakka, A. Zeng, P. Florence, J. Tompson, J. Bohg, and D. Dwibedi, “Xirl: Cross-embodiment inverse reinforcement learning,” in Conference on Robot Learning . PMLR, 2022, pp. 537–546
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
J. Borja-Diaz, O. Mees, G. Kalweit, L. Hermann, J. Boedecker, and W. Burgard, “Affordance learning from play for sample-efficient policy learning,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) , Philadelphia, USA, 2022
2022
Closest in time.