Fetching the paper…
Reading the bibliography…
We address a practical yet challenging problem of training robot agents to navigate in an environment following a path described by some language instructions.
Bidirectional recurrent neural networks
M. Schuster and K. K. Paliwal · 1997
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, Ç. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Mutual information-based exploration on continuous occupancy maps
M. G. Jadidi, J. V. Miró, and G. Dissanayake · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Learning deep features for discriminative localization
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2016
Earlier work this paper cites.
Cognitive mapping and planning for visual navigation
S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik · 2017
Earlier work this paper cites.
Hide-and-seek: Forcing a network to be meticulous for weakly-supervised object and action localization
K. Kumar Singh and Y. Jae Lee · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Soft proposal networks for weakly supervised object localization
Y. Zhu, Y. Zhou, Q. Ye, Q. Qiu, and J. Jiao · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. D. Reid, S. Gould, and A. van den Hengel · 2018
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. D. Reid, S. Gould, and A. van den Hengel · 2018
Earlier work this paper cites.
Geometry-aware recurrent neural networks for active visual recognition
R. Cheng, Z. Wang, and K. Fragkiadaki · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell · 2018
Earlier work this paper cites.
Mapnet: An allocentric spatial memory for mapping environments
J. F. Henriques and A. Vedaldi · 2018
Earlier work this paper cites.
Semantic SLAM based on object detection and improved octomap
L. Zhang, L. Wei, P. Shen, W. Wei, G. Zhu, and J. Song · 2018
Earlier work this paper cites.
Adversarial complementary learning for weakly supervised object localization
X. Zhang, Y. Wei, J. Feng, Y. Yang, and T. S. Huang · 2018
Earlier work this paper cites.
Interpreting deep visual representations via network dissection
B. Zhou, D. Bau, A. Oliva, and A. Torralba · 2018
Earlier work this paper cites.
Progressive representation adaptation for weakly supervised object localization
D. Li, J.-B. Huang, Y. Li, S. Wang, and M.-H. Yang · 2019
Cited alongside, same era.
Self-monitoring navigation agent via auxiliary progress estimation
C. Ma, J. Lu, Z. Wu, G. AlRegib, Z. Kira, R. Socher, and C. Xiong · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Cited alongside, same era.
Habitat: A platform for embodied AI research
M. Savva, J. Malik, D. Parikh, D. Batra, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, and V. Koltun · 2019
Cited alongside, same era.
Learning to navigate unseen environments: Back translation with environmental dropout
H. Tan, L. Yu, and M. Bansal · 2019
Cited alongside, same era.
Semantic mapnet: Building allocentric semantic maps and representations from egocentric views
V. Cartillier, Z. Ren, N. Jain, S. Lee, I. Essa, and D. Batra · 2021
Later among the works it cites.
SEAL: self-supervised embodied active learning using exploration and 3d consistency
D. S. Chaplot, M. Dalal, S. Gupta, J. Malik, and R. Salakhutdinov · 2021
Later among the works it cites.
Topological planning with transformers for vision-and-language navigation
K. Chen, J. K. Chen, J. Chuang, M. Vázquez, and S. Savarese · 2021
Later among the works it cites.
Topological and semantic map generation for mobile robot indoor navigation
Y. Chen, J. Zhang, and Y. Lou · 2021
Later among the works it cites.
Threedworld: A platform for interactive multi-modal physical simulation
C. Gan, J. Schwartz, S. Alter, M. Schrimpf, J. Traer, J. De Freitas, J. Kubilius, A. Bhandwaldar, N. Haber, M. Sano, et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y. Wang, W. Y. Wang, and L. Zhang · 2019
Cited alongside, same era.
Danet: Divergent activation for weakly supervised object localization
H. Xue, C. Liu, F. Wan, J. Jiao, X. Ji, and Q. Ye · 2019
Cited alongside, same era.
Breaking winner-takes-all: Iterative-winners-out networks for weakly supervised temporal action localization
R. Zeng, C. Gan, P. Chen, W. Huang, Q. Wu, and M. Tan · 2019
Cited alongside, same era.
Learning to plan with uncertain topological maps
E. Beeching, J. Dibangoye, O. Simonin, and C. Wolf · 2020
Cited alongside, same era.
Object goal navigation using goal-oriented semantic exploration
D. S. Chaplot, D. Gandhi, A. Gupta, and R. R. Salakhutdinov · 2020
Cited alongside, same era.
Neural topological SLAM for visual navigation
D. S. Chaplot, R. Salakhutdinov, A. Gupta, and S. Gupta · 2020
Cited alongside, same era.
Look, listen, and act: Towards audio-visual embodied navigation
C. Gan, Y. Zhang, J. Wu, B. Gong, and J. B. Tenenbaum · 2020
Cited alongside, same era.
Airbert: In-domain pretraining for vision-and-language navigation
P.-L. Guhur, M. Tapaswi, S. Chen, I. Laptev, and C. Schmid · 2021
Later among the works it cites.
VLN BERT: A recurrent vision-and-language BERT for navigation
Y. Hong, Q. Wu, Y. Qi, C. R. Opazo, and S. Gould · 2021
Later among the works it cites.
M. Z. Irshad, N. C. Mithun, Z. Seymour, H. Chiu, S. Samarasekera, and R. Kumar · 2021
Later among the works it cites.
High-speed robot navigation using predicted occupancy maps
K. D. Katyal, A. Polevoy, J. L. Moore, C. Knuth, and K. M. Popek · 2021
Later among the works it cites.
Waypoint models for instruction-guided navigation in continuous environments
J. Krantz, A. Gokaslan, D. Batra, S. Lee, and O. Maksymets · 2021
Later among the works it cites.
Language-aligned waypoint (LAW) supervision for vision-and-language navigation in continuous environments
S. Raychaudhuri, S. Wani, S. Patel, U. Jain, and A. X. Chang · 2021
Later among the works it cites.
Learning active camera for multi-object navigatio
P. Chen, D. Ji, K. Lin, W. Hu, W. Huang, T. H. Li, M. Tan, and C. Gan · 2022
Closest in time.
Finding fallen objects via asynchronous audio-visual integration
C. Gan, Y. Gu, S. Zhou, J. Schwartz, S. Alter, J. Traer, D. Gutfreund, J. B. Tenenbaum, J. H. McDermott, and A. Torralba · 2022
Closest in time.
The threedworld transport challenge: A visually guided task-and-motion planning benchmark towards physically realistic embodied ai
C. Gan, S. Zhou, J. Schwartz, S. Alter, A. Bhandwaldar, D. Gutfreund, D. L. Yamins, J. J. DiCarlo, J. McDermott, A. Torralba, et al · 2022
Closest in time.
Learning to map for active semantic goal navigation
G. Georgakis, B. Bucher, K. Schmeckpeper, S. Singh, and K. Daniilidis · 2022
Closest in time.
Cross-modal map learning for vision and language navigation
G. Georgakis, K. Schmeckpeper, K. Wanchoo, S. Dan, E. Miltsakaki, D. Roth, and K. Daniilidis · 2022
Closest in time.
Cross-modal map learning for vision and language navigation
G. Georgakis, K. Schmeckpeper, K. Wanchoo, S. Dan, E. Miltsakaki, D. Roth, and K. Daniilidis · 2022
Closest in time.
Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation
Y. Hong, Z. Wang, Q. Wu, and S. Gould · 2022
Closest in time.
Shape or texture: Understanding discriminative features in cnns
M. A. Islam, M. Kowal, P. Esser, S. Jia, B. Ommer, K. G. Derpanis, and N. Bruce · 2022
Closest in time.