Fetching the paper…
Reading the bibliography…
The vision-language navigation (VLN) task requires an agent to reach a target with the guidance of natural language instruction.
Learning video representations using contrastive bidirectional transformer
Sun, C.; Baradel, F.; Murphy, K.; and Schmid, C. 2019 · 1906
Earlier work this paper cites.
Dimensionality Reduction by Learning an Invariant Mapping
Hadsell, R.; Chopra, S.; and Lecun, Y. 2006 · 2006
Earlier work this paper cites.
Distance Metric Learning for Large Margin Nearest Neighbor Classification
Weinberger, K. Q.; Blitzer, J.; and Saul, L. 2006 · 2006
Earlier work this paper cites.
Discriminative Unsupervised Feature Learning with Convolutional Neural Networks
Dosovitskiy, A.; Springenberg, J. T.; Riedmiller, M.; and Brox, T. 2014 · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
Sutskever, I.; Vinyals, O.; and Le, Q. V. 2014 · 2014
Earlier work this paper cites.
Character-level Convolutional Networks for Text Classification
Zhang, X.; Zhao, J.; and LeCun, Y. 2015 · 2015
Earlier work this paper cites.
Asynchronous Methods for Deep Reinforcement Learning
Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016 · 2016
Earlier work this paper cites.
Improved Deep Metric Learning with Multi-class N-pair Loss Objective
Sohn, K. 2016 · 2016
Earlier work this paper cites.
Deep Metric Learning via Lifted Structured Feature Embedding
Song, H. O.; Xiang, Y.; Jegelka, S.; and Savarese, S. 2016 · 2016
Earlier work this paper cites.
Deep Metric Learning With Angular Loss
Wang, J.; Zhou, F.; Wen, S.; Liu, X.; and Lin, Y. 2017 · 2017
Earlier work this paper cites.
Speaker-Follower Models for Vision-and-Language Navigation
Fried, D.; Hu, R.; Cirik, V.; Rohrbach, A.; Andreas, J.; Morency, L.-P.; Berg-Kirkpatrick, T.; Saenko, K.; Klein, D.; and Darrell, T. 2018 · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d.; Li, Y.; and Vinyals, O. 2018 · 2018
Earlier work this paper cites.
Unsupervised Feature Learning via Non-Parametric Instance Discrimination
Wu, Z.; Xiong, Y.; Yu, S. X.; and Lin, D. 2018 · 2018
Earlier work this paper cites.
Learning Representations by Maximizing Mutual Information Across Views
Bachman, P.; Hjelm, R. D.; and Buchwalter, W. 2019 · 2019
Earlier work this paper cites.
TOUCHDOWN: Natural Language Navigation and Spatial Reasoning in Visual Street Environments
Chen, H.; Suhr, A.; Misra, D.; Snavely, N.; and Artzi, Y. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. N. 2019 · 2019
Earlier work this paper cites.
Learning deep representations by mutual information estimation and maximization
Hjelm, R. D.; Fedorov, A.; Lavoie-Marchildon, S.; Grewal, K.; Bachman, P.; Trischler, A.; and Bengio, Y. 2019 · 2019
Earlier work this paper cites.
Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation
Jain, V.; Magalhães, G.; Ku, A.; Vaswani, A.; Ie, E.; and Baldridge, J. 2019 · 2019
Earlier work this paper cites.
Tactical Rewind: Self-Correction via Backtracking in Vision-And-Language Navigation
Ke, L.; Li, X.; Bisk, Y.; Holtzman, A.; Gan, Z.; Liu, J.; Gao, J.; Choi, Y.; and Srinivasa, S. 2019 · 2019
Cited alongside, same era.
Perceive, Transform, and Act: Multi-Modal Attention Networks for Vision-and-Language Navigation
Landi, F.; Baraldi, L.; Cornia, M.; Corsini, M.; and Cucchiara, R. 2019 · 2019
Cited alongside, same era.
Robust Navigation with Language Pretraining and Stochastic Sampling
Li, X.; Li, C.; Xia, Q.; Bisk, Y.; Çelikyilmaz, A.; Gao, J.; Smith, N. A.; and Choi, Y. 2019 · 2019
Cited alongside, same era.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Cited alongside, same era.
General Evaluation for Instruction Conditioned Navigation using Dynamic Time Warping
Magalhaes, G. I.; Jain, V.; Ku, A.; Ie, E.; and Baldridge, J. 2019 · 2019
Cited alongside, same era.
Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-Training
Hao, W.; Li, C.; Li, X.; Carin, L.; and Gao, J. 2020 · 2020
Later among the works it cites.
Momentum Contrast for Unsupervised Visual Representation Learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020 · 2020
Later among the works it cites.
Data-Efficient Image Recognition with Contrastive Predictive Coding
Henaff, O. 2020 · 2020
Later among the works it cites.
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
Ku, A.; Anderson, P.; Patel, R.; Ie, E.; and Baldridge, J. 2020 · 2020
Later among the works it cites.
Data Augmentation using Pre-trained Transformer Models
Kumar, V.; Choudhary, A.; and Cho, E. 2020 · 2020
Later among the works it cites.
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Help, Anna! Visual Navigation with Natural Multimodal Assistance via Retrospective Curiosity-Encouraging Imitation Learning
Nguyen, K.; and Daumé III, H. 2019 · 2019
Cited alongside, same era.
Vision-Based Navigation With Language-Based Assistance via Imitation Learning With Indirect Intervention
Nguyen, K.; Dey, D.; Brockett, C.; and Dolan, B. 2019 · 2019
Cited alongside, same era.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Tan, H.; and Bansal, M. 2019 · 2019
Cited alongside, same era.
Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout
Tan, H.; Yu, L.; and Bansal, M. 2019 · 2019
Cited alongside, same era.
Vision-and-Dialog Navigation
Thomason, J.; Murray, M.; Cakmak, M.; and Zettlemoyer, L. 2019 · 2019
Cited alongside, same era.
Deep Metric Learning With Tuplet Margin Loss
Yu, B.; and Tao, D. 2019 · 2019
Cited alongside, same era.
Do Not Have Enough Data? Deep Learning to the Rescue!
Anaby-Tavor, A.; Carmeli, B.; Goldbraich, E.; Kantor, A.; Kour, G.; Shlomov, S.; Tepper, N.; and Zwerdling, N. 2020 · 2020
Cited alongside, same era.
Li, X.; Yin, X.; Li, C.; Hu, X.; Zhang, P.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; Choi, Y.; and Gao, J. 2020 · 2020
Later among the works it cites.
Improving Vision-and-Language Navigation with Image-Text Pairs from the Web
Majumdar, A.; Shrivastava, A.; Lee, S.; Anderson, P.; Parikh, D.; and Batra, D. 2020 · 2020
Later among the works it cites.
https://www.mindspore.cn/
MindSpore. 2020 · 2020
Later among the works it cites.
Self-Supervised Learning of Pretext-Invariant Representations
Misra, I.; and van der Maaten, L. 2020 · 2020
Later among the works it cites.
Circle Loss: A Unified Perspective of Pair Similarity Optimization
Sun, Y.; Cheng, C.; Zhang, Y.; Zhang, C.; Zheng, L.; Wang, Z.; and Wei, Y. 2020 · 2020
Later among the works it cites.
Soft Expert Reward Learning for Vision-and-Language Navigation
Wang, H.; Wu, Q.; and Shen, C. 2020 · 2020
Later among the works it cites.
Unsupervised Data Augmentation for Consistency Training
Xie, Q.; Dai, Z.; Hovy, E.; Luong, T.; and Le, Q. 2020 · 2020
Later among the works it cites.
PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding
Xie, S.; Gu, J.; Guo, D.; Qi, C. R.; Guibas, L. J.; and Litany, O. 2020 · 2020
Later among the works it cites.
Exploring Simple Siamese Representation Learning
Chen, X.; and He, K. 2021 · 2021
Closest in time.
A Recurrent Vision-and-Language BERT for Navigation
Hong, Y.; Wu, Q.; Qi, Y.; Rodriguez-Opazo, C.; and Gould, S. 2021 · 2021
Closest in time.
Improving Cross-Modal Alignment in Vision Language Navigation via Syntactic Information
Li, J.; Tan, H.; and Bansal, M. 2021 · 2021
Closest in time.
UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning
Li, W.; Gao, C.; Niu, G.; Xiao, X.; Liu, H.; Liu, J.; Wu, H.; and Wang, H. 2021 · 2021
Closest in time.
Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation Learning
Xie, Z.; Lin, Y.; Zhang, Z.; Cao, Y.; Lin, S.; and Hu, H. 2021 · 2021
Closest in time.