Fetching the paper…
Reading the bibliography…
In this work, we study the problem of Embodied Referring Expression Grounding, where an agent needs to navigate in a previously unseen environment and localize a remote object described by a concise high-level natural language instruction.
Language Models are Few-shot Learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Self-monitoring Navigation Agent via Auxiliary Progress Estimation
Ma, C.-Y.; Lu, J.; Wu, Z.; AlRegib, G.; Kira, Z.; Socher, R.; and Xiong, C. 2019 · 1901
Earlier work this paper cites.
A Note on Two Problems in Connexion with Graphs
Dijkstra, E. W. 1959 · 1959
Earlier work this paper cites.
ConceptNet—a practical commonsense reasoning tool-kit
Liu, H.; and Singh, P. 2004 · 2004
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-regret Online Learning
Ross, S.; Gordon, G.; and Bagnell, D. 2011 · 2011
Earlier work this paper cites.
Wikidata: A Free Collaborative Knowledge Base
Vrandečić, D.; and Krötzsch, M. 2014 · 2014
Earlier work this paper cites.
Matterport3D: Learning from RGB-D Data in Indoor Environments
Chang, A.; Dai, A.; Funkhouser, T.; Halber, M.; Niessner, M.; Savva, M.; Song, S.; Zeng, A.; and Zhang, Y. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Dream to Control: Learning Behaviors by Latent Imagination
Hafner, D.; Lillicrap, T.; Ba, J.; and Norouzi, M. 2019 · 2019
Earlier work this paper cites.
Vilbert: Pretraining Task-agnostic Visiolinguistic Representations for Vision-and-language Tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Earlier work this paper cites.
Language Models as Knowledge Bases?
Petroni, F.; Rocktäschel, T.; Riedel, S.; Lewis, P.; Bakhtin, A.; Wu, Y.; and Miller, A. 2019 · 2019
Earlier work this paper cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Su, W.; Zhu, X.; Cao, Y.; Li, B.; Lu, L.; Wei, F.; and Dai, J. 2019 · 2019
Cited alongside, same era.
Reinforced Cross-modal Matching and Self-supervised Imitation Learning for Vision-language Navigation
Wang, X.; Huang, Q.; Celikyilmaz, A.; Gao, J.; Shen, D.; Wang, Y.-F.; Wang, W. Y.; and Zhang, L. 2019 · 2019
Cited alongside, same era.
Mastering Atari with Discrete World Models
Hafner, D.; Lillicrap, T. P.; Norouzi, M.; and Ba, J. 2020 · 2020
Cited alongside, same era.
Towards Learning a Generic Agent for Vision-and-language Navigation via Pre-training
Hao, W.; Li, C.; Li, X.; Carin, L.; and Gao, J. 2020 · 2020
Cited alongside, same era.
Beyond the Nav-graph: Vision-and-language Navigation in Continuous Environments
Krantz, J.; Wijmans, E.; Majumdar, A.; Batra, D.; and Lee, S. 2020 · 2020
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Later among the works it cites.
Soon: Scenario Oriented Object Navigation with Graph-based Exploration
Zhu, F.; Liang, X.; Zhu, Y.; Yu, Q.; Chang, X.; and Liang, X. 2021 · 2021
Later among the works it cites.
A Persistent Spatial Semantic Representation for High-level Natural Language Instruction Execution
Blukis, V.; Paxton, C.; Fox, D.; Garg, A.; and Artzi, Y. 2022 · 2022
Closest in time.
Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation
Chen, S.; Guhur, P.-L.; Tapaswi, M.; Schmid, C.; and Laptev, I. 2022 · 2022
Closest in time.
HOLM: Hallucinating Objects with Language Models for Referring Expression Recognition in Partially-Observed Scenes
Cirik, V.; Morency, L.-P.; and Berg-Kirkpatrick, T. 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Majumdar, A.; Shrivastava, A.; Lee, S.; Anderson, P.; Parikh, D.; and Batra, D. 2020 · 2020
Cited alongside, same era.
REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments
Qi, Y.; Wu, Q.; Anderson, P.; Wang, X.; Wang, W. Y.; Shen, C.; and van den Hengel, A. 2020 · 2020
Cited alongside, same era.
No rl, No Simulation: Learning to Navigate without Navigating
Hahn, M.; Chaplot, D. S.; Tulsiani, S.; Mukadam, M.; Rehg, J. M.; and Gupta, A. 2021 · 2021
Cited alongside, same era.
A Recurrent Vision-and-Language BERT for Navigation
Hong, Y.; Wu, Q.; Qi, Y.; Rodriguez-Opazo, C.; and Gould, S. 2021 · 2021
Cited alongside, same era.
Pathdreamer: A World Model for Indoor Navigation
Koh, J. Y.; Lee, H.; Yang, Y.; Baldridge, J.; and Anderson, P. 2021 · 2021
Cited alongside, same era.
Scene-intuitive Agent for Remote Embodied Visual Grounding
Lin, X.; Li, G.; and Yu, Y. 2021 · 2021
Cited alongside, same era.
FILM: Following Instructions in Language with Modular Methods
Min, S. Y.; Chaplot, D. S.; Ravikumar, P. K.; Bisk, Y.; and Salakhutdinov, R. 2021 · 2021
Cited alongside, same era.
Cross-modal Map Learning for Vision and Language Navigation
Georgakis, G.; Schmeckpeper, K.; Wanchoo, K.; Dan, S.; Miltsakaki, E.; Roth, D.; and Daniilidis, K. 2022 · 2022
Closest in time.
SASRA: Semantically-aware Spatio-temporal Reasoning Agent for Vision-and-Language Navigation in Continuous Environments
Irshad, M. Z.; Mithun, N. C.; Seymour, Z.; Chiu, H.-P.; Samarasekera, S.; and Kumar, R. 2022 · 2022
Closest in time.
Large Language Models are Zero-Shot Reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022 · 2022
Closest in time.
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Nichol, A. Q.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; Mcgrew, B.; Sutskever, I.; and Chen, M. 2022 · 2022
Closest in time.
HOP: History-and-Order Aware Pre-training for Vision-and-Language Navigation
Qiao, Y.; Qi, Y.; Hong, Y.; Yu, Z.; Wang, P.; and Wu, Q. 2022 · 2022
Closest in time.
Hierarchical Text-conditional Image Generation with Clip Latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Closest in time.
One Step at a Time: Long-Horizon Vision-and-Language Navigation with Milestones
Song, C. H.; Kil, J.; Pan, T.-Y.; Sadler, B. M.; Chao, W.-L.; and Su, Y. 2022 · 2022
Closest in time.
Find a Way Forward: a Language-Guided Semantic Map Navigator
Wang, Z.; Li, M.; Wu, M.; Moens, M.-F.; and Tuytelaars, T. 2022 · 2022
Closest in time.