Fetching the paper…
Reading the bibliography…
Advances in learning and representations have reinvigorated work that connects language to other modalities.
A note on two problems in connexion with graphs
Edsger W Dijkstra. 1959 · 1959
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams. 1992 · 1992
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut. 1999 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K.K. Paliwal. 1997 · 1997
Earlier work this paper cites.
Walk the talk: Connecting language, knowledge, action in route instructions
Matt Macmahon, Brian Stankiewicz, and Benjamin Kuipers. 2006 · 2006
Earlier work this paper cites.
Designing, developing, and deploying systems to support human-robot teams in disaster response
G.J.M. Kruijff, Ivana Kruijff-Korbayova, Shanker Keshavdas, Benoit Larochelle, Miroslav Janicek, Francis Colas, Ming Liu, François Pomerleau, Roland Siegwart, Mark Neerincx, Rosemarijn Looije, Nanja Smets, Tina Mioch, Jurriaan Diggelen, Fiora Pirri, Mario Gianni, Federico Ferri, Matteo Menna, Rainer Worst, and Vaclav Hlavac. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
VQA: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh. 2015 · 2015
Earlier work this paper cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Earlier work this paper cites.
Learning transferable policies for monocular reactive MAV control
Shreyansh Daftry, J. Andrew Bagnell, and Martial Hebert. 2016 · 2016
Earlier work this paper cites.
Deployment of ground and aerial robots in earthquake-struck amatrice in italy (brief report)
I. Kruijff-Korbayová, L. Freda, M. Gianni, V. Ntouskos, V. Hlaváč, V. Kubelka, E. Zimmermann, H. Surmann, K. Dulic, W. Rottner, and E. Gissi. 2016 · 2016
Cited alongside, same era.
An integrated system for interactive continuous learning of categorical knowledge
D. Skočaj, A. Vrečko, M. Mahnič, M. Janíček, G.-J. M. Kruijff, M. Hanheide, N. Hawes, J. L. Wyatt, T. Keller, K. Zhou, M. Zillich, and M. Kristan. 2016 · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alexander J. Smola. 2016 · 2016
Cited alongside, same era.
Video paragraph captioning using hierarchical recurrent neural networks
Haonan Yu, Jiang Wang, Zhiheng Huang, Yi Yang, and Wei Xu. 2016 · 2016
Cited alongside, same era.
Visual Dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M.F. Moura, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Mapping instructions to actions in 3D environments with visual goal prediction
Dipendra Misra, Andrew Bennett, Valts Blukis, Eyvind Niklasson, Max Shatkhin, and Yoav Artzi. 2018 · 2018
Later among the works it cites.
FollowNet: Robot navigation by following natural language directions with deep reinforcement learning
Pararth Shah, Marek Fiser, Aleksandra Faust, Chase Kew, and Dilek Hakkani-Tur. 2018 · 2018
Later among the works it cites.
Guiding exploratory behaviors for multi-modal grounding of linguistic descriptions
Jesse Thomason, Jivko Sinapov, Raymond Mooney, and Peter Stone. 2018 · 2018
Later among the works it cites.
Video captioning via hierarchical reinforcement learning
Xin Wang, Wenhu Chen, Jiawei Wu, Yuan-Fang Wang, and William Yang Wang. 2018 · 2018
Later among the works it cites.
Learning to parse natural language to grounded reward functions with weak supervision
Edward C. Williams, Nakul Gopalan, Mina Rhee, and Stefanie Tellex. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, M. Rohrbach, S. Venugopalan, S. Guadarrama, K. Saenko, and T. Darrell. 2017 · 2017
Cited alongside, same era.
Points, paths, and playscapes: Large-scale spatial language understanding tasks set in the real world
Jason Baldridge, Tania Bedrax-Weiss, Daphne Luong, Srini Narayanan, Bo Pang, Fernando Pereira, Radu Soricut, Michael Tseng, and Yuan Zhang. 2018 · 2018
Cited alongside, same era.
Learning interpretable spatial operations in a rich 3d blocks world
Yonatan Bisk, Kevin Shih, Yejin Choi, and Daniel Marcu. 2018 · 2018
Cited alongside, same era.
Mapping navigation instructions to continuous control actions with position visitation prediction
Valts Blukis, Dipendra Misra, Ross A. Knepper, and Yoav Artzi. 2018 · 2018
Cited alongside, same era.
Following formulaic map instructions in a street simulation environment
Volkan Cirik, Yuan Zhang, and Jason Baldridge. 2018 · 2018
Cited alongside, same era.
Talk the walk: Navigating new york city through grounded dialogue
Harm de Vries, Kurt Shuster, Dhruv Batra, Devi Parikh, Jason Weston, and Douwe Kiela. 2018 · 2018
Cited alongside, same era.
Speaker-Follower models for Vision-and-Language Navigation
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. 2018 · 2018
Cited alongside, same era.
Claudia Yan, Dipendra Misra, Andrew Bennett, Aaron Walsman, Yonatan Bisk, and Yoav Artzi. 2018 · 2018
Later among the works it cites.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Suhr, Dipendra Misra, and Yoav Artzi. 2019 · 2019
Closest in time.
Multi-modal discriminative model for vision-and-language navigation
Haoshuo Huang, Vihan Jain, Harsh Mehta, Jason Baldridge, and Eugene Ie. 2019 · 2019
Closest in time.
Self-monitoring navigation agent via auxiliary progress estimation
Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan Alregib, Zsolt Kira, Richard Socher, and Caiming Xiong. 2019 · 2019
Closest in time.
Shifting the baseline: Single modality performance on visual navigation & QA
Jesse Thomason, Daniel Gordon, and Yonatan Bisk. 2019 · 2019
Closest in time.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
Xin Wang, Qiuyuan Huang, Asli Çelikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. 2019 · 2019
Closest in time.