Fetching the paper…
Reading the bibliography…
In this work, we propose a goal-driven collaborative task that combines language, perception, and action.
Unified pragmatic models for generating and following instructions
Daniel Fried, Jacob Andreas, and Dan Klein. 2018a · 1963
Earlier work this paper cites.
Convention: A Philosophical Study
David Lewis. 1969 · 1969
Earlier work this paper cites.
Languages and language
David Lewis. 1975 · 1975
Earlier work this paper cites.
The symbol grounding problem
Stevan Harnad. 1990 · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Perceptions of perceptual symbols
Lawrence W Barsalou. 1999 · 1999
Earlier work this paper cites.
An isu dialogue system exhibiting reinforcement learning of dialogue policies: generic slot-filling in the talk in-car system
Oliver Lemon, Kallirroi Georgila, James Henderson, and Matthew Stuttle. 2006 · 2006
Earlier work this paper cites.
Natural language processing with Python
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Learning to follow navigational directions
Adam Vogel and Daniel Jurafsky. 2010 · 2010
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
Stefanie A Tellex, Thomas Fleming Kollar, Steven R Dickerson, Matthew R Walter, Ashis Banerjee, Seth Teller, and Nicholas Roy. 2011 · 2011
Earlier work this paper cites.
A simple and generic belief tracking mechanism for the dialog state tracking challenge: On the believability of observed information
Zhuoran Wang and Oliver Lemon. 2013 · 2013
Earlier work this paper cites.
Bringing semantics into focus using visual abstraction
C. Lawrence Zitnick and Devi Parikh. 2013 · 2013
Earlier work this paper cites.
Learning the visual interpretation of sentences
C. Lawrence Zitnick, Devi Parikh, and Lucy Vanderwende. 2013 · 2013
Earlier work this paper cites.
A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
Mateusz Malinowski and Mario Fritz. 2014 · 2014
Earlier work this paper cites.
A framework for learning semantic maps from grounded natural language descriptions
Matthew R. Walter, Sachithra Hemachandra, Bianca Homberg, Stefanie Tellex, and Seth Teller. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Mind’s eye: A recurrent visual representation for image caption generation
Xinlei Chen and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015 · 2015
Cited alongside, same era.
Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. 2015 · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
Hongyuan Mei, Mohit Bansal, and Matthew R. Walter. 2015 · 2015
Cited alongside, same era.
Exploring Models and Data for Image Question Answering
Mengye Ren, Ryan Kiros, and Richard Zemel. 2015 · 2015
Cited alongside, same era.
Learning Multiagent Communication with Backpropagation
Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus. 2016 · 2016
Later among the works it cites.
MovieQA: Understanding Stories in Movies through Question-Answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016 · 2016
Later among the works it cites.
GuessWhat?! Visual object discovery through multi-modal dialogue
Harm de Vries, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, and Aaron Courville. 2016 · 2016
Later among the works it cites.
Learning Language Games through Interaction
Sida I. Wang, Percy Liang, and Christopher D. Manning. 2016 · 2016
Later among the works it cites.
Yin and Yang: Balancing and Answering Binary Visual Questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016 · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural Responding Machine for Short-Text Conversation
Lifeng Shang, Zhengdong Lu, and Hang Li. 2015 · 2015
Cited alongside, same era.
A Neural Network Approach to Context-Sensitive Generation of Conversational Responses
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, and William B. Dolan. 2015 · 2015
Cited alongside, same era.
Orioi Vinyals and Quoc V. Le. 2015 · 2015
Cited alongside, same era.
Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2017 · 2015
Cited alongside, same era.
Show, Attend and Tell : Neural Image Caption Generation with Visual Attention
Kelvin Xu, Aaron Courville, Richard S Zemel, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Visual Madlibs : Fill in the blank Description Generation and Question Answering
Licheng Yu, Eunbyung Park, Alexander C Berg, and Tamara L. Berg. 2015 · 2015
Cited alongside, same era.
Andrea F. Daniele, Mohit Bansal, and Matthew R. Walter. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Visual7W: Grounded Question Answering in Images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016 · 2016
Later among the works it cites.
Learning End-to-End Goal-Oriented Dialog
Antoine Bordes, Y-Lan Boureau, and Jason Weston. 2017 · 2017
Closest in time.
Learning symmetric collaborative dialogue agents with dynamic knowledge graph embeddings
He He, Anusha Balakrishnan, Mihail Eric, and Percy Liang. 2017 · 2017
Closest in time.
Natural language does not emerge ‘naturally’ in multi-agent dialog
Satwik Kottur, José Moura, Stefan Lee, and Dhruv Batra. 2017 · 2017
Closest in time.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Fei-Fei Li. 2017 · 2017
Closest in time.
Deal or no deal? end-to-end learning of negotiation dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017 · 2017
Closest in time.
Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher. 2017 · 2017
Closest in time.
Image-Grounded Conversations: Multimodal Context for Natural Question and Response Generation
Nasrin Mostafazadeh, Chris Brockett, Bill Dolan, Michel Galley, Jianfeng Gao, Georgios P. Spithourakis, and Lucy Vanderwende. 2017 · 2017
Closest in time.
End-to-end optimization of goal-driven and visually grounded dialogue systems
Florian Strub, Harm de Vries, Jeremie Mary, Bilal Piot, Aaron Courville, and Olivier Pietquin. 2017 · 2017
Closest in time.
Naturalizing a Programming Language via Interactive Learning
Sida I. Wang, Samuel Ginn, Percy Liang, and Christoper D. Manning. 2017 · 2017
Closest in time.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. 2018 · 2018
Closest in time.
Embodied question answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra. 2018 · 2018
Closest in time.