Fetching the paper…
Reading the bibliography…
Building computer systems that can converse about their visual environment is one of the oldest concerns of research in Artificial Intelligence and Computational Linguistics (see, for example, Winograd's 1972 SHRDLU system).
Speakers’ and Listeners’ Processes in a Word-Communication Task
Seymour Rosenberg and Bertram D. Cohen. 1964 · 1964
Earlier work this paper cites.
Referring as a collaborative process
Herbert H. Clark and Deanna Wilkes-Gibbs. 1986 · 1986
Earlier work this paper cites.
The symbol grounding problem
Stevan Harnard. 1990 · 1990
Earlier work this paper cites.
The hcrc map task corpus
Anne H Anderson, Miles Bader, Ellen Gurman Bard, Elizabeth Boyle, Gwyneth Doherty, Simon Garrod, Stephen Isard, Jacqueline Kowtko, Jan McAllister, Jim Miller, et al. 1991 · 1991
Earlier work this paper cites.
Grounding in communication
Herbert H Clark, Susan E Brennan, et al. 1991 · 1991
Earlier work this paper cites.
On Distinguishing Epistemic from Pragmatic Action
David Kirsh and Paul Maglio. 1994 · 1994
Earlier work this paper cites.
Using language. 1996
Herbert H Clark. 1996b · 1996
Earlier work this paper cites.
Reinforcement Learning
Richard S. Sutton and Andrew G. Barto. 1998 · 1998
Earlier work this paper cites.
Generating instructions in virtual environments (give): A challenge and an evaluation testbed for nlg
Donna Byron, Alexander Koller, Jon Oberlander, Laura Stoia, and Kristina Striegnitz. 2007 · 2007
Earlier work this paper cites.
Referring under restricted interactivity conditions
Raquel Fernández and David Schlangen. 2007 · 2007
Earlier work this paper cites.
Partially observable Markov decision processes for spoken dialog systems
Jason Williams and Steve Young. 2007 · 2007
Earlier work this paper cites.
Modelling sub-utterance phenomena in spoken dialogue systems
Okko Buß and David Schlangen. 2010 · 2010
Earlier work this paper cites.
Computational approaches to the production of referring expressions: Dialog changes (almost) everything
Amanda J Stent. 2011 · 2011
Earlier work this paper cites.
The rex corpora: A collection of multimodal corpora of referring expressions in collaborative problem solving dialogues
Takenobu Tokunaga, Ryu Iida, Asuka Terai, and Naoko Kuriyama. 2012 · 2012
Cited alongside, same era.
ReferItGame: Referring to Objects in Photographs of Natural Scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara L Berg. 2014 · 2014
Cited alongside, same era.
Learning image embeddings using convolutional neural networks for improved multi-modal semantics
Douwe Kiela and Léon Bottou. 2014 · 2014
Cited alongside, same era.
Learning grounded meaning representations with autoencoders
Carina Silberer and Mirella Lapata. 2014 · 2014
Cited alongside, same era.
Mind’s eye: A recurrent visual representation for image caption generation
Xinlei Chen and C Lawrence Zitnick. 2015 · 2015
Cited alongside, same era.
Language models for image captioning: The quirks and what works
Resolving references to objects in photographs using the words-as-classifiers model
David Schlangen, Sina Zarriess, and Casey Kennington. 2016 · 2016
Later among the works it cites.
Modeling Context in Referring Expressions , pages 69–85. Springer International Publishing, Cham
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg. 2016 · 2016
Later among the works it cites.
Pentoref: A corpus of spoken references in task-oriented dialogues
Sina Zarrieß, Julian Hough, Casey Kennington, Ramesh Manuvinakurike, David DeVault, Raquel Fernandez, and David Schlangen. 2016 · 2016
Later among the works it cites.
Guesswhat?! visual object discovery through multi-modal dialogue
Harm De Vries, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, and Aaron Courville. 2017 · 2017
Later among the works it cites.
MINOS: Multimodal indoor simulator for navigation in complex environments
Manolis Savva, Angel X. Chang, Alexey Dosovitskiy, Thomas Funkhouser, and Vladlen Koltun. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Devlin, Hao Cheng, Hao Fang, Saurabh Gupta, Li Deng, Xiaodong He, Geoffrey Zweig, and Margaret Mitchell. 2015 · 2015
Cited alongside, same era.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh Srivastava, Li Deng, Piotr Dollar, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John Platt, Lawrence Zitnick, and Geoffrey Zweig. 2015 · 2015
Cited alongside, same era.
Combining language and vision with a multimodal skip-gram model
Angeliki Lazaridou, Nghia The Pham, and Marco Baroni. 2015 · 2015
Cited alongside, same era.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L. Yuille, and Kevin Murphy. 2015 · 2015
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
Raffaella Bernardi, Ruket Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, and Barbara Plank. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. 2017 · 2017
Later among the works it cites.
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. 2018 · 2018
Later among the works it cites.
TextWorld: A Learning Environment for Text-based Games
Marc-Alexandre Côté, Ákos Kádár, Xingdi Yuan, Ben Kybartas, Tavian Barnes, Emery Fine, James Moore, Matthew Hausknecht, Layla El Asri, Mahmoud Adada, Wendy Tay, and Adam Trischler. 2018 · 2018
Later among the works it cites.
slurk – A Lightweight Interaction Server For Dialogue Experiments and Data Collection
David Schlangen, Tim Diekmann, Nikolai Ilinykh, and Sina Zarrieß. 2018 · 2018
Later among the works it cites.
The PhotoBook Dataset: Building Common Ground through Visually-Grounded Dialogue
Janosch Haber, Tim Baumgärtner, Ece Takmaz, Lieke Gelderloos, Elia Bruni, and Raquel Fernández. 2019 · 2019
Closest in time.
Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan AlRegib, Zsolt Kira, Richard Socher, and Caiming Xiong. 2019 · 2019
Closest in time.
Habitat: A Platform for Embodied AI Research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. 2019 · 2019
Closest in time.
Learning to Speak and Act in a Fantasy Text Adventure Game
Jack Urbanek, Angela Fan, Siddharth Karamcheti, Saachi Jain, Samuel Humeau, Emily Dinan, Tim Rocktäschel, Douwe Kiela, Arthur Szlam, and Jason Weston. 2019 · 2019
Closest in time.