Fetching the paper…
Reading the bibliography…
We introduce the task of Visual Dialog, which requires an AI agent to hold a meaningful dialog with humans in natural, conversational language about visual content.
Empirical methods for evaluating dialog systems
T. Paek · 2001
Earlier work this paper cites.
An ISU dialogue system exhibiting reinforcement learning of dialogue policies: generic slot-filling in the TALK in-car system
O. Lemon, K. Georgila, J. Henderson, and M. Stuttle · 2006
Earlier work this paper cites.
VizWiz: Nearly Real-time Answers to Visual Questions
J. P. Bigham, C. Jayant, H. Ji, G. Little, A. Miller, R. C. Miller, R. Miller, A. Tatarowicz, B. White, S. White, and T. Yeh · 2010
Earlier work this paper cites.
Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
C. Danescu-Niculescu-Mizil and L. Lee · 2011
Earlier work this paper cites.
A Visual Turing Test for Computer Vision Systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
Q. V. L. Ilya Sutskever, Oriol Vinyals · 2014
Earlier work this paper cites.
What are you talking about? text-to-image coreference
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Linking people with "their" names using coreference resolution
V. Ramanathan, A. Joulin, P. Liang, and L. Fei-Fei · 2014
Earlier work this paper cites.
Joint Video and Text Parsing for Understanding Events and Answering Queries
K. Tu, M. Meng, M. W. Lee, T. E. Choe, and S. C. Zhu · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Large-scale Simple Question Answering with Memory Networks
A. Bordes, N. Usunier, S. Chopra, and J. Weston · 2015
Earlier work this paper cites.
Long-term Recurrent Convolutional Networks for Visual Recognition and Description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
From Captions to Visual Concepts and Back
H. Fang, S. Gupta, F. N. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2015
Earlier work this paper cites.
Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. Kingma and J. Ba · 2015
Earlier work this paper cites.
The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems
R. Lowe, N. Pow, I. Serban, and J. Pineau · 2015
Earlier work this paper cites.
Deeper LSTM and Normalized CNN Visual Question Answering model
J. Lu, X. Lin, D. Batra, and D. Parikh · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Cited alongside, same era.
Exploring Models and Data for Image Question Answering
M. Ren, R. Kiros, and R. Zemel · 2015
Cited alongside, same era.
A dataset for movie description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Sequence to Sequence - Video to Text
How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation
C.-W. Liu, R. Lowe, I. V. Serban, M. Noseworthy, L. Charlin, and J. Pineau · 2016
Closest in time.
SSD: Single Shot MultiBox Detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg · 2016
Closest in time.
Hierarchical Question-Image Co-Attention for Visual Question Answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Closest in time.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
H. Mei, M. Bansal, and M. R. Walter · 2016
Closest in time.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Closest in time.
Question Relevance in VQA: Identifying Non-Visual And False-Premise Questions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Venugopalan, M. Rohrbach, J. Donahue, R. J. Mooney, T. Darrell, and K. Saenko · 2015
Cited alongside, same era.
Translating Videos to Natural Language Using Deep Recurrent Neural Networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. J. Mooney, and K. Saenko · 2015
Cited alongside, same era.
O. Vinyals and Q. Le · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Visual Madlibs: Fill in the blank Image Generation and Question Answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Cited alongside, same era.
Analyzing the Behavior of Visual Question Answering Models
A. Agrawal, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Sort story: Sorting jumbled images and captions into stories
H. Agrawal, A. Chandrasekaran, D. Batra, D. Parikh, and M. Bansal · 2016
Cited alongside, same era.
A. Ray, G. Christie, M. Bansal, D. Batra, and D. Parikh · 2016
Closest in time.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Closest in time.
Generating Factoid Questions With Recurrent Neural Networks: The 30M Factoid Question-Answer Corpus
I. V. Serban, A. García-Durán, Ç. Gülçehre, S. Ahn, S. Chandar, A. C. Courville, and Y. Bengio · 2016
Closest in time.
Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models
I. V. Serban, A. Sordoni, Y. Bengio, A. Courville, and J. Pineau · 2016
Closest in time.
A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues
I. V. Serban, A. Sordoni, R. Lowe, L. Charlin, J. Pineau, A. Courville, and Y. Bengio · 2016
Closest in time.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Closest in time.
MovieQA: Understanding Stories in Movies through Question-Answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Closest in time.
Knowledge Guided Disambiguation for Large-Scale Scene Classification with Multi-Resolution CNNs
L. Wang, S. Guo, W. Huang, Y. Xiong, and Y. Qiao · 2016
Closest in time.
Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
J. Weston, A. Bordes, S. Chopra, and T. Mikolov · 2016
Closest in time.
Using Artificial Intelligence to Help Blind People ‘See’ Facebook
S. Wu, H. Pique, and J. Wieland · 2016
Closest in time.
Stacked Attention Networks for Image Question Answering
Z. Yang, X. He, J. Gao, L. Deng, and A. J. Smola · 2016
Closest in time.
Yin and Yang: Balancing and Answering Binary Visual Questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Closest in time.
Visual7W: Grounded Question Answering in Images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Closest in time.
Measuring machine intelligence through visual question answering
C. L. Zitnick, A. Agrawal, S. Antol, M. Mitchell, D. Batra, and D. Parikh · 2016
Closest in time.
GuessWhat?! Visual object discovery through multi-modal dialogue
H. de Vries, F. Strub, S. Chandar, O. Pietquin, H. Larochelle, and A. C. Courville · 2017
Closest in time.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Closest in time.
Image-Grounded Conversations: Multimodal Context for Natural Question and Response Generation
N. Mostafazadeh, C. Brockett, B. Dolan, M. Galley, J. Gao, G. P. Spithourakis, and L. Vanderwende · 2017
Closest in time.