Fetching the paper…
Reading the bibliography…
Dialog is an effective way to exchange information, but subtle details and nuances are extremely important.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C.L. Zitnick · 2014
Earlier work this paper cites.
A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Memory networks
J. Weston, S. Chopra, and A. Bordes · 2014
Earlier work this paper cites.
VQA: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Are you talking to a machine? Dataset and Methods for Multilingual Image Question Answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Earlier work this paper cites.
Visual turing test for computer vision systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Loffe and C. Szegedy · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Earlier work this paper cites.
Deep Captioning with Multimodal Recurrent Neural Networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Earlier work this paper cites.
Deep compositional question answering with neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Deep reinforcement learning for mention-ranking coreference models
K. Clark and C. Manning · 2016
Earlier work this paper cites.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
A. Das, H. Agrawal, C. L. Zitnick, D. Parikh, and D. Batra · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Earlier work this paper cites.
Hadamard product for low-rank bilinear pooling
J. Kim, K. On, W. Lim, J. Kim, J. Ha, and B. Zhang · 2016
Cited alongside, same era.
How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
C.W. Liu, R. Lowe, I.V. Serban, M. Noseworthy, L. Charlin, and J. Pineau · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Generating natural questions about an image
N. Mostafazadeh, I. Misra, J. Devlin, M. Mitchell, X. He, and L. Vanderwende · 2016
Cited alongside, same era.
Where to look: Focus regions for visual question answering
K. J. Shih, S. Singh, and D. Hoiem · 2016
Cited alongside, same era.
Image captioning and visual question answering based on attributes and their related external knowledge
High-Order Attention Models for Visual Question Answering
I. Schwartz, A. G. Schwing, and T. Hazan · 2017
Later among the works it cites.
Visual reference resolution using attention memory for visual dialog
P. H. Seo, A. Lehrmann, B. Han, and L. Sigal · 2017
Later among the works it cites.
A hierarchical latent variable encoder-decoder model for generating dialogues
I. V. Serban, A. Sordoni, R. Lowe, L. Charlin, J. Pineau, A. Courville, and Y. Bengio · 2017
Later among the works it cites.
P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio · 2017
Later among the works it cites.
Diverse and Accurate Image Description Using a Variational Auto-Encoder with an Additive Gaussian Encoding Space
L. Wang, A. G. Schwing, and S. Lazebnik · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Q. Wu, C. Shen, A. van den Hengel, P. Wang, and A. Dick · 2016
Cited alongside, same era.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Cited alongside, same era.
Visual7W: Grounded Question Answering in Images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Cited alongside, same era.
Mutan: Multimodal tucker fusion for visual question answering
H. Ben-Younes, R. Cadene, M. Cord, and N. Thome · 2017
Cited alongside, same era.
Visual Dialog
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra · 2017
Cited alongside, same era.
Q. Wu, P. Wang, C. Shen, I. Reid, and A. van den Hengel · 2017
Later among the works it cites.
Structured attentions for visual question answering
C. Zhu, Y. Zhao, S. Huang, K. Tu, and Y. Ma · 2017
Later among the works it cites.
Audio visual scene-aware dialog (avsd) challenge at dstc7
H. Alamri, V Cartillier, A. Das, J. Wang, J. Essa, D. Batra, D. Parikh, A. Cherian, T. K. Marks, and C. Hori · 2018
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Later among the works it cites.
Convolutional Image Captioning
J. Aneja, A. Deshpande, and A. G. Schwing · 2018
Later among the works it cites.
Diverse and Coherent Paragraph Generation from Images
M. Chatterjee and A. G. Schwing · 2018
Later among the works it cites.
Diverse and Controllable Image Captioning with Part-of-Speech Guidance
A. Deshpande, J. Aneja, L. Wang, A. G. Schwing, and D. A. Forsyth · 2018
Later among the works it cites.
End-to-end audio visual scene-aware dialog using multimodal attention-based video features
C. Hori, H. Alamri, J. Wang, G. Winchern, T. Hori, A. Cherian, T.K. Marks, V. Cartillier, R.G. Lopes, A. Das, I. Essa, D. Batra, and D. Parikh · 2018
Later among the works it cites.
Two can play this game: Visual dialog with discriminative question generation and answering
U. Jain, S. Lazebnik, and A. G. Schwing · 2018
Later among the works it cites.
Visual coreference resolution in visual dialog using neural module networks
S. Kottur, J. Moura, D. Parikh, D. Batra, and M. Rohrbach · 2018
Later among the works it cites.
Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering
M. Narasimhan, S. Lazebnik, and A. G. Schwing · 2018
Later among the works it cites.
Straight to the Facts: Learning Knowledge Base Retrieval for Factual Visual Question Answering
M. Narasimhan and A. G. Schwing · 2018
Later among the works it cites.
Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models
A. K. Vijayakumar, M. Cogswell, R. R. Selvaraju, Q. Sun, S. Lee, D. Crandall, and D. Batra · 2018
Later among the works it cites.
A Simple Baseline for Audio-Visual Scene-Aware Dialog
I. Schwartz, A. G. Schwing, and T. Hazan · 2019
Closest in time.