Fetching the paper…
Reading the bibliography…
The image, question (combined with the history for de-referencing), and the corresponding answer are three vital components of visual dialog.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
S. Auer, C. Bizer, G. Kobilarov, J. Lehmann, R. Cyganiak, and Z. Ives · 2007
Earlier work this paper cites.
A theoretical analysis of ndcg ranking measures
Y. Wang, L. Wang, Y. Li, D. He, W. Chen, and T.-Y. Liu · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
E. Jang, S. Gu, and B. Poole · 2016
Cited alongside, same era.
Hadamard product for low-rank bilinear pooling
J.-H. Kim, K.-W. On, W. Lim, J. Kim, J.-W. Ha, and B.-T. Zhang · 2016
Cited alongside, same era.
Dual attention networks for multimodal reasoning and matching
H. Nam, J.-W. Ha, and J. Kim · 2016
Cited alongside, same era.
Improved deep metric learning with multi-class n-pair loss objective
K. Sohn · 2016
Cited alongside, same era.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
J. Lu, C. Xiong, D. Parikh, and R. Socher · 2017
Later among the works it cites.
Learning visual reasoning without strong priors
E. Perez, H. De Vries, F. Strub, V. Dumoulin, and A. Courville · 2017
Later among the works it cites.
Fvqa: Fact-based visual question answering
P. Wang, Q. Wu, C. Shen, A. Dick, and A. van den Hengel · 2017
Later among the works it cites.
Are you talking to me? reasoned visual dialog generation through adversarial learning
Q. Wu, P. Wang, C. Shen, I. Reid, and A. van den Hengel · 2017
Later among the works it cites.
Seqgan: Sequence generative adversarial nets with policy gradient
L. Yu, W. Zhang, J. Wang, and Y. Yu · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ask me anything: Free-form visual question answering based on knowledge from external sources
Q. Wu, P. Wang, C. Shen, A. Dick, and A. van den Hengel · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Cited alongside, same era.
Visual dialog
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra · 2017
Cited alongside, same era.
Guesswhat?! visual object discovery through multi-modal dialogue
H. De Vries, F. Strub, S. Chandar, O. Pietquin, H. Larochelle, and A. C. Courville · 2017
Cited alongside, same era.
Modulating early visual processing by language
H. De Vries, F. Strub, J. Mary, H. Larochelle, O. Pietquin, and A. C. Courville · 2017
Cited alongside, same era.
Best of both worlds: Transferring knowledge from discriminative learning to a generative visual dialog model
J. Lu, A. Kannan, J. Yang, D. Parikh, and D. Batra · 2017
Cited alongside, same era.
Multi-modal factorized bilinear pooling with co-attention learning for visual question answering
Z. Yu, J. Yu, J. Fan, and D. Tao · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Later among the works it cites.
Two can play this game: Visual dialog with discriminative question generation and answering
U. Jain, S. Lazebnik, and A. G. Schwing · 2018
Later among the works it cites.
J.-H. Kim, J. Jun, and B.-T. Zhang · 2018
Later among the works it cites.
Straight to the facts: Learning knowledge base retrieval for factual visual question answering
M. Narasimhan and A. G. Schwing · 2018
Later among the works it cites.
Rethinking diversified and discriminative proposal generation for visual grounding
Z. Yu, J. Yu, C. Xiang, Z. Zhao, Q. Tian, and D. Tao · 2018
Later among the works it cites.