Fetching the paper…
Reading the bibliography…
In this paper, we consider the problem of solving semantic tasks such as `Visual Question Answering' (VQA), where one aims to answers related to an image and `Visual Question Generation' (VQG), where one aims to generate a natural question pertaining to an image.
A. Tversky and D. Kahneman, “Availability: A heuristic for judging frequency and probability,”
1973
Earlier work this paper cites.
R. Shepard, “Toward a universal law of generalization for psychological science,”
1987
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in
2002
Earlier work this paper cites.
K. Barnard, P. Duygulu, and D. Forsyth, “N. de freitas, d,”
2003
Earlier work this paper cites.
A. R. Chappell, A. J. Cowell, D. A. Thurman, and J. R. Thomson, “Supporting mutual understanding in a visual dialogue between analyst and computer,” in
2004
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in
2005
Earlier work this paper cites.
J. Demšar, “Statistical comparisons of classifiers over multiple data sets,”
2006
Earlier work this paper cites.
A. Frome, Y. Singer, F. Sha, and J. Malik, “Learning globally-consistent local distance functions for shape-based image retrieval and classification,” in
2007
Earlier work this paper cites.
A. Frome, Y. Singer, F. Sha, and J. Malik, “Learning globally-consistent local distance functions for shape-based image retrieval and classification,” in
2007
Earlier work this paper cites.
J. V. Davis, B. Kulis, P. Jain, S. Sra, and I. S. Dhillon, “Information-theoretic metric learning,” in
2007
Earlier work this paper cites.
F. Jäkel, B. Schölkopf, and F. Wichmann, “Generalization and similarity in exemplar models of categorization: Insights from machine learning,”
2008
Earlier work this paper cites.
T. Judd, K. Ehinger, F. Durand, and A. Torralba, “Learning to predict where humans look,” in
2009
Earlier work this paper cites.
K. Q. Weinberger and L. K. Saul, “Distance metric learning for large margin nearest neighbor classification,”
2009
Earlier work this paper cites.
J. H. McDonald,
2009
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth, “Every picture tells a story: Generating sentences from images,” in
2010
Earlier work this paper cites.
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg, “Baby talk: Understanding and generating image descriptions,” in
2011
Earlier work this paper cites.
M. Malinowski and M. Fritz, “A multi-world approach to question answering about real-world scenes based on uncertain input,” in
2014
Earlier work this paper cites.
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng, “Grounded compositional semantics for finding and describing images with sentences,”
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in
2014
Earlier work this paper cites.
A. Karpathy, A. Joulin, and F. F. F. Li, “Deep fragment embeddings for bidirectional image sentence mapping,” in
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual Question Answering,” in
2015
Cited alongside, same era.
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu, “Are you talking to a machine? dataset and methods for multilingual image question,” in
2015
Cited alongside, same era.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in
2015
Cited alongside, same era.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in
2015
Cited alongside, same era.
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt
2015
Cited alongside, same era.
J. Johnson, A. Karpathy, and L. Fei-Fei, “Densecap: Fully convolutional localization networks for dense captioning,” in
2016
Later among the works it cites.
X. Yan, J. Yang, K. Sohn, and H. Lee, “Attribute2image: Conditional image generation from visual attributes,” in
2016
Later among the works it cites.
A. Das, H. Agrawal, C. L. Zitnick, D. Parikh, and D. Batra, “Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?” in
2016
Later among the works it cites.
L. Ma, Z. Lu, and H. Li, “Learning to answer questions from image using convolutional neural network,” in
2016
Later among the works it cites.
K. J. Shih, S. Singh, and D. Hoiem, “Where to look: Focus regions for visual question answering,” in
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Chen and C. Lawrence Zitnick, “Mind’s eye: A recurrent visual representation for image caption generation,” in
2015
Cited alongside, same era.
D. Geman, S. Geman, N. Hallonquist, and L. Younes, “Visual turing test for computer vision systems,”
2015
Cited alongside, same era.
M. Malinowski, M. Rohrbach, and M. Fritz., “Ask your neurons: A neural-based approach to answering questions about images bibtex,” in
2015
Cited alongside, same era.
M. Ren, R. Kiros, and R. Zemel, “Exploring models and data for image question answering,” in
2015
Cited alongside, same era.
2015
Cited alongside, same era.
J. Lu, X. Lin, D. Batra, and D. Parikh, “Deeper lstm and normalized cnn visual question answering model,”
2015
Cited alongside, same era.
E. Hoffer and N. Ailon, “Deep metric learning using triplet network,” in
2015
Cited alongside, same era.
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei, “Visual7w: Grounded question answering in images,” in
2016
Later among the works it cites.
H. Xu and K. Saenko, “Ask, attend and answer: Exploring question-guided spatial attention for visual question answering,” in
2016
Later among the works it cites.
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein, “Learning to compose neural networks for question answering,” in
2016
Later among the works it cites.
2016
Later among the works it cites.
C. Huang, Y. Li, C. Change Loy, and X. Tang, “Learning deep representation for imbalanced classification,” in
2016
Later among the works it cites.
D. Fišer, T. Erjavec, and N. Ljubešić, “Janes v0. 4: Korpus slovenskih spletnih uporabniških vsebin,”
2016
Later among the works it cites.
R. Zhang, P. Isola, and A. A. Efros, “Colorful image colorization,” in
2016
Later among the works it cites.
J.-H. Kim, K. W. On, W. Lim, J. Kim, J.-W. Ha, and B.-T. Zhang, “Hadamard Product for Low-rank Bilinear Pooling,” in
2017
Later among the works it cites.
U. Jain, Z. Zhang, and A. G. Schwing, “Creativity: Generating diverse questions using variational autoencoders.” in
2017
Later among the works it cites.
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra, “Visual Dialog,” in
2017
Later among the works it cites.
H. De Vries, F. Strub, S. Chandar, O. Pietquin, H. Larochelle, and A. Courville, “Guesswhat?! visual object discovery through multi-modal dialogue,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
B. Patro and V. P. Namboodiri, “Differential attention for visual question answering,” in
2018
Later among the works it cites.
B. N. Patro, S. Kumar, V. K. Kurmi, and V. Namboodiri, “Multimodal differential network for visual question generation,” in
2018
Later among the works it cites.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in
2057
Closest in time.