Fetching the paper…
Reading the bibliography…
Generating natural, diverse, and meaningful questions from images is an essential task for multimodal assistants as it confirms whether they have understood the object and scene in the images properly.
1906
Earlier work this paper cites.
1906
Earlier work this paper cites.
1908
Earlier work this paper cites.
1908
Earlier work this paper cites.
Pennington, J., Socher, R., & Manning, C. D. (2014). GloVe: Global vectors for word representation. EMNLP 2014 - 2014 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference. https://doi.org/10.3115/v1/d14-1162
2014
Earlier work this paper cites.
Simonyan, K., & Zisserman, A. (2015). Very deep convolutional networks for large-scale image recognition. 3rd International Conference on Learning Representations, ICLR 2015 - Conference Track Proceedings
2015
Earlier work this paper cites.
Gao, H., Mao, J., Zhou, J., Huang, Z., Wang, L., & Xu, W. (2015). Are you talking to a machine? Dataset and methods for multilingual image question answering. Advances in Neural Information Processing Systems
2015
Earlier work this paper cites.
Malinowski, M., Rohrbach, M., & Fritz, M. (2015). Ask your neurons: A neural-based approach to answering questions about images. Proceedings of the IEEE International Conference on Computer Vision. https://doi.org/10.1109/ICCV.2015.9
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Vinyals, O., Toshev, A., Bengio, S., & Erhan, D. (2015). Show and tell: A neural image caption generator. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR.2015.7298935
2015
Cited alongside, same era.
Simonyan, K., & Zisserman, A. (2015). Very deep convolutional networks for large-scale image recognition. 3rd International Conference on Learning Representations, ICLR 2015 - Conference Track Proceedings
2015
Cited alongside, same era.
Mostafazadeh, N., Misra, I., Devlin, J., Mitchell, M., He, X., & Vanderwende, L. (2016). Generating natural questions about an image. 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016 - Long Papers, 3, 1802–1813. https://doi.org/10.18653/v1/p16-1170
2016
Cited alongside, same era.
Wu, Q., Wang, P., Shen, C., Dick, A., & Van Den Hengel, A. (2016). Ask me anything: Free-form visual question answering based on knowledge from external sources. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR.2016.500
Huang, G., Liu, Z., Van Der Maaten, L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
Zhang, S., Qu, L., You, S., Yang, Z., & Zhang, J. (2016). Automatic Generation of Grounded Visual Questions. IJCAI International Joint Conference on Artificial Intelligence. https://doi.org/10.24963/ijcai.2017/592
2017
Later among the works it cites.
Huang, G., Liu, Z., Van Der Maaten, L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017. https://doi.org/10.1109/CVPR.2017.243
2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR.2016.90
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Jain, U., Zhang, Z., & Schwing, A. (2017). Creativity: Generating Diverse Questions using Variational Autoencoders. Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017. https://doi.org/10.1109/CVPR.2017.575
2017
Cited alongside, same era.
Wang, P., Wu, Q., Shen, C., Dick, A., & Van Den Hengel, A. (2017). KB-VQA Explicit knowledge-based reasoning for visual question answering. IJCAI International Joint Conference on Artificial Intelligence, 1290–1296. https://doi.org/10.24963/ijcai.2017/179
2017
Cited alongside, same era.
Wang, P., Wu, Q., Shen, C., Hengel, A. van den, & Dick, A. (2016). FVQA: Fact-based Visual Question Answering. IEEE Transactions on Pattern Analysis and Machine Intelligence. https://doi.org/10.1109/TPAMI.2017.2754246
2017
Cited alongside, same era.
Later among the works it cites.
Agrawal, A., Lu, J., Antol, S., Mitchell, M., Zitnick, C. L., Parikh, D., & Batra, D. (2017). VQA: Visual Question Answering: www.visualqa.org. International Journal of Computer Vision. https://doi.org/10.1007/s11263-016-0966-6
2017
Later among the works it cites.
2018
Later among the works it cites.
Narasimhan, M., Lazebnik, S., & Schwing, A. G. (2018). Out of the box: Reasoning with graph convolution nets for factual visual question answering. Advances in Neural Information Processing Systems
2018
Later among the works it cites.
Shahbaz, M., Suresh, L., Rexford, J., Feamster, N., Rottenstreich, O., & Hira, M. (2019). Elmo. https://doi.org/10.1145/3341302.3342066
2019
Later among the works it cites.
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, J. D. (2020). VL-BERT: Pre-training of Generic Visual-Linguistic Representations. 1–16
2020
Closest in time.