Fetching the paper…
Reading the bibliography…
Visual Question Generation (VQG) is the task of generating natural questions based on an image.
Triplet-Center Loss for Multi-view 3D Object Retrieval
Xinwei He, Yang Zhou, Zhichao Zhou, Song Bai, and Xiang Bai. 2018 · 1954
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out . Association for Computational Linguistics, Barcelona, Spain, 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling. 2013 · 2013
Earlier work this paper cites.
A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
Mateusz Malinowski and Mario Fritz. 2014 · 2014
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Devi Parikh, and Dhruv Batra. 2015 · 2015
Earlier work this paper cites.
VQA: Visual Question Answering. In International Conference on Computer Vision (ICCV)
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Neural Self Talk: Image Understanding via Continuous Questioning and Answering
Yezhou Yang, Yi Li, Cornelia Fermüller, and Yiannis Aloimonos. 2015 · 2015
Earlier work this paper cites.
Visual7W: Grounded Question Answering in Images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Generating Natural Questions About an Image
Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Margaret Mitchell, Xiaodong He, and Lucy Vanderwende. 2016 · 2016
Earlier work this paper cites.
A Discriminative Feature Learning Approach for Deep Face Recognition. In ECCV
Yandong Wen, Kaipeng Zhang, Zhifeng Li, and Yu Qiao. 2016 · 2016
Earlier work this paper cites.
Automatic Generation of Grounded Visual Questions
Shijie Zhang, Lizhen Qu, Shaodi You, Zhenglu Yang, and Jiawan Zhang. 2016 · 2016
Earlier work this paper cites.
Creativity: Generating Diverse Questions Using Variational Autoencoders
Unnat Jain, Ziyu Zhang, and Alexander G. Schwing. 2017 · 2017
Cited alongside, same era.
Visual Question Generation as Dual Task of Visual Question Answering
Yikang Li, Nan Duan, Bolei Zhou, X. R. Chu, Wanli Ouyang, and Xiaogang Wang. 2017 · 2017
Cited alongside, same era.
Image-Grounded Conversations: Multimodal Context for Natural Question and Response Generation. In IJCNLP
Nasrin Mostafazadeh, Chris Brockett, William B. Dolan, Michel Galley, Jianfeng Gao, Georgios P. Spithourakis, and Lucy Vanderwende. 2017 · 2017
Cited alongside, same era.
Multimodal Analysis of User-Generated Multimedia Content (1st ed.)
Rajiv Shah and Roger Zimmermann. 2017 · 2017
Cited alongside, same era.
A Joint Model for Question Answering and Question Generation
Tong Wang, Xingdi Yuan, and Adam Trischler. 2017 · 2017
Cited alongside, same era.
Dual learning for visual question generation. In 2018 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 1–6
Xing Xu, Jingkuan Song, Huimin Lu, Li He, Yang Yang, and Fumin Shen. 2018 · 2018
Later among the works it cites.
Visual Curiosity: Learning to Ask Questions to Learn Visual Recognition. In CoRL
Jianwei Yang, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh. 2018 · 2018
Later among the works it cites.
Deep Learning for Video Captioning: A Review. In IJCAI
Shaoxiang Chen, Ting Yao, and Yu-Gang Jiang. 2019 · 2019
Later among the works it cites.
A Comprehensive Survey of Deep Learning for Image Captioning
MD. Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga. 2019 · 2019
Later among the works it cites.
Bayes-Factor-VAE: Hierarchical Bayesian Deep Auto-Encoder Models for Factor Disentanglement
Minyoung Kim, Yuting Wang, Pritish Sahu, and Vladimir Pavlovic. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hyperprior Induced Unsupervised Disentanglement of Latent Representations. In AAAI
Abdul Fatir Ansari and Harold Soh. 2018 · 2018
Cited alongside, same era.
Understanding Center Loss Based Network for Image Retrieval with Few Training Data. In ECCV Workshops
Pallabi Ghosh and Larry S. Davis. 2018 · 2018
Cited alongside, same era.
Attribute-Centered Loss for Soft-Biometrics Guided Face Sketch-Photo Recognition
Hadi Kazemi, Sobhan Soleymani, Ali Dabouei, Seyed Mehdi Iranmanesh, and Nasser M. Nasrabadi. 2018 · 2018
Cited alongside, same era.
Learning Latent Subspaces in Variational Autoencoders
Jack Klys, Jake Snell, and Richard S. Zemel. 2018 · 2018
Cited alongside, same era.
Information Maximizing Visual Question Generation
Ranjay Krishna, Michael Bernstein, and Li Fei-Fei. 2019 · 2018
Cited alongside, same era.
iVQA: Inverse visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 8611–8619
Feng Liu, Tao Xiang, Timothy M Hospedales, Wankou Yang, and Changyin Sun. 2018 · 2018
Cited alongside, same era.
A Comprehensive Study on Center Loss for Deep Face Recognition
Yandong Wen, Kaipeng Zhang, Zhifeng Li, and Yu Qiao. 2018 · 2018
Cited alongside, same era.
A Logic-Driven Framework for Consistency of Neural Models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing
Tao Li, Vivek Gupta, Maitrey Mehta, and Vivek Srikumar. 2019 · 2019
Later among the works it cites.
Text2FaceGAN: Face Generation from Fine Grained Textual Descriptions
Osaid Rehman Nasir, S. K. Jha, M. S. Grover, Y. Yu, Ajit Kumar, and R. Shah. 2019 · 2019
Later among the works it cites.
User Input Based Style Transfer While Retaining Facial Attributes
Sharan Pai, Nikhil Sachdeva, R. Shah, and R. Zimmermann. 2019 · 2019
Later among the works it cites.
Product of Orthogonal Spheres Parameterization for Disentangled Representation Learning. In BMVC
Ankita Shukla, Sarthak Bhagat, Shagun Uppal, Saket Anand, and Pavan K. Turaga. 2019 · 2019
Later among the works it cites.
Disentangling Multiple Features in Video Sequences Using Gaussian Processes in Variational Autoencoders. In ECCV
Sarthak Bhagat, Shagun Uppal, Vivian T. Yin, and N. Lim. 2020 · 2020
Closest in time.
Emerging Trends of Multimodal Research in Vision and Language
Shagun Uppal, Sarthak Bhagat, Devamanyu Hazarika, Navonil Majumdar, Soujanya Poria, R. Zimmermann, and Amir Zadeh. 2020a · 2020
Closest in time.
Two-Step Classification using Recasted Data for Low Resource Settings. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing . Association for Computational Linguistics, Suzhou, China, 706–719
Shagun Uppal, Vivek Gupta, Avinash Swaminathan, Haimin Zhang, Debanjan Mahata, Rakesh Gosangi, Rajiv Ratn Shah, and Amanda Stent. 2020b · 2020
Closest in time.