Fetching the paper…
Reading the bibliography…
Aided by recent advances in Deep Learning, Image Caption Generation has seen tremendous progress over the last few years.
LeCun, Yann, Bernhard Boser, John Denker, Donnie Henderson, R. Howard, Wayne Hubbard, and Lawrence Jackel. ”Handwritten digit recognition with a back-propagation network.” Advances in neural information processing systems 2 (1989): 396-404
1989
Earlier work this paper cites.
Elman, Jeffrey L. ”Finding structure in time.” Cognitive science 14, no. 2 (1990): 179-211
1990
Earlier work this paper cites.
Hochreiter, Sepp, and Jürgen Schmidhuber. ”Long short-term memory.” Neural computation 9, no. 8 (1997): 1735-1780
1997
Earlier work this paper cites.
Kojima, Atsuhiro, Takeshi Tamura, and Kunio Fukunaga. ”Natural language description of human activities from video images based on concept hierarchy of actions.” International Journal of Computer Vision 50, no. 2 (2002): 171-184
2002
Earlier work this paper cites.
Deng, Jia, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ”Imagenet: A large-scale hierarchical image database.” In 2009 IEEE conference on computer vision and pattern recognition, pp. 248-255. Ieee, 2009
2009
Earlier work this paper cites.
Farhadi, Ali, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth. ”Every picture tells a story: Generating sentences from images.” In European conference on computer vision, pp. 15-29. Springer, Berlin, Heidelberg, 2010
2010
Earlier work this paper cites.
Li, Siming, Girish Kulkarni, Tamara Berg, Alexander Berg, and Yejin Choi. ”Composing simple image descriptions using web-scale n-grams.” In Proceedings of the Fifteenth Conference on Computational Natural Language Learning, pp. 220-228. 2011
2011
Earlier work this paper cites.
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E. Imagenet classification with deep convolutional neural networks. In NIPS, pp. 1097–1105, 2012
2012
Earlier work this paper cites.
Kulkarni, Girish, Visruth Premraj, Vicente Ordonez, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C. Berg, and Tamara L. Berg. ”Babytalk: Understanding and generating simple image descriptions.” IEEE Transactions on Pattern Analysis and Machine Intelligence 35, no. 12 (2013): 2891-2903
2013
Earlier work this paper cites.
Sutskever, Ilya, Oriol Vinyals, and Quoc V. Le. ”Sequence to sequence learning with neural networks.” In Advances in neural information processing systems, pp. 3104-3112. 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Mason, Rebecca, and Eugene Charniak. ”Nonparametric method for data-driven image captioning.” In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 592-598. 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Young, Peter, Alice Lai, Micah Hodosh, and Julia Hockenmaier. ”From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.” Transactions of the Association for Computational Linguistics 2 (2014): 67-78
2014
Cited alongside, same era.
Lin, Tsung-Yi, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. ”Microsoft coco: Common objects in context.” In European conference on computer vision, pp. 740-755. Springer, Cham, 2014
Huang, Gao, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q. Weinberger. ”Densely connected convolutional networks.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700-4708. 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. ”Attention is all you need.” In Advances in neural information processing systems, pp. 5998-6008. 2017
2017
Later among the works it cites.
Chen, Yunpeng, Jianan Li, Huaxin Xiao, Xiaojie Jin, Shuicheng Yan, and Jiashi Feng. ”Dual path networks.” In Advances in neural information processing systems, pp. 4467-4475. 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
Karpathy, Andrej, and Li Fei-Fei. ”Deep visual-semantic alignments for generating image descriptions.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3128-3137. 2015
2015
Cited alongside, same era.
Xu, Kelvin, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. ”Show, attend and tell: Neural image caption generation with visual attention.” In International conference on machine learning, pp. 2048-2057. 2015
2015
Cited alongside, same era.
Russakovsky, Olga, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang et al. ”Imagenet large scale visual recognition challenge.” International journal of computer vision 115, no. 3 (2015): 211-252
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Vinyals, Oriol, Alexander Toshev, Samy Bengio, and Dumitru Erhan. ”Show and tell: Lessons learned from the 2015 mscoco image captioning challenge.” IEEE transactions on pattern analysis and machine intelligence 39, no. 4 (2016): 652-663
2016
Cited alongside, same era.
Bernardi, Raffaella, Ruket Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, and Barbara Plank. ”Automatic description generation from images: A survey of models, datasets, and evaluation measures.” Journal of Artificial Intelligence Research 55 (2016): 409-442
2016
Cited alongside, same era.
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. ”Deep residual learning for image recognition.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778. 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Xie, Saining, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. ”Aggregated residual transformations for deep neural networks.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1492-1500. 2017
2017
Later among the works it cites.
Zhang, Xingcheng, Zhizhong Li, Chen Change Loy, and Dahua Lin. ”Polynet: A pursuit of structural diversity in very deep networks.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 718-726. 2017
2017
Later among the works it cites.
Aneja, Jyoti, Aditya Deshpande, and Alexander G. Schwing. ”Convolutional image captioning.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5561-5570. 2018
2018
Later among the works it cites.
Ma, Ningning, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. ”Shufflenet v2: Practical guidelines for efficient cnn architecture design.” In Proceedings of the European conference on computer vision (ECCV), pp. 116-131. 2018
2018
Later among the works it cites.
Sandler, Mark, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. ”Mobilenetv2: Inverted residuals and linear bottlenecks.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4510-4520. 2018
2018
Later among the works it cites.
Zoph, Barret, Vijay Vasudevan, Jonathon Shlens, and Quoc V. Le. ”Learning transferable architectures for scalable image recognition.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 8697-8710. 2018
2018
Later among the works it cites.
Hu, Jie, Li Shen, and Gang Sun. ”Squeeze-and-excitation networks.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7132-7141. 2018
2018
Later among the works it cites.
Hossain, MD Zakir, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga. ”A comprehensive survey of deep learning for image captioning.” ACM Computing Surveys (CSUR) 51, no. 6 (2019): 1-36
2019
Later among the works it cites.
Yu, Jun, Jing Li, Zhou Yu, and Qingming Huang. ”Multimodal transformer with multi-view visual representation for image captioning.” IEEE Transactions on Circuits and Systems for Video Technology (2019)
2019
Later among the works it cites.
Tan, Mingxing, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V. Le. ”Mnasnet: Platform-aware neural architecture search for mobile.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2820-2828. 2019
2019
Later among the works it cites.