Fetching the paper…
Reading the bibliography…
Visual language grounding is widely studied in modern neural image captioning systems, which typically adopts an encoder-decoder framework consisting of two principal components: a convolutional neural network (CNN) for image feature extraction and a recurrent neural network (RNN) for language caption generation.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Alon Lavie and Abhaya Agarwal. 2005 · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
VQA: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Mind’s eye: A recurrent visual representation for image caption generation
Xinlei Chen and C. Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Language models for image captioning: The quirks and what works
Jacob Devlin, Hao Cheng, Hao Fang, Saurabh Gupta, Li Deng, Xiaodong He, Geoffrey Zweig, and Margaret Mitchell. 2015 · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C Platt, et al. 2015 · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Guiding the long-short term memory model for image caption generation
Xu Jia, Efstratios Gavves, Basura Fernando, and Tinne Tuytelaars. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li. 2015 · 2015
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, and Alan L. Yuille. 2015 · 2015
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond J. Mooney, and Kate Saenko. 2015 · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Uncovering the temporal context for video question answering
Linchao Zhu, Zhongwen Xu, Yi Yang, and Alexander G Hauptmann. 2017 · 2015
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016 · 2016
Cited alongside, same era.
Visual storytelling
Ting-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, et al. 2016 · 2016
Cited alongside, same era.
Image captioning with semantic attention
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo. 2016 · 2016
Later among the works it cites.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner. 2017 · 2017
Closest in time.
ZOO: zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh. 2017 · 2017
Closest in time.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra. 2017 · 2017
Closest in time.
Guesswhat?! visual object discovery through multi-modal dialogue
Harm De Vries, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, and Aaron Courville. 2017 · 2017
Closest in time.
Semantic compositional networks for visual captioning
Zhe Gan, Chuang Gan, Xiaodong He, Yunchen Pu, Kenneth Tran, Jianfeng Gao, Lawrence Carin, and Li Deng. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016 · 2016
Cited alongside, same era.
Generating images from captions with attention
Elman Mansimov, Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov. 2016 · 2016
Cited alongside, same era.
Generating natural questions about an image
Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Margaret Mitchell, Xiaodong He, and Lucy Vanderwende. 2016 · 2016
Cited alongside, same era.
Distillation as a defense to adversarial perturbations against deep neural networks
Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami. 2016b · 2016
Cited alongside, same era.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016 · 2016
Cited alongside, same era.
Rich image captioning in the wild
Kenneth Tran, Xiaodong He, Lei Zhang, and Jian Sun. 2016 · 2016
Cited alongside, same era.
What value do explicit high level concepts have in vision to language problems?
Qi Wu, Chunhua Shen, Lingqiao Liu, Anthony Dick, and Anton van den Hengel. 2016 · 2016
Cited alongside, same era.
Closest in time.
Adversarial machine learning at scale
Alexey Kurakin, Ian Goodfellow, and Samy Bengio. 2017 · 2017
Closest in time.
Universal adversarial perturbations
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard. 2017 · 2017
Closest in time.
Image-grounded conversations: Multimodal context for natural question and response generation
Nasrin Mostafazadeh, Chris Brockett, Bill Dolan, Michel Galley, Jianfeng Gao, Georgios Spithourakis, and Lucy Vanderwende. 2017 · 2017
Closest in time.
Multi-task video captioning with video and entailment generation
Ramakanth Pasunuru and Mohit Bansal. 2017 · 2017
Closest in time.
Foil it! Find one mismatch between image and language caption
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurelie Herbelot, Moin Nabi, Enver Sangineto, Raffaella Bernardi, et al. 2017 · 2017
Closest in time.
EAD: elastic-net attacks to deep neural networks via adversarial examples
Pin-Yu Chen, Yash Sharma, Huan Zhang, Jinfeng Yi, and Cho-Jui Hsieh. 2018 · 2018
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.