Fetching the paper…
Reading the bibliography…
Given an image, generating its natural language description (i.e., caption) is a well studied problem.
Learning structured embeddings of knowledge bases
Antoine Bordes, Jason Weston, Ronan Collobert, and Yoshua Bengio · 2011
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher · 2016
Cited alongside, same era.
Rdf2vec: Rdf graph embeddings for data mining
Petar Ristoski and Heiko Paulheim · 2016
Cited alongside, same era.
Diverse beam search: Decoding diverse solutions from neural sequence models
Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra · 2016
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and vqa
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2017
Cited alongside, same era.
Show, adapt and tell: Adversarial training of cross-domain image captioner
Generating diverse and accurate visual captions by comparative adversarial learning
Dianqi Li, Qiuyuan Huang, Xiaodong He, Lei Zhang, and Ming-Ting Sun · 2018
Later among the works it cites.
Knowledge guided attention and inference for describing images containing unseen objects
Aditya Mogadala, Umanga Bista, Lexing Xie, and Achim Rettinger · 2018
Later among the works it cites.
Discovering connotations as labels for weakly supervised image-sentence data
Aditya Mogadala, Bhargav Kanuparthi, Achim Rettinger, and York Sure-Vetter · 2018
Later among the works it cites.
Cnn+ cnn: Convolutional decoders for image captioning
Qingzhong Wang and Antoni B Chan · 2018
Later among the works it cites.
Object counts! bringing explicit detections back into image captioning
Josiah Wang, Pranava Swaroop Madhyastha, and Lucia Specia · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tseng-Hung Chen, Yuan-Hong Liao, Ching-Yao Chuang, Wan-Ting Hsu, Jianlong Fu, and Min Sun · 2017
Cited alongside, same era.
Estimation of gap between current language models and human performance
Xiaoyu Shen, Youssef Oualil, Clayton Greenberg, Mittul Singh, and Dietrich Klakow · 2017
Cited alongside, same era.
Speaking the same language: Matching machine to human captions by adversarial training
Rakshith Shetty, Marcus Rohrbach, Lisa Anne Hendricks, Mario Fritz, and Bernt Schiele · 2017
Cited alongside, same era.
Convolutional image captioning
Jyoti Aneja, Aditya Deshpande, and Alexander G Schwing · 2018
Cited alongside, same era.
Show, control and tell: a framework for generating controllable and grounded captions
Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara · 2019
Later among the works it cites.
Fast, diverse and accurate image captioning guided by part-of-speech
Aditya Deshpande, Jyoti Aneja, Liwei Wang, Alexander G Schwing, and David Forsyth · 2019
Later among the works it cites.
Select and attend: Towards controllable content selection in text generation
Xiaoyu Shen, Jun Suzuki, Kentaro Inui, Hui Su, Dietrich Klakow, and Satoshi Sekine · 2019
Later among the works it cites.