Fetching the paper…
Reading the bibliography…
This paper explores image caption generation using conditional variational auto-encoders (CVAEs).
Mixture density networks
C. M. Bishop · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Approximating the kullback leibler divergence between gaussian mixture models
J. R. Hershey and P. A. Olsen · 2007
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Diverse M-Best Solutions in Markov Random Fields
D. Batra, P. Yadollahpour, A. Guzman-Rivera, and G. Shakhnarovich · 2012
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. Berg, K. Yamaguchi, T. Berg, K. Stratos, and H. Daumé III · 2012
Earlier work this paper cites.
Babytalk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2013
Earlier work this paper cites.
Generalizing image captions for image-text parallel corpus
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi · 2013
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Cited alongside, same era.
Language models for image captioning: The quirks and what works
J. Devlin, H. Cheng, H. Fang, S. Gupta, L. Deng, X. He, G. Zweig, and M. Mitchell · 2015
Cited alongside, same era.
Exploring nearest neighbor approaches for image captioning
J. Devlin, S. Gupta, R. Girshick, M. Mitchell, and C. L. Zitnick · 2015
Structured vaes: Composing probabilistic graphical models and variational autoencoders
M. J. Johnson, D. Duvenaud, A. Wiltschko, S. Datta, and R. Adams · 2016
Later among the works it cites.
Diverse beam search: Decoding diverse solutions from neural sequence models
A. K. Vijayakumar, M. Cogswell, R. R. Selvaraju, Q. Sun, S. Lee, D. Crandall, and D. Batra · 2016
Later among the works it cites.
Show and tell: Lessons learned from the 2015 mscoco image captioning challenge
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2016
Later among the works it cites.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Later among the works it cites.
Diverse image captioning via grouptalk
Z. Wang, F. Wu, W. Lu, J. Xiao, X. Li, Z. Zhang, and Y. Zhuang · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Learning structured output representation using deep conditional generative models
K. Sohn, H. Lee, and X. Yan · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Cited alongside, same era.
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Later among the works it cites.
Towards diverse and natural image descriptions via a conditional gan
B. Dai, D. Lin, R. Urtasun, and S. Fidler · 2017
Closest in time.
Learning diverse image colorization
A. Deshpande, J. Lu, M.-C. Yeh, and D. Forsyth · 2017
Closest in time.
Creativity: Generating diverse questions using variational autoencoders
U. Jain, Z. Zhang, and A. Schwing · 2017
Closest in time.
Categorical reparameterization with gumbel-softmax
E. Jang, S. Gu, and B. Poole · 2017
Closest in time.
Improved image captioning via policy gradient optimization of spider
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy · 2017
Closest in time.
Speaking the same language: Matching machine to human captions by adversarial training
R. Shetty, M. Rohrbach, L. A. Hendricks, M. Fritz, and B. Schiele · 2017
Closest in time.