Fetching the paper…
Reading the bibliography…
Current image captioning methods are usually trained via (penalized) maximum likelihood estimation.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. Mc Allester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
C.-Y. Lin and E. Hovy · 2003
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
How (not) to train your generative model: Scheduled sampling, likelihood, adversary?
F. Huszár · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Reinforcement learning neural turing machines-revised
W. Zaremba and I. Sutskever · 2015
Cited alongside, same era.
Reward augmented maximum likelihood for neural structured prediction
M. Norouzi, S. Bengio, Z. Chen, N. Jaitly, M. Schuster, Y. Wu, and D. Schuurmans · 2016
Closest in time.
Minimum risk training for neural machine translation
S. Shen, Y. Cheng, Z. He, W. He, H. Wu, M. Sun, and Y. Liu · 2016
Closest in time.
Rich image captioning in the wild
K. Tran, X. He, L. Zhang, J. Sun, C. Carapcea, C. Thrasher, C. Buehler, and C. Sienkiewicz · 2016
Closest in time.
Captioning images with diverse objects
S. Venugopalan, L. A. Hendricks, M. Rohrbach, R. Mooney, T. Darrell, and K. Saenko · 2016
Closest in time.
What value high level concepts in vision to language problems?
Q. Wu, C. Shen, A. van den Hengel, L. Liu, and A. Dick · 2016
Closest in time.
Review networks for caption generation
Z. Yang, Y. Yuan, Y. Wu, R. Salakhutdinov, and W. W. Cohen · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Cited alongside, same era.
An Actor-Critic algorithm for sequence prediction
D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. Courville, and Y. Bengio · 2016
Cited alongside, same era.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
R. Bernardi, R. Cakici, D. Elliott, A. Erdem, E. Erdem, N. Ikizler-Cinbis, F. Keller, A. Muscat, and B. Plank · 2016
Cited alongside, same era.
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra · 2016
Cited alongside, same era.
Generating visual explanations
L. A. Hendricks, Z. Akata, M. Rohrbach, J. Donahue, B. Schiele, and T. Darrell · 2016
Cited alongside, same era.
Boosting image captioning with attributes
T. Yao, Y. Pan, Y. Li, Z. Qiu, and T. Mei · 2016
Closest in time.
Image captioning with semantic attention
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Closest in time.
SeqGAN: Sequence generative adversarial nets with policy gradient
L. Yu, W. Zhang, J. Wang, and Y. Yu · 2016
Closest in time.
Automatic alt-text: Computer-generated image descriptions for blind users on a social network service
S. Wu, J. Wieland, O. Farivar, and J. Schiller · 2017
Closest in time.