Fetching the paper…
Reading the bibliography…
Video captioning is the task of automatically generating a textual description of the actions in a video.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K. K. Paliwal · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
The Stanford CoreNLP natural language processing toolkit
C. D. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. J. Bethard, and D. McClosky · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail
A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
deltableu: A discriminative metric for generation tasks with intrinsically diverse targets
M. Galley, C. Brockett, A. Sordoni, Y. Ji, M. Auli, C. Quirk, M. Mitchell, J. Gao, and B. Dolan · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Sequence to sequence-video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Cited alongside, same era.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko · 2015
Cited alongside, same era.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Later among the works it cites.
Video paragraph captioning using hierarchical recurrent neural networks
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu · 2016
Later among the works it cites.
Video captioning and retrieval models with semantic attention
Y. Yu, H. Ko, J. Choi, and G. Kim · 2016
Later among the works it cites.
Hierarchical boundary-aware neural encoder for video captioning
L. Baraldi, C. Grana, and R. Cucchiara · 2017
Closest in time.
Dense-captioning events in videos
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles · 2017
Closest in time.
Improved image captioning via policy gradient optimization of spider
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Zaremba and I. Sutskever · 2015
Cited alongside, same era.
An actor-critic algorithm for sequence prediction
D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. Courville, and Y. Bengio · 2016
Cited alongside, same era.
Semantic compositional networks for visual captioning
Z. Gan, C. Gan, X. He, Y. Pu, K. Tran, J. Gao, L. Carin, and L. Deng · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Describing videos using multi-modal fusion
Q. Jin, J. Chen, S. Chen, Y. Xiong, and A. Hauptmann · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum · 2016
Cited alongside, same era.
Hierarchical recurrent neural encoder for video representation with application to captioning
P. Pan, Z. Xu, Y. Yang, F. Wu, and Y. Zhuang · 2016
Cited alongside, same era.
Closest in time.
Multi-task video captioning with video and entailment generation
R. Pasunuru and M. Bansal · 2017
Closest in time.
Reinforced video captioning with entailment rewards
R. Pasunuru and M. Bansal · 2017
Closest in time.
Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning
B. Peng, X. Li, L. Li, J. Gao, A. Celikyilmaz, S. Lee, and K.-F. Wong · 2017
Closest in time.
Deep reinforcement learning-based image captioning with embedding reward
Z. Ren, X. Wang, N. Zhang, X. Lv, and L.-J. Li · 2017
Closest in time.
Self-critical sequence training for image captioning
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel · 2017
Closest in time.
Weakly supervised dense video captioning
Z. Shen, J. Li, Z. Su, M. Li, Y. Chen, Y.-G. Jiang, and X. Xue · 2017
Closest in time.
Hierarchical lstm with adjusted temporal attention for video captioning
J. Song, Z. Guo, L. Gao, W. Liu, D. Zhang, and H. T. Shen · 2017
Closest in time.
Feudal networks for hierarchical reinforcement learning
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Closest in time.
Actor-critic sequence training for image captioning
L. Zhang, F. Sung, F. Liu, T. Xiang, S. Gong, Y. Yang, and T. M. Hospedales · 2017
Closest in time.
Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning
X. Wang, Y.-F. Wang, and W. Y. Wang · 2018
Closest in time.