Fetching the paper…
Reading the bibliography…
Given the features of a video, recurrent neural networks can be used to automatically generate a caption for the video.
R. J. Williams and D. Zipser, “A learning algorithm for continually running fully recurrent neural networks,” Neural Computation , vol. 1, no. 2, pp. 270–280, 1989
1989
Earlier work this paper cites.
J. L. Elman, “Finding structure in time,” Cognitive Science , vol. 14, no. 2, pp. 179–211, 1990
1990
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, no. 3, pp. 229–256, 1992
1992
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
G. Tsoumakas and I. Katakis, “Multi-label classification: An overview,” International Journal of Data Warehousing and Mining , vol. 3, no. 3, pp. 1–13, 2007
2007
Earlier work this paper cites.
D. L. Chen and W. B. Dolan, “Collecting highly parallel data for paraphrase evaluation,” in Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics (ACL-2011) , Portland, OR, June 2011
2011
Earlier work this paper cites.
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. J. Mooney, T. Darrell, and K. Saenko, “Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition,” in IEEE International Conference on Computer Vision, ICCV, Sydney, Australia, December 1-8, 2013 , 2013, pp. 2712–2719
2013
Earlier work this paper cites.
K. Cho, B. van Merrienboer, Ç. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” in EMNLP, October 25-29, 2014, Doha, Qatar, A meeting of SIGDAT, a Special Interest Group of the ACL , 2014, pp. 1724–1734
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
S. Venugopalan, M. Rohrbach, J. Donahue, R. J. Mooney, T. Darrell, and K. Saenko, “Sequence to sequence - video to text,” in ICCV, Santiago, Chile, December 7-13, 2015 , 2015, pp. 4534–4542
2015
Earlier work this paper cites.
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, T. Darrell, and K. Saenko, “Long-term recurrent convolutional networks for visual recognition and description,” in IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, June 7-12, 2015 , 2015, pp. 2625–2634
2015
Earlier work this paper cites.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems, December 7-12, 2015, Montreal, Quebec, Canada , 2015, pp. 1171–1179
2015
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, June 7-12, 2015 , 2015, pp. 3156–3164
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui, “Jointly modeling embedding and translation to bridge video and language,” in IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, June 27-30, 2016 , 2016, pp. 4594–4602
2016
Cited alongside, same era.
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo, “Image captioning with semantic attention,” in IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, June 27-30, 2016 , 2016, pp. 4651–4659. [Online]. Available: https://doi.org/10.1109/CVPR.2016.503
2016
Cited alongside, same era.
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu, “Video paragraph captioning using hierarchical recurrent neural networks,” in IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, June 27-30, 2016 , 2016, pp. 4584–4593
2016
Cited alongside, same era.
A. Goyal, A. Lamb, Y. Zhang, S. Zhang, A. C. Courville, and Y. Bengio, “Professor forcing: A new algorithm for training recurrent networks,” in Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems, December 5-10, 2016, Barcelona, Spain , 2016, pp. 4601–4609
X. Wang, Y. Wang, and W. Y. Wang, “Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 2 (Short Papers) , 2018, pp. 795–801
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Zolfaghari, K. Singh, and T. Brox, “ECO: efficient convolutional network for online video understanding,” in Computer Vision - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part II , 2018, pp. 713–730. [Online]. Available: https://doi.org/10.1007/978-3-030-01216-8_43
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, J. Klingner, A. Shah, M. Johnson, X. Liu, Ł. Kaiser, S. Gouws, Y. Kato, T. Kudo, H. Kazawa, K. Stevens, G. Kurian, N. Patil, W. Wang, C. Young, J. Smith, J. Riesa, A. Rudnick, O. Vinyals, G. Corrado, M. Hughes, and J. Dean, “Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation,” Sep. 2016
2016
Cited alongside, same era.
J. Xu, T. Mei, T. Yao, and Y. Rui, “MSR-VTT: A large video description dataset for bridging video and language,” in IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, June 27-30, 2016 , 2016, pp. 5288–5296
2016
Cited alongside, same era.
Z. Gan, C. Gan, X. He, Y. Pu, K. Tran, J. Gao, L. Carin, and L. Deng, “Semantic compositional networks for visual captioning,” in IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, July 21-26, 2017 , 2017, pp. 1141–1150. [Online]. Available: https://doi.org/10.1109/CVPR.2017.127
2017
Cited alongside, same era.
L. Gao, Z. Guo, H. Zhang, X. Xu, and H. T. Shen, “Video captioning with attention-based lstm and semantic consistency,” IEEE Transactions on Multimedia , vol. 19, no. 9, pp. 2045–2055, 2017
2017
Cited alongside, same era.
R. Pasunuru and M. Bansal, “Reinforced video captioning with entailment rewards,” in EMNLP, Copenhagen, Denmark, September 9-11, 2017 , 2017, pp. 979–985
2017
Cited alongside, same era.
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017 , 2017, pp. 1179–1195. [Online]. Available: https://doi.org/10.1109/CVPR.2017.131
2017
Cited alongside, same era.
V. Ramanishka, A. Das, J. Zhang, and K. Saenko, “Top-down visual saliency guided by captions,” in IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, July 21-26, 2017 , 2017, pp. 3135–3144. [Online]. Available: https://doi.org/10.1109/CVPR.2017.334
2017
Cited alongside, same era.
R. Pasunuru and M. Bansal, “Multi-task video captioning with video and entailment generation,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Vancouver, Canada, July 30, 2017 - August 4, Volume 1: Long Papers , 2017, pp. 1273–1283. [Online]. Available: https://doi.org/10.18653/v1/P17-1117
2017
Cited alongside, same era.
S. Liu, Z. Ren, and J. Yuan, “Sibnet: Sibling convolutional encoder for video captioning,” in ACM Multimedia Conference on Multimedia Conference, Seoul, Republic of Korea, October 22-26, 2018 , 2018, pp. 1425–1434. [Online]. Available: https://doi.org/10.1145/3240508.3240667
2018
Later among the works it cites.
J. Yu, J. Li, Z. Yu, and Q. Huang, “Multimodal transformer with multi-view visual representation for image captioning,” IEEE Transactions on Circuits and Systems for Video Technology , pp. 1–1, 2019
2019
Closest in time.
M. Cornia, M. Stefanini, L. Baraldi, and R. Cucchiara, “Meshed-Memory Transformer for Image Captioning,” arXiv e-prints , Dec. 2019
2019
Closest in time.
B. Wang, L. Ma, W. Zhang, W. Jiang, J. Wang, and W. Liu, “Controllable video captioning with pos sequence guidance based on gated fusion network,” in The IEEE International Conference on Computer Vision (ICCV) , October 2019
2019
Closest in time.
W. Pei, J. Zhang, X. Wang, L. Ke, X. Shen, and Y.-W. Tai, “Memory-attended recurrent network for video captioning,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2019
2019
Closest in time.
J. Hou, X. Wu, W. Zhao, J. Luo, and Y. Jia, “Joint syntax representation learning and visual cue translation for video captioning,” in The IEEE International Conference on Computer Vision (ICCV) , October 2019
2019
Closest in time.
N. Aafaq, N. Akhtar, W. Liu, S. Z. Gilani, and A. Mian, “Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2019
2019
Closest in time.
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in IEEE/CVF International Conference on Computer Vision, Seoul, Korea (South), October 27 - November 2, 2019 . IEEE, 2019, pp. 7463–7472
2019
Closest in time.
X. Wang, J. Wu, D. Zhang, Y. Su, and W. Y. Wang, “Learning to compose topic-aware mixture of experts for zero-shot video captioning,” in The Thirty-Third AAAI Conference on Artificial Intelligence, The Thirty-First Innovative Applications of Artificial Intelligence Conference, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, Honolulu, Hawaii, USA, January 27 - February 1, 2019 . AAAI Press, 2019, pp. 8965–8972
2019
Closest in time.
Q. Zheng, C. Wang, and D. Tao, “Syntax-aware action targeting for video captioning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
Closest in time.
B. Pan, H. Cai, D.-A. Huang, K.-H. Lee, A. Gaidon, E. Adeli, and J. C. Niebles, “Spatio-temporal graph for video captioning with knowledge distillation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
Closest in time.
Z. Zhang, Y. Shi, C. Yuan, B. Li, P. Wang, W. Hu, and Z.-J. Zha, “Object relational graph with teacher-recommended learning for video captioning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
Closest in time.
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in Proceedings of the 32nd International Conference on Machine Learning, Lille, France, 6-11 July 2015 , 2015, pp. 2048–2057
2057
Closest in time.