Fetching the paper…
Reading the bibliography…
Video captioning automatically generates short descriptions of the video content, usually in form of a single sentence.
Binary codes capable of correcting deletions, insertions, and reversals, in: Soviet physics doklady, pp. 707–710
Levenshtein, V.I., 1966 · 1966
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation, in: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, ACL, pp. 311–318
Papineni, K., Roukos, S., Ward, T., Zhu, W.J., 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries, in: Text Summarization Branches Out, pp. 74–81
Lin, C.Y., 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments, in: Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pp. 65–72
Banerjee, S., Lavie, A., 2005 · 2005
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition, in: Proceedings of the IEEE international conference on computer vision, ICCV, pp. 2712–2719
Guadarrama, S., Krishnamoorthy, N., Malkarnenkar, G., Venugopalan, S., Mooney, R., Darrell, T., Saenko, K., 2013 · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation, in: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP, pp. 1724–1734
Cho, K., van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., Bengio, Y., 2014 · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation, in: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 4566–4575
Vedantam, R., Zitnick, C.L., Parikh, D., 2015 · 2015
Earlier work this paper cites.
The 1st video to language challenge
Xu, J., Mei, T., Yao, T., Rui, Y., 2016a · 2016
Cited alongside, same era.
Semantic compositional networks for visual captioning, in: Proceedings of the IEEE conference on computer vision and pattern recognition, CVPR, pp. 5630–5639
Gan, Z., Gan, C., He, X., Pu, Y., Tran, K., Gao, J., Carin, L., Deng, L., 2017 · 2017
Cited alongside, same era.
Reinforced video captioning with entailment rewards, in: Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP, pp. 979–985
Pasunuru, R., Bansal, M., 2017b · 2017
Cited alongside, same era.
Self-critical sequence training for image captioning, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 7008–7024
Rennie, S.J., Marcheret, E., Mroueh, Y., Ross, J., Goel, V., 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering, in: Proceedings of the IEEE conference on computer vision and pattern recognition, CVPR, pp. 6077–6086
Hierarchical attention-based multimodal fusion for video captioning
Wu, C., Wei, Y., Chu, X., Sun, W., Su, F., Wang, L., 2018 · 2018
Later among the works it cites.
Topic-oriented image captioning based on order-embedding
Yu, N., Hu, X., Song, B., Yang, J., Zhang, J., 2018 · 2018
Later among the works it cites.
ECO: efficient convolutional network for online video understanding, in: Proceedings of the European conference on computer vision ECCV, pp. 713–730
Zolfaghari, M., Singh, K., Brox, T., 2018 · 2018
Later among the works it cites.
Memory-attended recurrent network for video captioning, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 8347–8356
Pei, W., Zhang, J., Wang, X., Ke, L., Shen, X., Tai, Y.W., 2019 · 2019
Later among the works it cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research, in: 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 4580–4590
Wang, X., Wu, J., Chen, J., Li, L., Wang, Y.F., Wang, W.Y., 2019c · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L., 2018 · 2018
Cited alongside, same era.
Sibnet: Sibling convolutional encoder for video captioning, in: Proceedings of the 26th ACM International Conference on Multimedia, p. 1425–1434
Liu, S., Ren, Z., Yuan, J., 2018 · 2018
Cited alongside, same era.
Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning, in: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 795–801
Wang, X., Wang, Y., Wang, W.Y., 2018 · 2018
Cited alongside, same era.
A semantics-assisted video captioning model trained with scheduled sampling
Chen, H., Lin, K., Maye, A., Li, J., Hu, X., 2020b
Cited in the paper.
Multi-task video captioning with video and entailment generation, in: Barzilay, R., Kan, M. (Eds.), Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL, pp. 1273–1283
Pasunuru, R., Bansal, M., 2017a
Cited in the paper.
Controllable video captioning with pos sequence guidance based on gated fusion network, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV, pp. 2641–2650
Wang, B., Ma, L., Zhang, W., Jiang, W., Wang, J., Liu, W., 2019a
Cited in the paper.
Controllable video captioning with pos sequence guidance based on gated fusion network, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV, pp. 2641–2650
Wang, B., Ma, L., Zhang, W., Jiang, W., Wang, J., Liu, W., 2019b
Cited in the paper.
Learning to compose topic-aware mixture of experts for zero-shot video captioning, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 8965–8972
Wang, X., Wu, J., Zhang, D., Su, Y., Wang, W.Y., 2019d
Cited in the paper.
Later among the works it cites.
Delving deeper into the decoder for video captioning, in: ECAI 2020 - 24th European Conference on Artificial Intelligence, pp. 1079–1086
Chen, H., Li, J., Hu, X., 2020a · 2020
Later among the works it cites.
Open-book video captioning with retrieve-copy-generate network, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9837–9846
Zhang, Z., Qi, Z., Yuan, C., Shan, Y., Li, B., Deng, Y., Hu, W., 2021 · 2021
Closest in time.