Fetching the paper…
Reading the bibliography…
Video captioning is the process of describing the content of a sequence of images capturing its semantic relationships and meanings.
Bertscore: Evaluating text generation with bert
Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., Artzi, Y., 2019a · 1904
Earlier work this paper cites.
Weakly supervised dense video captioning, in: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 1916–1924
Shen, Z., Li, J., Su, Z., Li, M., Chen, Y., Jiang, Y.G., Xue, X., 2017 · 1924
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Marcinkiewicz, M.A., 1994 · 1994
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp. 311–318
Papineni, K., Roukos, S., Ward, T., Zhu, W.J., 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries, in: Text summarization branches out, pp. 74–81
Lin, C.Y., 2004 · 2004
Earlier work this paper cites.
Understanding inverse document frequency: on theoretical arguments for idf
Robertson, S., 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments, in: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pp. 65–72
Banerjee, S., Lavie, A., 2005 · 2005
Earlier work this paper cites.
Using information theoretic vector quantization for inverted mfcc based speaker verification, in: 2009 2nd International Conference on Computer, Control and Communication, IEEE. pp. 1–5
Memon, S., Lech, M., He, L., 2009 · 2009
Earlier work this paper cites.
On the use of distributed dct in speaker identification, in: 2009 Annual IEEE India Conference, IEEE. pp. 1–4
Sahidullah, M., Saha, G., 2009 · 2009
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation, in: Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pp. 190–200
Chen, D., Dolan, W.B., 2011 · 2011
Earlier work this paper cites.
Natural language descriptions of visual scenes corpus generation and analysis, in: Proceedings of the Joint Workshop on Exploiting Synergies Between Information Retrieval and Machine Translation (ESIRMT) and Hybrid Approaches to Machine Translation (HyTra), pp. 38–47
Khan, M.U.G., Nawab, R.M.A., Gotoh, Y., 2012 · 2012
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition, in: Proceedings of the IEEE international conference on computer vision, pp. 2712–2719
Guadarrama, S., Krishnamoorthy, N., Malkarnenkar, G., Venugopalan, S., Mooney, R., Darrell, T., Saenko, K., 2013 · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Regneri, M., Rohrbach, M., Wetzel, D., Thater, S., Schiele, B., Pinkal, M., 2013 · 2013
Earlier work this paper cites.
Action recognition with improved trajectories, in: Proceedings of the IEEE international conference on computer vision, pp. 3551–3558
Wang, H., Schmid, C., 2013 · 2013
Earlier work this paper cites.
Distributed representations of sentences and documents, in: International conference on machine learning, PMLR. pp. 1188–1196
Le, Q., Mikolov, T., 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context, in: European conference on computer vision, Springer. pp. 740–755
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L., 2014 · 2014
Earlier work this paper cites.
The stanford corenlp natural language processing toolkit, in: Proceedings of 52nd annual meeting of the association for computational linguistics: system demonstrations, pp. 55–60
Manning, C.D., Surdeanu, M., Bauer, J., Finkel, J.R., Bethard, S., McClosky, D., 2014 · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail, in: German conference on pattern recognition, Springer. pp. 184–195
Rohrbach, A., Rohrbach, M., Qiu, W., Friedrich, A., Pinkal, M., Schiele, B., 2014 · 2014
Earlier work this paper cites.
Evaluation of video activity localizations integrating quality and quantity measurements
Wolf, C., Lombardi, E., Mille, J., Celiktutan, O., Jiu, M., Dogan, E., Eren, G., Baccouche, M., Dellandréa, E., Bichot, C.E., et al., 2014 · 2014
Earlier work this paper cites.
Delving deeper into convolutional networks for learning video representations
Ballas, N., Yao, L., Pal, C., Courville, A., 2015 · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding, in: Proceedings of the ieee conference on computer vision and pattern recognition, pp. 961–970
Caba Heilbron, F., Escorcia, V., Ghanem, B., Carlos Niebles, J., 2015 · 2015
Earlier work this paper cites.
Using descriptive video services to create a large data source for video annotation research
Torabi, A., Pal, C., Larochelle, H., Courville, A., 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4566–4575
Vedantam, R., Lawrence Zitnick, C., Parikh, D., 2015 · 2015
Earlier work this paper cites.
Early embedding and late reranking for video captioning, in: Proceedings of the 24th ACM international conference on Multimedia, pp. 1082–1086
Dong, J., Li, X., Lan, W., Huo, Y., Snoek, C.G., 2016 · 2016
Earlier work this paper cites.
Memory-augmented attention modelling for videos
Fakoor, R., Mohamed, A.r., Mitchell, M., Kang, S.B., Kohli, P., 2016 · 2016
Earlier work this paper cites.
Describing videos using multi-modal fusion, in: Proceedings of the 24th ACM international conference on Multimedia, pp. 1087–1091
Jin, Q., Chen, J., Chen, S., Xiong, Y., Hauptmann, A., 2016 · 2016
Earlier work this paper cites.
Multimodal video description, in: Proceedings of the 24th ACM international conference on Multimedia, pp. 1092–1096
Ramanishka, V., Das, A., Park, D.H., Venugopalan, S., Hendricks, L.A., Rohrbach, M., Saenko, K., 2016 · 2016
Earlier work this paper cites.
Frame-and segment-level features and candidate pool evaluation for video caption generation, in: Proceedings of the 24th ACM international conference on Multimedia, pp. 1073–1076
Shetty, R., Laaksonen, J., 2016 · 2016
Earlier work this paper cites.
Beyond caption to narrative: Video captioning with multiple sentences, in: 2016 IEEE International Conference on Image Processing (ICIP), IEEE. pp. 3364–3368
Shin, A., Ohnishi, K., Harada, T., 2016 · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding, in: European Conference on Computer Vision, Springer. pp. 510–526
Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., Gupta, A., 2016 · 2016
Earlier work this paper cites.
Improving lstm-based video description with linguistic knowledge mined from text
Venugopalan, S., Hendricks, L.A., Mooney, R., Saenko, K., 2016 · 2016
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5288–5296
Xu, J., Mei, T., Yao, T., Rui, Y., 2016 · 2016
Earlier work this paper cites.
Video paragraph captioning using hierarchical recurrent neural networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4584–4593
Yu, H., Wang, J., Huang, Z., Yang, Y., Xu, W., 2016 · 2016
Earlier work this paper cites.
Spatio-temporal attention models for grounded video captioning, in: asian conference on computer vision, Springer. pp. 104–119
Zanfir, M., Marinoiu, E., Sminchisescu, C., 2016 · 2016
Earlier work this paper cites.
Automatic video description generation via lstm with joint two-stream encoding, in: 2016 23rd International Conference on Pattern Recognition (ICPR), IEEE. pp. 2924–2929
Zhang, C., Tian, Y., 2016 · 2016
Earlier work this paper cites.
Hierarchical boundary-aware neural encoder for video captioning, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1657–1666
Baraldi, L., Grana, C., Cucchiara, R., 2017 · 2017
Earlier work this paper cites.
Video captioning with guidance of multimodal latent topics, in: Proceedings of the 25th ACM international conference on Multimedia, pp. 1838–1846
Chen, S., Chen, J., Jin, Q., Hauptmann, A., 2017 · 2017
Earlier work this paper cites.
Improving interpretability of deep neural networks with semantic information, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4306–4314
Dong, Y., Su, H., Zhu, J., Zhang, B., 2017 · 2017
Cited alongside, same era.
Semantic compositional networks for visual captioning, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5630–5639
Gan, Z., Gan, C., He, X., Pu, Y., Tran, K., Gao, J., Carin, L., Deng, L., 2017 · 2017
Cited alongside, same era.
Attention-based multimodal fusion for video description, in: Proceedings of the IEEE international conference on computer vision, pp. 4193–4202
Hori, C., Hori, T., Lee, T.Y., Zhang, Z., Harsham, B., Hershey, J.R., Marks, T.K., Sumi, K., 2017 · 2017
Cited alongside, same era.
Recurrent memory addressing for describing videos., in: CVPR Workshops, pp. 1–8
Jain, A.K., Agarwalla, A., Agrawal, K.K., Mitra, P., 2017 · 2017
Cited alongside, same era.
Dense-captioning events in videos, in: Proceedings of the IEEE international conference on computer vision, pp. 706–715
End-to-end video captioning with multitask reinforcement learning, in: 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE. pp. 339–348
Li, L., Gong, B., 2019 · 2019
Later among the works it cites.
Residual attention-based lstm for video captioning
Li, X., Zhou, Z., Chen, L., Gao, L., 2019 · 2019
Later among the works it cites.
Putting evaluation in context: Contextual embeddings improve machine translation evaluation, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 2799–2808
Mathur, N., Baldwin, T., Cohn, T., 2019 · 2019
Later among the works it cites.
Streamlined dense video captioning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6588–6597
Mun, J., Yang, L., Ren, Z., Xu, N., Han, B., 2019 · 2019
Later among the works it cites.
Memory-attended recurrent network for video captioning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8347–8356
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Krishna, R., Hata, K., Ren, F., Fei-Fei, L., Carlos Niebles, J., 2017 · 2017
Cited alongside, same era.
Mam-rnn: Multi-level attention model based rnn for video captioning., in: IJCAI, pp. 2208–2214
Li, X., Zhao, B., Lu, X., et al., 2017 · 2017
Cited alongside, same era.
Learning explicit video attributes from mid-level representation for video captioning
Nian, F., Li, T., Wang, Y., Wu, X., Ni, B., Xu, C., 2017 · 2017
Cited alongside, same era.
Video captioning with transferred semantic attributes, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6504–6512
Pan, Y., Yao, T., Li, H., Mei, T., 2017 · 2017
Cited alongside, same era.
Multi-task video captioning with video and entailment generation
Pasunuru, R., Bansal, M., 2017 · 2017
Cited alongside, same era.
Consensus-based sequence training for video captioning
Phan, S., Henter, G.E., Miyao, Y., Satoh, S., 2017 · 2017
Cited alongside, same era.
Movie description
Rohrbach, A., Torabi, A., Rohrbach, M., Tandon, N., Pal, C., Larochelle, H., Courville, A., Schiele, B., 2017 · 2017
Cited alongside, same era.
Hierarchical lstm with adjusted temporal attention for video captioning
Song, J., Guo, Z., Gao, L., Liu, W., Zhang, D., Shen, H.T., 2017 · 2017
Cited alongside, same era.
Pei, W., Zhang, J., Wang, X., Ke, L., Shen, X., Tai, Y.W., 2019 · 2019
Later among the works it cites.
Sports video captioning via attentive motion representation and group relationship modeling
Qi, M., Wang, Y., Li, A., Luo, J., 2019 · 2019
Later among the works it cites.
Watch it twice: Video captioning with a refocused video encoder, in: Proceedings of the 27th ACM International Conference on Multimedia, pp. 818–826
Shi, X., Cai, J., Joty, S., Gu, J., 2019 · 2019
Later among the works it cites.
Rich visual and language representation with complementary semantics for video captioning
Tang, P., Wang, H., Li, Q., 2019 · 2019
Later among the works it cites.
Convolutional reconstruction-to-sequence for video captioning
Wu, A., Han, Y., Yang, Y., Hu, Q., Wu, F., 2019 · 2019
Later among the works it cites.
Video captioning with adaptive attention and mixed loss optimization
Xiao, H., Shi, J., 2019 · 2019
Later among the works it cites.
Multi-guiding long short-term memory for video captioning
Xu, N., Liu, A.A., Nie, W., Su, Y., 2019 · 2019
Later among the works it cites.
Stat: Spatial-temporal attention mechanism for video captioning
Yan, C., Tu, Y., Wang, X., Zhang, Y., Hao, X., Zhang, Y., Dai, Q., 2019 · 2019
Later among the works it cites.
Cam-rnn: Co-attention model based rnn for video captioning
Zhao, B., Li, X., Lu, X., 2019 · 2019
Later among the works it cites.
Multi-modal dense video captioning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 958–959
Iashin, V., Rahtu, E., 2020 · 2020
Later among the works it cites.
A review of methods for video captioning
Kumar, A., Mathew, R., 2020 · 2020
Later among the works it cites.
Sibnet: Sibling convolutional encoder for video captioning
Liu, S., Ren, Z., Yuan, J., 2020 · 2020
Later among the works it cites.
Spatio-temporal graph for video captioning with knowledge distillation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10870–10879
Pan, B., Cai, H., Huang, D.A., Lee, K.H., Gaidon, A., Adeli, E., Niebles, J.C., 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, Online. pp. 7881–7892
Sellam, T., Das, D., Parikh, A., 2020 · 2020
Later among the works it cites.
Video captioning with boundary-aware hierarchical language decoding and joint video prediction
Shi, X., Cai, J., Gu, J., Joty, S., 2020 · 2020
Later among the works it cites.
An efficient framework for dense video captioning, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 12039–12046
Suin, M., Rajagopalan, A., 2020 · 2020
Later among the works it cites.
Exploiting the local temporal information for video captioning
Wei, R., Mi, L., Hu, Y., Chen, Z., 2020 · 2020
Later among the works it cites.
Video captioning with text-based dynamic attention and step-by-step learning
Xiao, H., Shi, J., 2020 · 2020
Later among the works it cites.
Controllable video captioning with an exemplar sentence, in: Proceedings of the 28th ACM International Conference on Multimedia, pp. 1085–1093
Yuan, Y., Ma, L., Wang, J., Zhu, W., 2020 · 2020
Later among the works it cites.
Object relational graph with teacher-recommended learning for video captioning, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 13278–13288
Zhang, Z., Shi, Y., Yuan, C., Li, B., Wang, P., Hu, W., Zha, Z.J., 2020 · 2020
Later among the works it cites.
Syntax-aware action targeting for video captioning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13096–13105
Zheng, Q., Wang, C., Tao, D., 2020 · 2020
Later among the works it cites.
Understanding objects in video: Object-oriented video captioning via structured trajectory and adversarial learning
Zhu, F., Hwang, J.N., Ma, Z., Chen, G., Guo, J., 2020 · 2020
Later among the works it cites.
Motion guided region message passing for video captioning, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1543–1552
Chen, S., Jiang, Y.G., 2021 · 2021
Later among the works it cites.
Natural language description of videos for smart surveillance
Dilawari, A., Khan, M.U.G., Al-Otaibi, Y.D., Rehman, Z.u., Rahman, A.u., Nam, Y., 2021 · 2021
Later among the works it cites.
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J., Holtzman, A., Forbes, M., Bras, R.L., Choi, Y., 2021 · 2021
Later among the works it cites.
Osvidcap: A framework for the simultaneous recognition and description of concurrent actions in videos in an open-set scenario
Inácio, A.D.S., Gutoski, M., Lazzaretti, A.E., Lopes, H.S., 2021 · 2021
Later among the works it cites.
Exploring video captioning techniques: A comprehensive survey on deep learning methods
Islam, S., Dash, A., Seum, A., Raj, A.H., Hossain, T., Shah, F.M., 2021 · 2021
Later among the works it cites.
A multi-instance multi-label dual learning approach for video captioning
Ji, W., Wang, R., 2021 · 2021
Later among the works it cites.
Video captioning based on channel soft attention and semantic reconstructor
Lei, Z., Huang, Y., 2021 · 2021
Later among the works it cites.
Improving video captioning with temporal composition of a visual-syntactic embedding, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 3039–3049
Perez-Martin, J., Bustos, B., Pérez, J., 2021 · 2021
Later among the works it cites.
Accelerated masked transformer for dense video captioning
Yu, Z., Han, N., 2021 · 2021
Later among the works it cites.
Video captioning: a review of theory, techniques and practices
Jain, V., Al-Turjman, F., Chaudhary, G., Nayar, D., Gupta, V., Kumar, A., 2022 · 2022
Closest in time.
Video captioning with attention-based lstm and semantic consistency
Gao, L., Guo, Z., Zhang, H., Xu, X., Shen, H.T., 2017 · 2055
Closest in time.