Fetching the paper…
Reading the bibliography…
Video captioning, i.e.
Learning to forget: Continual prediction with LSTM
Gers, F. A., Schmidhuber, J. A., and Cummins, F. A. (2000) · 2000
Earlier work this paper cites.
Delving deeper into the decoder for video captioning
Chen, H., Li, J., and Hu, X. (2020) · 2001
Earlier work this paper cites.
Event detection and summarization in sports video
Li, B. and Sezan, M. I. (2001) · 2001
Earlier work this paper cites.
Natural language description of human activities from video images based on concept hierarchy of actions
Kojima, A., Tamura, T., and Fukunaga, K. (2002) · 2002
Earlier work this paper cites.
Feature extraction and a database strategy for video fingerprinting
Oostveen, J., Kalker, T., and Haitsma, J. (2002) · 2002
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002) · 2002
Earlier work this paper cites.
Methods of feature extraction of video sequences
Divakaran, A., Sun, H., and Ito, H. (2003) · 2003
Earlier work this paper cites.
Automatic soccer video analysis and summarization
Ekin, A., Tekalp, A. M., and Mehrotra, R. (2003) · 2003
Earlier work this paper cites.
Generic play-break event detection for summarization and hierarchical sports video analysis
Ekin, A. and Tekalp, M. (2003) · 2003
Earlier work this paper cites.
A rule-based style and grammar checker
Naber, D. (2003) · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. (2004) · 2004
Earlier work this paper cites.
Meteor: An automatic metric for MT evaluation with improved correlation with human judgments
Banerjee, S. and Lavie, A. (2005) · 2005
Earlier work this paper cites.
Columbia university TRECVID-2005 video search and high-level feature extraction
Chang, S.-F., Hsu, W. H., Kennedy, L. S., Xie, L., Yanagawa, A., Zavesky, E., and Zhang, D.-Q. (2005) · 2005
Earlier work this paper cites.
Re-evaluation the role of BLEU in machine translation research
Callison-Burch, C., Osborne, M., and Koehn, P. (2006) · 2006
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D. and Dolan, W. B. (2011) · 2011
Earlier work this paper cites.
Translating video content to natural language descriptions
Rohrbach, M., Qiu, W., Titov, I., Thater, S., Pinkal, M., and Schiele, B. (2013) · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Efficient feature extraction, encoding and classification for action recognition
Kantorov, V. and Laptev, I. (2014) · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D. (2014) · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V. (2014) · 2014
Cited alongside, same era.
Re-evaluating automatic summarization with BLEU and 192 shades of rouge
Graham, Y. (2015) · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D. (2015) · 2015
Cited alongside, same era.
Sequence to sequence-video to text
Venugopalan, S., Rohrbach, M., Donahue, J., Mooney, R., Darrell, T., and Saenko, K. (2015) · 2015
Hierarchical LSTM with adjusted temporal attention for video captioning
Song, J., Guo, Z., Gao, L., Liu, W., Zhang, D., and Shen, H. T. (2017) · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Later among the works it cites.
Sequence-to-sequence models can directly translate foreign speech
Weiss, R. J., Chorowski, J., Jaitly, N., Wu, Y., and Chen, Z. (2017) · 2017
Later among the works it cites.
Learning multimodal attention LSTM networks for video captioning
Xu, J., Yao, T., Zhang, Y., and Mei, T. (2017) · 2017
Later among the works it cites.
Trecvid 2018: Benchmarking video activity detection, video captioning and matching, video storytelling linking and video search
Awad, G., Butt, A., Curtis, K., Lee, Y., Fiscus, J., Godil, A., Joy, D., Delgado, A., Smeaton, A., and Graham, Y. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Convolutional LSTM network: A machine learning approach for precipitation nowcasting
Xingjian, S., Chen, Z., Wang, H., Yeung, D.-Y., Wong, W.-K., and Woo, W.-c. (2015) · 2015
Cited alongside, same era.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
Xu, R., Xiong, C., Chen, W., and Corso, J. J. (2015) · 2015
Cited alongside, same era.
Delving deeper into convolutional networks for learning video representations
Ballas, N., Yao, L., Pal, C., and Courville, A. C. (2016) · 2016
Cited alongside, same era.
Re-evaluating automatic metrics for image captioning
Kilickaya, M., Erdem, A., Ikizler-Cinbis, N., and Erdem, E. (2016) · 2016
Cited alongside, same era.
Professor forcing: A new algorithm for training recurrent networks
Lamb, A. M., Goyal, A. G. A. P., Zhang, Y., Zhang, S., Courville, A. C., and Bengio, Y. (2016) · 2016
Cited alongside, same era.
Jointly modeling embedding and translation to bridge video and language
Pan, Y., Mei, T., Yao, T., Li, H., and Rui, Y. (2016) · 2016
Cited alongside, same era.
Later among the works it cites.
Describing video with attention-based bidirectional LSTM
Bin, Y., Yang, Y., Shen, F., Xie, N., Shen, H. T., and Li, X. (2018) · 2018
Later among the works it cites.
Evaluation of automatic video captioning using direct assessment
Graham, Y., Awad, G., and Smeaton, A. (2018) · 2018
Later among the works it cites.
Video captioning with multi-faceted attention
Long, X., Gan, C., and de Melo, G. (2018) · 2018
Later among the works it cites.
A call for clarity in reporting BLEU scores
Post, M. (2018) · 2018
Later among the works it cites.
Interpretable video captioning via trajectory structured localization
Wu, X., Li, G., Cao, Q., Ji, Q., and Lin, L. (2018) · 2018
Later among the works it cites.
Video captioning by adversarial LSTM
Yang, Y., Zhou, J., Ai, J., Bin, Y., Hanjalic, A., Shen, H. T., and Ji, Y. (2018) · 2018
Later among the works it cites.
Video captioning using deep learning: An overview of methods, datasets and metrics
Amaresh, M. and Chitrakala, S. (2019) · 2019
Later among the works it cites.
A gentle introduction to calculating the BLEU score for text in python
Brownlee, J. (2019) · 2019
Later among the works it cites.
Memory-attended recurrent network for video captioning
Pei, W., Zhang, J., Wang, X., Ke, L., Shen, X., and Tai, Y.-W. (2019) · 2019
Later among the works it cites.
Stat: spatial-temporal attention mechanism for video captioning
Yan, C., Tu, Y., Wang, X., Zhang, Y., Hao, X., Zhang, Y., and Dai, Q. (2019) · 2019
Later among the works it cites.
Show, tell and summarize: Dense video captioning using visual cue aided sentence summarization
Zhang, Z., Xu, D., Ouyang, W., and Tan, C. (2019) · 2019
Later among the works it cites.
Object relational graph with teacher-recommended learning for video captioning
Zhang, Z., Shi, Y., Yuan, C., Li, B., Wang, P., Hu, W., and Zha, Z. J. (2020) · 2020
Closest in time.
Video captioning with attention-based LSTM and semantic consistency
Gao, L., Guo, Z., Zhang, H., Xu, X., and Shen, H. T. (2017) · 2055
Closest in time.