Fetching the paper…
Reading the bibliography…
It is encouraged to see that progress has been made to bridge videos and natural language.
A generalized framework of sequence generation with application to undirected sequence models
Mansimov, E.; Wang, A.; and Cho, K. 2019 · 1905
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Banerjee, S.; and Lavie, A. 2005 · 2005
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Bird, S.; Klein, E.; and Loper, E. 2009 · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D.; and Dolan, W. B. 2011 · 2011
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S.; Vinyals, O.; Jaitly, N.; and Shazeer, N. 2015 · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X.; Fang, H.; Lin, T.-Y.; Vedantam, R.; Gupta, S.; Dollár, P.; and Zitnick, C. L. 2015 · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe, S.; and Szegedy, C. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Training very deep networks
Srivastava, R. K.; Greff, K.; and Schmidhuber, J. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Translating Videos to Natural Language Using Deep Recurrent Neural Networks
Venugopalan, S.; Xu, H.; Donahue, J.; Rohrbach, M.; Mooney, R.; and Saenko, K. 2015 · 2015
Earlier work this paper cites.
Describing videos by exploiting temporal structure
Yao, L.; Torabi, A.; Cho, K.; Ballas, N.; Pal, C.; Larochelle, H.; and Courville, A. 2015 · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J.; Mei, T.; Yao, T.; and Rui, Y. 2016 · 2016
Cited alongside, same era.
Attention-based multimodal fusion for video description
Hori, C.; Hori, T.; Lee, T.-Y.; Zhang, Z.; Harsham, B.; Hershey, J. R.; Marks, T. K.; and Sumi, K. 2017 · 2017
Cited alongside, same era.
The kinetics human action video dataset
Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; et al. 2017 · 2017
Cited alongside, same era.
Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning
Aafaq, N.; Akhtar, N.; Liu, W.; Gilani, S. Z.; and Mian, A. 2019 · 2019
Closest in time.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Closest in time.
Mask-predict: Parallel decoding of conditional masked language models
Ghazvininejad, M.; Levy, O.; Liu, Y.; and Zettlemoyer, L. 2019 · 2019
Closest in time.
Levenshtein transformer
Gu, J.; Wang, C.; and Zhao, J. 2019 · 2019
Closest in time.
Non-autoregressive neural machine translation with enhanced decoder input
Guo, J.; Tan, X.; He, D.; Qin, T.; Xu, L.; and Liu, T.-Y. 2019 · 2019
Closest in time.
Aligning visual regions and textual concepts for semantic-grounded image representations
Liu, F.; Liu, Y.; Ren, X.; He, X.; and Sun, X. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Song, J.; Guo, Z.; Gao, L.; Liu, W.; Zhang, D.; and Shen, H. T. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Learning multimodal attention LSTM networks for video captioning
Xu, J.; Yao, T.; Zhang, Y.; and Mei, T. 2017 · 2017
Cited alongside, same era.
A neural compositional paradigm for image captioning
Dai, B.; Fidler, S.; and Lin, D. 2018 · 2018
Cited alongside, same era.
A Dataset for Telling the Stories of Social Media Videos
Gella, S.; Lewis, M.; and Rohrbach, M. 2018 · 2018
Cited alongside, same era.
Non-autoregressive neural machine translation
Gu, J.; Bradbury, J.; Xiong, C.; Li, V. O.; and Socher, R. 2018 · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Hara, K.; Kataoka, H.; and Satoh, Y. 2018 · 2018
Cited alongside, same era.
Closest in time.
Memory-attended recurrent network for video captioning
Pei, W.; Zhang, J.; Wang, X.; Ke, L.; Shen, X.; and Tai, Y.-W. 2019 · 2019
Closest in time.
STAT: Spatial-Temporal Attention Mechanism for Video Captioning
Yan, C.; Tu, Y.; Wang, X.; Zhang, Y.; Hao, X.; Zhang, Y.; and Dai, Q. 2019 · 2019
Closest in time.
Image captioning with end-to-end attribute detection and subsequent attributes prediction
Huang, Y.; Chen, J.; Ouyang, W.; Wan, W.; and Xue, Y. 2020 · 2020
Closest in time.
Spatio-Temporal Graph for Video Captioning with Knowledge Distillation
Pan, B.; Cai, H.; Huang, D.-A.; Lee, K.-H.; Gaidon, A.; Adeli, E.; and Niebles, J. C. 2020 · 2020
Closest in time.
Show, Edit and Tell: A Framework for Editing Image Captions
Sammani, F.; and Melas-Kyriazi, L. 2020 · 2020
Closest in time.
Minimizing the bag-of-ngrams difference for non-autoregressive neural machine translation
Shao, C.; Zhang, J.; Feng, Y.; Meng, F.; and Zhou, J. 2020 · 2020
Closest in time.