Fetching the paper…
Reading the bibliography…
With the advent of rich visual representations and pre-trained language models, video captioning has seen continuous improvement over time.
Papineni, K., Roukos, S., Ward, T., Zhu, W.: Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA. pp. 311–318. ACL (2002). https://doi.org/10.3115/1073083.1073135, https://aclanthology.org/P02-1040/
2002
Earlier work this paper cites.
Lin, C.: Rouge: A package for automatic evaluation of summaries. In Text summarization branches out: Proceedings of the ACL-04 workshop, volume 8. Barcelona, Spain (2004)
2004
Earlier work this paper cites.
Banerjee, S., Lavie, A.: METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In: Proceedings of the Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization@ACL 2005, Ann Arbor, Michigan, USA, June 29, 2005. pp. 65–72. Association for Computational Linguistics (2005), https://aclanthology.org/W05-0909/
2005
Earlier work this paper cites.
Bird, S., Klein, E., Loper, E.: Natural Language Processing with Python. O’Reilly (2009), http://www.oreilly.de/catalog/9780596516499/index.html
2009
Earlier work this paper cites.
Řehůřek, R., Sojka, P.: Software Framework for Topic Modelling with Large Corpora. In: Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks. pp. 45–50. ELRA, Valletta, Malta (May 2010), http://is.muni.cz/publication/884893/en
2010
Earlier work this paper cites.
Chen, D.L., Dolan, W.B.: Collecting highly parallel data for paraphrase evaluation. In: The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference, 19-24 June, 2011, Portland, Oregon, USA. pp. 190–200. The Association for Computer Linguistics (2011), https://aclanthology.org/P11-1020/
2011
Earlier work this paper cites.
Chen, D., Manning, C.D.: A fast and accurate dependency parser using neural networks. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, Qatar (2014)
2014
Earlier work this paper cites.
Bengio, S., Vinyals, O., Jaitly, N., Shazeer, N.: Scheduled sampling for sequence prediction with recurrent neural networks. Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015,Montreal, Quebec, Canada (2015)
2015
Earlier work this paper cites.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA (2015)
2015
Earlier work this paper cites.
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M.S., Berg, A.C., Li, F.: Imagenet large scale visual recognition challenge. International Journal of Computer Vision, (2015)
2015
Earlier work this paper cites.
Vedantam, R., Zitnick, C.L., Parikh, D.: Cider: Consensus-based image description evaluation. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015. pp. 4566–4575. IEEE Computer Society (2015). https://doi.org/10.1109/CVPR.2015.7299087, https://doi.org/10.1109/CVPR.2015.7299087
2015
Earlier work this paper cites.
Venugopalan, S., Rohrbach, M., Donahue, J., Mooney, R.J., Darrell, T., Saenko, K.: Sequence to sequence - video to text. In: 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015. pp. 4534–4542. IEEE Computer Society (2015). https://doi.org/10.1109/ICCV.2015.515, https://doi.org/10.1109/ICCV.2015.515
2015
Earlier work this paper cites.
Venugopalan, S., Xu, H., Donahue, J., Rohrbach, M., Mooney, R.J., Saenko, K.: Translating videos to natural language using deep recurrent neural networks. In: NAACL HLT 2015, The 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Denver, Colorado, USA, May 31 - June 5, 2015. pp. 1494–1504. The Association for Computational Linguistics (2015). https://doi.org/10.3115/v1/n15-1173, https://doi.org/10.3115/v1/n15-1173
2015
Earlier work this paper cites.
Yao, L., Torabi, A., Cho, K., Ballas, N., Pal, C.J., Larochelle, H., Courville, A.C.: Describing videos by exploiting temporal structure. In: 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015. pp. 4507–4515. IEEE Computer Society (2015). https://doi.org/10.1109/ICCV.2015.512, https://doi.org/10.1109/ICCV.2015.512
2015
Earlier work this paper cites.
Goyal, A., Lamb, A., Zhang, Y., Zhang, S., Courville, A.C., Bengio, Y.: Professor forcing: A new algorithm for training recurrent networks. Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, Barcelona, Spain (2016)
2016
Earlier work this paper cites.
Guillaume, L., Miguel, B., Sandeep, S., Kazuya, K., Chris, D.: Neural architectures for named entity recognition. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (2016)
2016
Earlier work this paper cites.
Pan, P., Xu, Z., Yang, Y., Wu, F., Zhuang, Y.: Hierarchical recurrent neural encoder for video representation with application to captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. (2016)
2016
Cited alongside, same era.
Xu, J., Mei, T., Yao, T., Rui, Y.: MSR-VTT: A large video description dataset for bridging video and language. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016. pp. 5288–5296. IEEE Computer Society (2016). https://doi.org/10.1109/CVPR.2016.571, https://doi.org/10.1109/CVPR.2016.571
2016
Cited alongside, same era.
Yu, H., Wang, J., Huang, Z., Yang, Y., Xu, W.: Video paragraph captioning using hierarchical recurrent neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016)
2016
Cited alongside, same era.
Bojanowski, P., Grave, E., Joulin, A., Mikolov, T.: Enriching word vectors with subword information. Trans. Assoc. Comput. Linguistics (2017)
Wang, J., Wang, W., Huang, Y., Wang, L., Tan, T.: M3: Multimodal memory modelling for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2018)
2018
Later among the works it cites.
Aafaq, N., Akhtar, N., Liu, W., Gilani, S.Z., Mian, A.: Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, pages 12487–12496 (2019)
2019
Later among the works it cites.
Pei, W., Zhang, J., Wang, X., Ke, L., Shen, X., Tai, Y.: Memory-attended recurrent network for video captioning. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019. pp. 8347–8356. Computer Vision Foundation / IEEE (2019)
2019
Later among the works it cites.
Wang, B., Ma, L., Zhang, W., Jiang, W., Wang, J., Liu, W.: Controllable video captioning with POS sequence guidance based on gated fusion network. In The IEEE International Conference on Computer Vision (ICCV) (2019)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Honnibal, M., Montani, I.: spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing (2017), to appear
2017
Cited alongside, same era.
Koehn, P., Knowles, R.: Six challenges for neural machine translation. In: Proceedings of the First Workshop on Neural Machine Translation, NMT@ACL 2017, Vancouver, Canada, August 4, 2017. pp. 28–39. Association for Computational Linguistics (2017). https://doi.org/10.18653/v1/w17-3204, https://doi.org/10.18653/v1/w17-3204
2017
Cited alongside, same era.
Ren, S., He, K., Girshick, R.B., Sun, J.: Faster R-CNN: towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. (2017)
2017
Cited alongside, same era.
Song, J., Gao, L., Guo, Z., Liu, W., Zhang, D., Shen, H.T.: Hierarchical lstm with adjusted temporal attention for video captioning. In Proceedings of the 26th International Joint Conference on Artificial Intelligence (2017)
2017
Cited alongside, same era.
Tu, Z., Liu, Y., Lu, Z., Liu, X., Li, H.: Context gates for neural machine translation. Trans. Assoc. Comput. Linguistics 5
2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In Advances in Neural Information Processing Systems (2017)
2017
Cited alongside, same era.
Xie, S., Girshick, R.B., Dollár, P., Tu, Z., He, K.: Aggregated residual transformations for deep neural networks. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, pages 5987–5995 (2017)
2017
Cited alongside, same era.
Bairui, W., Lin, M., Wei, Z., Wei, L.: Reconstruction network for video captioning. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2018)
2018
Cited alongside, same era.
2019
Later among the works it cites.
Wu, Y., Kirillov, A., Massa, F., Lo, W.Y., Girshick, R.: Detectron2. https://github.com/facebookresearch/detectron2 (2019)
2019
Later among the works it cites.
Zhang, J., Peng, Y.: Object-aware aggregation with bidirectional temporal graph for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2019)
2019
Later among the works it cites.
Müller, M., Rios, A., Sennrich, R.: Domain robustness in neural machine translation. In: Proceedings of the 14th Conference of the Association for Machine Translation in the Americas, AMTA 2020, Virtual, October 6-9, 2020. pp. 151–164. Association for Machine Translation in the Americas (2020), https://aclanthology.org/2020.amta-research.14/
2020
Later among the works it cites.
Pan, B., Cai, H., Huang, D., Lee, K., Gaidon, A., Adeli, E., Niebles, J.C.: Spatio-temporal graph for video captioning with knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2020)
2020
Later among the works it cites.
Wang, C., Sennrich, R.: On exposure bias, hallucination and domain shift in neural machine translation. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020. pp. 3544–3552. Association for Computational Linguistics (2020). https://doi.org/10.18653/v1/2020.acl-main.326, https://doi.org/10.18653/v1/2020.acl-main.326
2020
Later among the works it cites.
Zheng, Q., Wang, C., Tao, D.: Syntax-aware action targeting for video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2020)
2020
Later among the works it cites.
Ziqi, Z., Shi, Y., Yuan, C., Li, B., Wang, P., Hu, W., Zha, Z.: Object relational graph with teacher- recommended learning for video captioning. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2020)
2020
Later among the works it cites.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An image is worth 16x16 words: Transformers for image recognition at scale. 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria (2021)
2021
Later among the works it cites.
Raunak, V., Menezes, A., Junczys-Dowmunt, M.: The curious case of hallucinations in neural machine translation. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021. pp. 1172–1183. Association for Computational Linguistics (2021). https://doi.org/10.18653/v1/2021.naacl-main.92, https://doi.org/10.18653/v1/2021.naacl-main.92
2021
Later among the works it cites.
Ryu, H., Kang, S., Kang, H., Yoo, C.D.: Semantic grouping network for video captioning. Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021 (2021)
2021
Later among the works it cites.
Voita, E., Sennrich, R., Titov, I.: Analyzing the source and target contributions to predictions in neural machine translation. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021. pp. 1126–1140. Association for Computational Linguistics (2021). https://doi.org/10.18653/v1/2021.acl-long.91, https://doi.org/10.18653/v1/2021.acl-long.91
2021
Later among the works it cites.