Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Later among the works it cites.
Streamlined dense video captioning
Jonghwan Mun, L. Yang, Zhou Ren, N. Xu, and Bohyung Han. 2019 · 2019
Later among the works it cites.
Dense procedure captioning in narrated instructional videos
Botian Shi, Lei Ji, Yaobo Liang, Nan Duan, Peng Chen, Zhendong Niu, and M. Zhou. 2019 · 2019
Later among the works it cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin P. Murphy, and Cordelia Schmid. 2019 · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Multi-modal dense video captioning
Vladimir Iashin and Esa Rahtu. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Sketch, ground, and refine: Top-down dense video captioning
Chaorui Deng, Shizhe Chen, Da Chen, Yuan He, and Qi Wu. 2021 · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021 · 2021
Later among the works it cites.
End-to-end dense video captioning with parallel decoding
Original
Teng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng, Ran Cheng, and Ping Luo. 2021 · 2021
Later among the works it cites.
Pix2seq: A language modeling framework for object detection
Ting Chen, Saurabh Saxena, Lala Li, David J. Fleet, and Geoffrey Hinton. 2022 · 2022
Closest in time.