Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Original
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations, 2019
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Later among the works it cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang · 2019
Later among the works it cites.
Fine-grained action retrieval through multiple parts-of-speech embeddings
Michael Wray, Diane Larlus, Gabriela Csurka, and Dima Damen · 2019
Later among the works it cites.
Billion-scale semi-supervised learning for image classification, 2019
I. Zeki Yalniz, Hervé Jégou, Kan Chen, Manohar Paluri, and Dhruv Mahajan · 2019
Later among the works it cites.
Self-supervised multimodal versatile networks
Original
Jean-Baptiste Alayrac, A. Recasens, Rosália G. Schneider, R. Arandjelović, Jason Ramapuram, J. Fauw, Lucas Smaira, S. Dieleman, and Andrew Zisserman · 2020
Closest in time.
Self-supervised learning by cross-modal audio-video clustering
Humam Alwassel, Bruno Korbar, Dhruv Mahajan, Lorenzo Torresani, Bernard Ghanem, and Du Tran · 2020
Closest in time.
Noise estimation using density estimation for self-supervised multimodal learning
Original
Elad Amrani, Rami Ben-Ari, Daniel Rotman, and Alex Bronstein · 2020
Closest in time.
Unsupervised learning of visual features by contrasting cluster assignments
M. Caron, I. Misra, J. Mairal, Priya Goyal, P. Bojanowski, and Armand Joulin · 2020
Closest in time.
Electra: Pre-training text encoders as discriminators rather than generators, 2020
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning · 2020
Closest in time.
Virtex: Learning visual representations from textual annotations
Original
Karan Desai and Justin Johnson · 2020
Closest in time.
Omni-sourced webly-supervised learning for video recognition
Haodong Duan, Yue Zhao, Yuanjun Xiong, Wentao Liu, and Dahua Lin · 2020
Closest in time.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Closest in time.
Momentum contrast for unsupervised visual representation learning, 2020
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Closest in time.
Hard negative mixing for contrastive learning
Yannis Kalantidis, Mert Bulent Sariyildiz, Noe Pion, Philippe Weinzaepfel, and Diane Larlus · 2020
Closest in time.
Supervised contrastive learning, 2020
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan · 2020
Closest in time.
Video understanding as machine translation, 2020
Bruno Korbar, Fabio Petroni, Rohit Girdhar, and Lorenzo Torresani · 2020
Closest in time.
Univl: A unified video and language pre-training model for multimodal understanding and generation, 2020
Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Jason Li, Taroon Bharti, and Ming Zhou · 2020
Closest in time.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Closest in time.
Self-supervised learning of pretext-invariant representations
Ishan Misra and Laurens van der Maaten · 2020
Closest in time.
Audio-visual instance discrimination with cross-modal agreement
Original
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra · 2020
Closest in time.
Multi-modal self-supervision from generalized data transformations, 2020
Mandela Patrick, Yuki M. Asano, Polina Kuznetsova, Ruth Fong, João F. Henriques, Geoffrey Zweig, and Andrea Vedaldi · 2020
Closest in time.
Evolving losses for unsupervised video representation learning
AJ Piergiovanni, Anelia Angelova, and Michael S. Ryoo · 2020
Closest in time.
Avlnet: Learning audio-visual language representations from instructional videos
Original
Andrew Rouditchenko, Angie Boggust, David Harwath, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Rogerio Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, et al · 2020
Closest in time.
Learning visual representations with caption annotations, 2020
Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus · 2020
Closest in time.
Object relational graph with teacher-recommended learning for video captioning
Ziqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li, Peijin Wang, Weiming Hu, and Zheng-Jun Zha · 2020
Closest in time.
Actbert: Learning global-local video-text representations
Linchao Zhu and Yi Yang · 2020
Closest in time.