Clip2video: Mastering video-text retrieval via image clip
Original
Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
Transformers in vision: A survey
Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Original
Xiang Lisa Li and Percy Liang · 2021
Later among the works it cites.
The clear benchmark: Continual learning on real-world imagery
Zhiqiu Lin, Jia Shi, Deepak Pathak, and Deva Ramanan · 2021
Later among the works it cites.
Avalanche: an end-to-end library for continual learning
Vincenzo Lomonaco, Lorenzo Pellegrini, Andrea Cossu, Antonio Carta, Gabriele Graffieti, Tyler L. Hayes, Matthias De Lange, Marc Masana, Jary Pomponi, Gido van de Ven, Martin Mundt, Qi She, Keiland Cooper, Jeremy Forest, Eden Belouadah, Simone Calderara, German I. Parisi, Fabio Cuzzolin, Andreas Tolias, Simone Scardapane, Luca Antiga, Subutai Amhad, Adrian Popescu, Christopher Kanan, Joost van de Weijer, Tinne Tuytelaars, Davide Bacciu, and Davide Maltoni · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Encoders and ensembles for task-free continual learning
Original
Murray Shanahan, Christos Kaplanis, and Jovana Mitrović · 2021
Later among the works it cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Original
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Der: Dynamically expandable representation for class incremental learning
Shipeng Yan, Jiangwei Xie, and Xuming He · 2021
Later among the works it cites.
Filip: Fine-grained interactive language-image pre-training
Original
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Original
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al · 2021
Later among the works it cites.
Prototype augmentation and self-supervision for incremental learning
Original
Fei Zhu, Xu-Yao Zhang, Chuang Wang, Fei Yin, and Cheng-Lin Liu · 2021
Later among the works it cites.
Cvpr 2022 clear challenge: Leaderboards, 2022
AIcrowd · 2022
Closest in time.
Dytox: Transformers for continual learning with dynamic token expansion
Arthur Douillard, Alexandre Ramé, Guillaume Couairon, and Matthieu Cord · 2022
Closest in time.
Bridging the gap between object and image-level representations for open-vocabulary detection
Hanoona Rasheed, Muhammad Maaz, Muhammad Uzair Khattak, Salman Khan, and Fahad Shahbaz Khan · 2022
Closest in time.
Deep bayesian unsupervised lifelong learning
Tingting Zhao, Zifeng Wang, Aria Masoomi, and Jennifer Dy · 2022
Closest in time.