Fetching the paper…
Reading the bibliography…
Fine-tuning visual models has been widely shown promising performance on many downstream visual tasks.
Language models are few-shot learners. In Proc. NeurIPS , Vol. 33. 1877–1901
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Receptive fields of single neurones in the cat’s striate cortex
David H Hubel and Torsten N Wiesel. 1959 · 1959
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal. 1997 · 1997
Earlier work this paper cites.
A multilinear singular value decomposition
Lieven De Lathauwer, Bart De Moor, and Joos Vandewalle. 2000 · 2000
Earlier work this paper cites.
PCA-SIFT: A more distinctive representation for local image descriptors. In Proc. CVPR , Vol. 2. IEEE, II–II
Yan Ke and Rahul Sukthankar. 2004 · 2004
Earlier work this paper cites.
A Modern Introduction to Probability and Statistics: Understanding why and how . Vol. 488
Frederik Michel Dekking, Cornelis Kraaikamp, Hendrik Paul Lopuhaä, and Ludolf Erwin Meester. 2005 · 2005
Earlier work this paper cites.
Model compression. In ACM SIGKDD . 535–541
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. 2006 · 2006
Earlier work this paper cites.
Scaling and assessment of data quality
Philip Evans. 2006 · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. In 2009 CVPR . IEEE, 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Brief: Binary robust independent elementary features. In Proc. ECCV . Springer, 778–792
Michael Calonder, Vincent Lepetit, Christoph Strecha, and Pascal Fua. 2010 · 2010
Earlier work this paper cites.
Vision: A computational investigation into the human representation and processing of visual information
David Marr. 2010 · 2010
Earlier work this paper cites.
Artificial intelligence a modern approach
Stuart J Russell. 2010 · 2010
Earlier work this paper cites.
ORB: An efficient alternative to SIFT or SURF. In Proc. ICCV . IEEE, 2564–2571
Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. 2011 · 2011
Earlier work this paper cites.
Data security and privacy protection issues in cloud computing. In 2012 international conference on computer science and electronics engineering , Vol. 1. IEEE, 647–651
Deyan Chen and Hong Zhao. 2012 · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent. 2013 · 2013
Earlier work this paper cites.
Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. 2013 · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?. In Proc. NeurIPS , Vol. 27
Jimmy Ba and Rich Caruana. 2014 · 2014
Earlier work this paper cites.
The ecological approach to visual perception: classic edition
James J Gibson. 2014 · 2014
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Yunchao Gong, Liu Liu, Ming Yang, and Lubomir Bourdev. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context. In Proc. ECCV . Springer, 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Data security and privacy in cloud computing
Yunchuan Sun, Junsheng Zhang, Yongping Xiong, and Guangyu Zhu. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Brenden M Lake, Ruslan Salakhutdinov, and Joshua B Tenenbaum. 2015 · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets. In Proc. ICLR
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Going deeper with convolutions. In Proc. CVPR . 1–9
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks. In Proc. ICCV . 4489–4497
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015 · 2015
Earlier work this paper cites.
Net2net: Accelerating learning via knowledge transfer. In Proc. ICLR
Tianqi Chen, Ian Goodfellow, and Jonathon Shlens. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proc. CVPR . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Fast algorithms for convolutional neural networks. In Proc. CVPR . 4013–4021
Andrew Lavin and Scott Gray. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision. In Proc. CVPR . 2818–2826
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
Learning structured sparsity in deep neural networks. In Proc. NeurIPS , Vol. 29
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2016 · 2016
Earlier work this paper cites.
Understanding of a convolutional neural network. In Proc. ICET . IEEE, 1–6
Saad Albawi, Tareq Abed Mohammed, and Saad Al-Zawi. 2017 · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset. In Proc. CVPR . 6299–6308
Joao Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2017 · 2017
Earlier work this paper cites.
Mimicking very efficient network for object detection. In Proc. CVPR . 6356–6364
Quanquan Li, Shengying Jin, and Junjie Yan. 2017 · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters. In Proc. NeurIPS
Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017 · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era. In Proc. ICCV . 843–852
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. 2017 · 2017
Earlier work this paper cites.
N2n learning: Network to network compression via policy gradient reinforcement learning. In Proc. ICLR
Anubhav Ashok, Nicholas Rhinehart, Fares Beainy, and Kris M Kitani. 2018 · 2018
Earlier work this paper cites.
Simple and efficient architecture search for convolutional neural networks. In Proc. ICLR, Workshop Track
Thomas Elsken, Jan-Hendrik Metzen, and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
Paraphrasing complex network: Network compression via factor transfer. In Proc. NeurIPS , Vol. 31
Jangho Kim, SeongUk Park, and Nojun Kwak. 2018 · 2018
Earlier work this paper cites.
Efficient neural architecture search via parameters sharing. In Proc. ICML . PMLR, 4095–4104
Hieu Pham, Melody Guan, Barret Zoph, Quoc Le, and Jeff Dean. 2018 · 2018
Earlier work this paper cites.
Efficient parametrization of multi-domain deep neural networks. In Proc. CVPR . 8119–8127
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2018 · 2018
Earlier work this paper cites.
Incremental learning through deep adaptation
Amir Rosenfeld and John K Tsotsos. 2018 · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks. In Proc. CVPR . 4510–4520
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018 · 2018
Earlier work this paper cites.
A survey on deep transfer learning. In International conference on artificial neural networks . Springer, 270–279
Chuanqi Tan, Fuchun Sun, Tao Kong, Wenchang Zhang, Chao Yang, and Chunfang Liu. 2018 · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In Proc. ECCV . 305–321
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. 2018 · 2018
Earlier work this paper cites.
Spatial temporal graph convolutional networks for skeleton-based action recognition. In Proc. AAAI , Vol. 32
Sijie Yan, Yuanjun Xiong, and Dahua Lin. 2018 · 2018
Earlier work this paper cites.
DATA: Differentiable ArchiTecture Approximation. In Proc. NeurIPS . 874–884
Jianlong Chang, Xinbang Zhang, Yiwen Guo, Gaofeng Meng, Shiming Xiang, and Chunhong Pan. 2019 · 2019
Earlier work this paper cites.
Progressive differentiable architecture search: Bridging the depth gap between search and evaluation. In Proc. CVPR . 1294–1303
Xin Chen, Lingxi Xie, Jun Wu, and Qi Tian. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP. In Proc. ICML . PMLR, 2790–2799
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Darts+: Improved differentiable architecture search with early stopping
Hanwen Liang, Shifeng Zhang, Jiacheng Sun, Xingqiu He, Weiran Huang, Kechen Zhuang, and Zhenguo Li. 2019 · 2019
Earlier work this paper cites.
Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding
Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot. 2019c · 2019
Earlier work this paper cites.
Relational knowledge distillation. In Proc. CVPR . 3967–3976
Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho. 2019 · 2019
Earlier work this paper cites.
Correlation congruence for knowledge distillation. In Proc. CVPR . 5007–5016
Baoyun Peng, Xiao Jin, Jiaheng Liu, Dongsheng Li, Yichao Wu, Yu Liu, Shunfeng Zhou, and Zhaoning Zhang. 2019 · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks. In Proc. ICML . PMLR, 6105–6114
Mingxing Tan and Quoc Le. 2019 · 2019
Earlier work this paper cites.
SNAS: stochastic neural architecture search. In Proc. ICLR
Sirui Xie, Hehui Zheng, Chunxiao Liu, and Liang Lin. 2019 · 2019
Earlier work this paper cites.
Bridging theory and algorithm for domain adaptation. In Proc. ICML . PMLR, 7404–7413
Yuchen Zhang, Tianle Liu, Mingsheng Long, and Michael Jordan. 2019 · 2019
Earlier work this paper cites.
Learning student networks via feature embedding
Hanting Chen, Yunhe Wang, Chang Xu, Chao Xu, and Dacheng Tao. 2020b · 2020
Earlier work this paper cites.
FNA++: Fast network adaptation via parameter remapping and architecture search
Jiemin Fang, Yuzhu Sun, Qian Zhang, Kangjian Peng, Yuan Li, Wenyu Liu, and Xinggang Wang. 2020b · 2020
Earlier work this paper cites.
X3d: Expanding architectures for efficient video recognition. In Proc. CVPR . 203–213
Christoph Feichtenhofer. 2020 · 2020
Earlier work this paper cites.
Differentiable feature aggregation search for knowledge distillation. In ECCV 16 . Springer, 469–484
Yushuo Guan, Pengyu Zhao, Bingxuan Wang, Yuanxing Zhang, Cong Yao, Kaigui Bian, and Jian Tang. 2020 · 2020
Earlier work this paper cites.
Milenas: Efficient neural architecture search via mixed-level reformulation. In Proc. CVPR . 11993–12002
Chaoyang He, Haishan Ye, Li Shen, and Tong Zhang. 2020 · 2020
Earlier work this paper cites.
Secure, privacy-preserving and federated machine learning in medical imaging
Georgios A Kaissis, Marcus R Makowski, Daniel Rückert, and Rickmer F Braren. 2020 · 2020
Earlier work this paper cites.
Sgas: Sequential greedy architecture search. In Proc. CVPR . 1620–1630
Guohao Li, Guocheng Qian, Itzel C Delgadillo, Matthias Muller, Ali Thabet, and Bernard Ghanem. 2020 · 2020
Earlier work this paper cites.
Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning. In Findings of the Association for Computational Linguistics: EMNLP 2020 . 441–459
Zhaojiang Lin, Andrea Madotto, and Pascale Fung. 2020 · 2020
Earlier work this paper cites.
Structured knowledge distillation for dense prediction
Yifan Liu, Changyong Shu, Jingdong Wang, and Chunhua Shen. 2020c · 2020
Earlier work this paper cites.
Heterogeneous knowledge distillation using information flow modeling. In Proc. CVPR . 2339–2348
Nikolaos Passalis, Maria Tzelepi, and Anastasios Tefas. 2020 · 2020
Earlier work this paper cites.
Adapterhub: A framework for adapting transformers
Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020 · 2020
Earlier work this paper cites.
Robust re-identification by multiple views knowledge distillation. In Proc. ECCV . Springer, 93–110
Angelo Porrello, Luca Bergamini, and Simone Calderara. 2020 · 2020
Earlier work this paper cites.
On the theory of transfer learning: The importance of task diversity. In Proc. NeurIPS , Vol. 33. 7852–7862
Nilesh Tripuraneni, Michael Jordan, and Chi Jin. 2020 · 2020
Earlier work this paper cites.
Knowledge distillation via adaptive instance normalization
Jing Yang, Brais Martinez, Adrian Bulat, and Georgios Tzimiropoulos. 2020 · 2020
Earlier work this paper cites.
A comprehensive survey on transfer learning
Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He. 2020 · 2020
Earlier work this paper cites.
Hyperparameter optimization: Foundations, algorithms, best practices, and open challenges
Bernd Bischl, Martin Binder, Michel Lang, Tobias Pielok, Jakob Richter, Stefan Coors, Janek Thomas, Theresa Ullmann, Marc Becker, Anne-Laure Boulesteix, et al · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Conditional positional encodings for vision transformers
Xiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. 2021b · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale. In Proc. ICLR
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Earlier work this paper cites.
MAGMA–Multimodal Augmentation of Generative Models through Adapter-based Finetuning
Constantin Eichenberg, Sidney Black, Samuel Weinbach, Letitia Parcalabescu, and Anette Frank. 2021 · 2021
Earlier work this paper cites.
Clip-adapter: Better vision-language models with feature adapters
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. 2021 · 2021
Earlier work this paper cites.
Open-vocabulary object detection via vision and language knowledge distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. 2021 · 2021
Earlier work this paper cites.
Distilling object detectors via decoupled features. In Proc. CVPR . 2154–2164
Jianyuan Guo, Kai Han, Yunhe Wang, Han Wu, Xinghao Chen, Chunjing Xu, and Chang Xu. 2021 · 2021
Earlier work this paper cites.
Dynamic neural networks: A survey
Yizeng Han, Gao Huang, Shiji Song, Le Yang, Honghui Wang, and Yulin Wang. 2021a · 2021
Cited alongside, same era.
Boosting light-weight depth estimation via knowledge distillation
Junjie Hu, Chenyou Fan, Hualie Jiang, Xiyue Guo, Yuan Gao, Xiangyong Lu, and Tin Lun Lam. 2021 · 2021
Cited alongside, same era.
Shuffle transformer: Rethinking spatial shuffle for vision transformer
Zilong Huang, Youcheng Ben, Guozhong Luo, Pei Cheng, Gang Yu, and Bin Fu. 2021 · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision. In Proc. ICML . PMLR, 4904–4916
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Cited alongside, same era.
Compacter: Efficient low-rank hypercomplex adapter layers. In Proc. NeurIPS , Vol. 34. 1022–1035
Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. 2021 · 2021
Knowledge distillation for multi-target domain adaptation in real-time person re-identification. In ICIP . IEEE, 3853–3557
Félix Remigereau, Djebril Mekhazni, Sajjad Abdoli, Rafael MO Cruz, Eric Granger, et al · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models. In Proc. CVPR . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Later among the works it cites.
Prefix conditioning unifies language and label supervision
Kuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li, Chen-Yu Lee, Kate Saenko, and Tomas Pfister. 2022 · 2022
Later among the works it cites.
Multitask Vision-Language Prompt Tuning
Sheng Shen, Shijia Yang, Tianjun Zhang, Bohan Zhai, Joseph E Gonzalez, Kurt Keutzer, and Trevor Darrell. 2022 · 2022
Later among the works it cites.
Towards Parameter-Efficient Integration of Pre-Trained Language Models In Temporal Video Grounding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
How to Adapt Your Large-Scale Vision-and-Language Model
Konwoo Kim, Michael Laskin, Igor Mordatch, and Deepak Pathak. 2021 · 2021
Cited alongside, same era.
Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution. In Proc. ICLR
Ananya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma, and Percy Liang. 2021 · 2021
Cited alongside, same era.
Benchmarking detection transfer learning with vision transformers
Yanghao Li, Saining Xie, Xinlei Chen, Piotr Dollar, Kaiming He, and Ross Girshick. 2021c · 2021
Cited alongside, same era.
A survey of convolutional neural networks: analysis, applications, and prospects
Zewen Li, Fan Liu, Wenjie Yang, Shouheng Peng, and Jun Zhou. 2021b · 2021
Cited alongside, same era.
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021d · 2021
Cited alongside, same era.
Exploring a large-scale multi-modal transportation recommendation system
Yang Liu, Cheng Lyu, Zhiyuan Liu, and Jinde Cao. 2021c · 2021
Cited alongside, same era.
A simple long-tailed recognition baseline via vision-language model
Teli Ma, Shijie Geng, Mengmeng Wang, Jing Shao, Jiasen Lu, Hongsheng Li, Peng Gao, and Yu Qiao. 2021 · 2021
Cited alongside, same era.
Erica K Shimomoto, Edison Marrese-Taylor, Hiroya Takamura, Ichiro Kobayashi, Hideki Nakayama, and Yusuke Miyao. 2022 · 2022
Later among the works it cites.
Test-time prompt tuning for zero-shot generalization in vision-language models
Manli Shu, Weili Nie, De-An Huang, Zhiding Yu, Tom Goldstein, Anima Anandkumar, and Chaowei Xiao. 2022 · 2022
Later among the works it cites.
Visual Prompt Tuning for Generative Transfer Learning
Kihyuk Sohn, Yuan Hao, José Lezama, Luisa Polania, Huiwen Chang, Han Zhang, Irfan Essa, and Lu Jiang. 2022 · 2022
Later among the works it cites.
Dualcoop: Fast adaptation to multi-label recognition with limited annotations
Ximeng Sun, Ping Hu, and Kate Saenko. 2022 · 2022
Later among the works it cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2022 · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022 · 2022
Later among the works it cites.
Cheng-Hao Tu, Zheda Mai, and Wei-Lun Chao. 2022 · 2022
Later among the works it cites.
Mojtaba Valipour, Mehdi Rezagholizadeh, Ivan Kobyzev, and Ali Ghodsi. 2022 · 2022
Later among the works it cites.
Language models generalize beyond natural proteins
Robert Verkuil, Ori Kabeli, Yilun Du, Basile IM Wicky, Lukas F Milles, Justas Dauparas, David Baker, Sergey Ovchinnikov, Tom Sercu, and Alexander Rives. 2022 · 2022
Later among the works it cites.
Learning to Decompose Visual Features with Latent Textual Prompts
Feng Wang, Manling Li, Xudong Lin, Hairong Lv, Alexander G Schwing, and Heng Ji. 2022d · 2022
Later among the works it cites.
Fine-grained Retrieval Prompt Tuning
Shijie Wang, Jianlong Chang, Zhihui Wang, Haojie Li, Wanli Ouyang, and Qi Tian. 2022b · 2022
Later among the works it cites.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al · 2022
Later among the works it cites.
Images Speak in Images: A Generalist Painter for In-Context Visual Learning
Xinlong Wang, Wen Wang, Yue Cao, Chunhua Shen, and Tiejun Huang. 2022e · 2022
Later among the works it cites.
S-Prompts Learning with Pre-trained Transformers: An Occam’s Razor for Domain Incremental Learning
Yabin Wang, Zhiwu Huang, and Xiaopeng Hong. 2022c · 2022
Later among the works it cites.
P2p: Tuning pre-trained image models for point cloud analysis with point-to-pixel prompting
Ziyi Wang, Xumin Yu, Yongming Rao, Jie Zhou, and Jiwen Lu. 2022f · 2022
Later among the works it cites.
Pruning Adapters with Lottery Ticket
Jiarun Wu and Qingliang Chen. 2022 · 2022
Later among the works it cites.
Unleashing the Power of Visual Prompting At the Pixel Level
Junyang Wu, Xianhang Li, Chen Wei, Huiyu Wang, Alan Yuille, Yuyin Zhou, and Cihang Xie. 2022a · 2022
Later among the works it cites.
Class-aware visual prompt tuning for vision-language pre-trained model
Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang, and Yanning Zhang. 2022 · 2022
Later among the works it cites.
Towards a Unified View on Visual Parameter-Efficient Transfer Learning
Bruce XB Yu, Jianlong Chang, Lingbo Liu, Qi Tian, and Chang Wen Chen. 2022a · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. 2022b · 2022
Later among the works it cites.
Volo: Vision outlooker for visual recognition
Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, and Shuicheng Yan. 2022a · 2022
Later among the works it cites.
Sha Yuan, Hanyu Zhao, Shuai Zhao, Jiahong Leng, Yangxiao Liang, Xiaozhi Wang, Jifan Yu, Xin Lv, Zhou Shao, Jiaao He, et al · 2022
Later among the works it cites.
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models. In Proc. ACL (Volume 2: Short Papers) . 1–9
Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. 2022 · 2022
Later among the works it cites.
Unified vision and language prompt learning
Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, and Chen Change Loy. 2022 · 2022
Later among the works it cites.
Feature-Proxy Transformer for Few-Shot Segmentation
Jian-Wei Zhang, Yifan Sun, Yi Yang, and Wei Chen. 2022e · 2022
Later among the works it cites.
Collaboration of Pre-trained Models Makes Better Few-shot Learner
Renrui Zhang, Hanqiu Deng, Bohao Li, Wei Zhang, Hao Dong, Hongsheng Li, Peng Gao, and Yu Qiao. 2022a · 2022
Later among the works it cites.
Prompting through Prototype: A Prototype-based Prompt Learning on Pretrained Vision-Language Models
Yue Zhang, Hongliang Fei, Dingcheng Li, Tan Yu, and Ping Li. 2022b · 2022
Later among the works it cites.
Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. 2022f · 2022
Later among the works it cites.
Zhengkun Zhang, Wenya Guo, Xiaojun Meng, Yasheng Wang, Yadao Wang, Xin Jiang, Qun Liu, and Zhenglu Yang. 2022c · 2022
Later among the works it cites.
Learning Domain Invariant Prompt for Vision-Language Models
Cairong Zhao, Yubin Wang, Xinyang Jiang, Yifei Shen, Kaitao Song, Dongsheng Li, and Duoqian Miao. 2022b · 2022
Later among the works it cites.
Localization distillation for dense object detection. In Proc. CVPR . 9407–9416
Zhaohui Zheng, Rongguang Ye, Ping Wang, Dongwei Ren, Wangmeng Zuo, Qibin Hou, and Ming-Ming Cheng. 2022 · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022b · 2022
Later among the works it cites.
ZegCLIP: Towards Adapting CLIP for Zero-shot Semantic Segmentation
Ziqin Zhou, Bowen Zhang, Yinjie Lei, Lingqiao Liu, and Yifan Liu. 2022c · 2022
Later among the works it cites.
Prompt-aligned gradient for prompt tuning
Beier Zhu, Yulei Niu, Yucheng Han, Yue Wu, and Hanwang Zhang. 2022a · 2022
Later among the works it cites.
PointCLIP V2: Adapting CLIP for Powerful 3D Open-world Learning
Xiangyang Zhu, Renrui Zhang, Bowei He, Ziyao Zeng, Shanghang Zhang, and Peng Gao. 2022b · 2022
Later among the works it cites.
\ \backslash A-la-carte Prompt Tuning (APT): Combining Distinct Data Via Composable Prompting
Benjamin Bowman, Alessandro Achille, Luca Zancato, Matthew Trager, Pramuditha Perera, Giovanni Paolini, and Stefano Soatto. 2023 · 2023
Closest in time.
NVIDIA Hopper H100 GPU: Scaling Performance
Jack Choquette. 2023 · 2023
Closest in time.
Scaling Vision Transformers to 22 Billion Parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al · 2023
Closest in time.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al · 2023
Closest in time.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al · 2023
Closest in time.
Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection
Kaifeng Gao, Long Chen, Hanwang Zhang, Jun Xiao, and Qianru Sun. 2023 · 2023
Closest in time.
Consolidator: Mergable Adapter with Group Connections for Visual Adaptation. In Proc. ICLR
Tianxiang Hao, Hui Chen, Yuchen Guo, and Guiguang Ding. 2023 · 2023
Closest in time.
ZegOT: Zero-shot Segmentation Through Optimal Transport of Text Prompts
Kwanyoung Kim, Yujin Oh, and Jong Chul Ye. 2023b · 2023
Closest in time.
Prompt Tuning of Deep Neural Networks for Speaker-adaptive Visual Speech Recognition
Minsu Kim, Hyung-Il Kim, and Yong Man Ro. 2023a · 2023
Closest in time.
Being Comes from Not-being: Open-vocabulary Text-to-Motion Generation with Wordless Training. In Proc. CVPR
Junfan Lin, Jianlong Chang, Lingbo Liu, Guanbin Li, Liang Lin, Qi Tian, and Chang-wen Chen. 2023 · 2023
Closest in time.
Shiwei Liu and Zhangyang Wang. 2023 · 2023
Closest in time.
UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling
Haoyu Lu, Mingyu Ding, Yuqi Huo, Guoxing Yang, Zhiwu Lu, Masayoshi Tomizuka, and Wei Zhan. 2023 · 2023
Closest in time.
Towards Efficient Visual Adaption via Structural Re-parameterization
Gen Luo, Minglang Huang, Yiyi Zhou, Xiaoshuai Sun, Guannan Jiang, Zhiyu Wang, and Rongrong Ji. 2023a · 2023
Closest in time.
A survey on deep hashing methods
Xiao Luo, Haixin Wang, Daqing Wu, Chong Chen, Minghua Deng, Jianqiang Huang, and Xian-Sheng Hua. 2023b · 2023
Closest in time.
Understanding and Mitigating Overfitting in Prompt Tuning for Vision-Language Models
Chengcheng Ma, Yang Liu, Jiankang Deng, Lingxi Xie, Weiming Dong, and Changsheng Xu. 2023 · 2023
Closest in time.
Tiny Adapters for Vision Transformers
Imad Eddine MAROUF, Enzo Tartaglione, and Stéphane Lathuilière. 2023 · 2023
Closest in time.
Latent-nerf for shape-guided generation of 3d shapes and textures. In Proc. CVPR . 12663–12673
Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, and Daniel Cohen-Or. 2023 · 2023
Closest in time.
Role of Bias Terms in Dot-Product Attention
Mahdi Namazifar, Devamanyu Hazarika, and Dilek Hakkani-Tur. 2023 · 2023
Closest in time.
ClimaX: A foundation model for weather and climate
Tung Nguyen, Johannes Brandstetter, Ashish Kapoor, Jayesh K Gupta, and Aditya Grover. 2023 · 2023
Closest in time.
Scalable diffusion models with transformers. In Proc. CVPR . 4195–4205
William Peebles and Saining Xie. 2023 · 2023
Closest in time.
Tool learning with foundation models
Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Yufei Huang, Chaojun Xiao, Chi Han, et al · 2023
Closest in time.
Low-Rank Winograd Transformation for 3D Convolutional Neural Networks
Ziran Qin, Mingbao Lin, and Weiyao Lin. 2023b · 2023
Closest in time.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proc. CVPR . 22500–22510
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023 · 2023
Closest in time.
Lossless Adaptation of Pretrained Vision Models For Robotic Manipulation. In Proc. ICLR
Mohit Sharma, Claudio Fantacci, Yuxiang Zhou, Skanda Koppula, Nicolas Heess, Jon Scholz, and Yusuf Aytar. 2023 · 2023
Closest in time.
Open-transmind: A new baseline and benchmark for 1st foundation model challenge of intelligent transportation. In Proc. CVPR Workshop . 6327–6334
Yifeng Shi, Feng Lv, Xinliang Wang, Chunlong Xia, Shaojie Li, Shujie Yang, Teng Xi, and Gang Zhang. 2023 · 2023
Closest in time.
FiT: Parameter Efficient Few-shot Transfer Learning for Personalized and Federated Image Classification. In Proc. ICLR
Aliaksandra Shysheya, John F Bronskill, Massimiliano Patacchiola, Sebastian Nowozin, and Richard E Turner. 2023 · 2023
Closest in time.
GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis
Ming Tao, Bing-Kun Bao, Hao Tang, and Changsheng Xu. 2023 · 2023
Closest in time.
Universal Deep Image Compression via Content-Adaptive Optimization with Adapters. In Proc. WACV . 2529–2538
Koki Tsubota, Hiroaki Akutsu, and Kiyoharu Aizawa. 2023 · 2023
Closest in time.
Chatgpt for robotics: Design principles and model abilities
Sai Vemprala, Rogerio Bonatti, Arthur Bucker, and Ashish Kapoor. 2023 · 2023
Closest in time.
LION: Implicit Vision Prompt Tuning
Haixin Wang, Jianlong Chang, Xiao Luo, Jinan Sun, Zhouchen Lin, and Qi Tian. 2023b · 2023
Closest in time.
Mode Approximation Makes Good Vision-Language Prompts
Haixin Wang, Xinlong Yang, Jianlong Chang, Dian Jin, Jinan Sun, Shikun Zhang, Xiao Luo, and Qi Tian. 2023c · 2023
Closest in time.
Open-Set Fine-Grained Retrieval via Prompting Vision-Language Evaluator. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 . IEEE, 19381–19391
Shijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang, Wanli Ouyang, and Qi Tian. 2023a · 2023
Closest in time.
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. 2023b · 2023
Closest in time.
Exploring Efficient Few-shot Adaptation for Vision Transformers
Chengming Xu, Siqian Yang, Yabiao Wang, Zhanxiong Wang, Yanwei Fu, and Xiangyang Xue. 2023a · 2023
Closest in time.
Side Adapter Network for Open-Vocabulary Semantic Segmentation
Mengde Xu, Zheng Zhang, Fangyun Wei, Han Hu, and Xiang Bai. 2023b · 2023
Closest in time.
Diffusion models: A comprehensive survey of methods and applications
Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2023a · 2023
Closest in time.
Prompting and Tuning: A Two-Stage Unsupervised Domain Adaptive Person Re-identification Method on Vision Transformer Backbone
Shengming Yu, Zhaopeng Dou, and Shengjin Wang. 2023b · 2023
Closest in time.
Rethinking Vision Transformer and Masked Autoencoder in Multimodal Face Anti-Spoofing
Zitong Yu, Rizhao Cai, Yawen Cui, Xin Liu, Yongjian Hu, and Alex Kot. 2023a · 2023
Closest in time.
Multimodal Video Adapter for Parameter Efficient Video Text Retrieval
Bowen Zhang, Xiaojie Jin, Weibo Gong, Kai Xu, Zhao Zhang, Peng Wang, Xiaohui Shen, and Jiashi Feng. 2023a · 2023
Closest in time.
What Makes Good Examples for Visual In-Context Learning?
Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. 2023b · 2023
Closest in time.
A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT
Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, et al · 2023
Closest in time.
EGCN++: A New Fusion Strategy for Ensemble Learning in Skeleton-Based Rehabilitation Exercise Assessment
XB Bruce, Yan Liu, Keith CC Chan, and Chang Wen Chen. 2024 · 2024
Closest in time.
Agent ai: Surveying the horizons of multimodal interaction
Zane Durante, Qiuyuan Huang, Naoki Wake, Ran Gong, Jae Sung Park, Bidipta Sarkar, Rohan Taori, Yusuke Noda, Demetri Terzopoulos, Yejin Choi, et al · 2024
Closest in time.
Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, et al · 2024
Closest in time.