Fetching the paper…
Reading the bibliography…
Despite recent competitive performance across a range of vision tasks, vision Transformers still have an issue of heavy computational costs.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Refinenet: Multi-path refinement networks for high-resolution semantic segmentation
Lin, G.; Milan, A.; Shen, C.; and Reid, I. 2017 · 1934
Earlier work this paper cites.
Exploiting cloze questions for few shot text classification and natural language inference
Schick, T.; and Schütze, H. 2020 · 2001
Earlier work this paper cites.
AdapterFusion: Non-destructive task composition for transfer learning
Pfeiffer, J.; Kamath, A.; Rücklé, A.; Cho, K.; and Gurevych, I. 2020a · 2005
Earlier work this paper cites.
Adapterhub: A framework for adapting transformers
Pfeiffer, J.; Rücklé, A.; Poth, C.; Kamath, A.; Vulić, I.; Ruder, S.; Cho, K.; and Gurevych, I. 2020b · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E.; and Zisserman, A. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A.; Hinton, G.; et al. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Novel dataset for fine-grained image categorization: Stanford dogs
Khosla, A.; Jayadevaprakash, N.; Yao, B.; and Li, F.-F. 2011 · 2011
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Gao, T.; Fisch, A.; and Chen, D. 2020 · 2012
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M.; Vedaldi, A.; Zisserman, A.; and Jawahar, C. 2012 · 2012
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Bossard, L.; Guillaumin, M.; and Van Gool, L. 2014 · 2014
Earlier work this paper cites.
Recurrent models of visual attention
Mnih, V.; Heess, N.; Graves, A.; et al. 2014 · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J.; Clune, J.; Bengio, Y.; and Lipson, H. 2014 · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Optnet: Differentiable optimization as a layer in neural networks
Amos, B.; and Kolter, J. Z. 2017 · 2017
Cited alongside, same era.
Fine-grained car detection for visual census estimation
Gebru, T.; Krause, J.; Wang, Y.; Chen, D.; Deng, J.; and Fei-Fei, L. 2017 · 2017
Cited alongside, same era.
Sparsemap: Differentiable sparse structured inference
Niculae, V.; Martins, A.; Blondel, M.; and Cardie, C. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Cited alongside, same era.
Joint optimization framework for learning with noisy labels
Tanaka, D.; Ikami, D.; Yamasaki, T.; and Aizawa, K. 2018 · 2018
Cited alongside, same era.
Deep equilibrium models
Bai, S.; Kolter, J. Z.; and Koltun, V. 2019 · 2019
Cited alongside, same era.
Robust early-learning: Hindering the memorization of noisy labels
Xia, X.; Liu, T.; Han, B.; Gong, C.; Wang, N.; Ge, Z.; and Chang, Y. 2021 · 2021
Later among the works it cites.
Optimization induced equilibrium networks
Xie, X.; Wang, Q.; Ling, Z.; Li, X.; Wang, Y.; Liu, G.; and Lin, Z. 2021 · 2021
Later among the works it cites.
Pruning by explaining: A novel criterion for deep neural network pruning
Yeom, S.-K.; Seegerer, P.; Lapuschkin, S.; Binder, A.; Wiedemann, S.; Müller, K.-R.; and Samek, W. 2021 · 2021
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B.; Ravfogel, S.; and Goldberg, Y. 2021 · 2021
Later among the works it cites.
Adversarial Robustness against Multiple and Single l _ p l\_p -Threat Models via Quick Fine-Tuning of Robust Classifiers
Croce, F.; and Hein, M. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J.; and Carbin, M. 2019 · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Cited alongside, same era.
Importance estimation for neural network pruning
Molchanov, P.; Mallya, A.; Tyree, S.; Frosio, I.; and Kautz, J. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Cited alongside, same era.
Satnet: Bridging deep learning and logical reasoning using a differentiable satisfiability solver
Wang, P.-W.; Donti, P.; Wilder, B.; and Kolter, Z. 2019 · 2019
Cited alongside, same era.
Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning
Lin, Z.; Madotto, A.; and Fung, P. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Survey on Efficient Training of Large Neural Networks
Gusak, J.; Cherniuk, D.; Shilova, A.; Katrutsa, A.; Bershatsky, D.; Zhao, X.; Eyraud-Dubois, L.; Shliazhko, O.; Dimitrov, D.; Oseledets, I.; and Beaumont, O. 2022 · 2022
Later among the works it cites.
Jia, M.; Tang, L.; Chen, B.-C.; Cardie, C.; Belongie, S.; Hariharan, B.; and Lim, S.-N. 2022 · 2022
Later among the works it cites.
Video swin transformer
Liu, Z.; Ning, J.; Cao, Y.; Wei, Y.; Zhang, Z.; Lin, S.; and Hu, H. 2022 · 2022
Later among the works it cites.
Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks
Sung, Y.-L.; Cho, J.; and Bansal, M. 2022 · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M.; Ilharco, G.; Gadre, S. Y.; Roelofs, R.; Gontijo-Lopes, R.; Morcos, A. S.; Namkoong, H.; Farhadi, A.; Carmon, Y.; Kornblith, S.; et al. 2022 · 2022
Later among the works it cites.
Diffusion models: A comprehensive survey of methods and applications
Yang, L.; Zhang, Z.; Song, Y.; Hong, S.; Xu, R.; Zhao, Y.; Zhang, W.; Cui, B.; and Yang, M.-H. 2022 · 2022
Later among the works it cites.
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
Zaken, E. B.; Goldberg, Y.; and Ravfogel, S. 2022 · 2022
Later among the works it cites.
Parameter-efficient Tuning of Large-scale Multimodal Foundation Model
Wang, H.; Yang, X.; Chang, J.; Jin, D.; Sun, J.; Zhang, S.; Luo, X.; and Tian, Q. 2023 · 2023
Closest in time.
Improving Diffusion-Based Image Synthesis with Context Prediction
Yang, L.; Liu, J.; Hong, S.; Zhang, Z.; Huang, Z.; Cai, Z.; Zhang, W.; and CUI, B. 2023 · 2023
Closest in time.
Yu, B. X.; Chang, J.; Wang, H.; Liu, L.; Wang, S.; Wang, Z.; Lin, J.; Xie, L.; Li, H.; Lin, Z.; et al. 2023 · 2023
Closest in time.