Fetching the paper…
Reading the bibliography…
In computer vision, it has achieved great transfer learning performance via adapting large-scale pretrained vision models (e.g., vision transformers) to downstream tasks.
Parameter-Efficient Transfer Learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; de Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 1902
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Knowledge guided text retrieval and reading for open domain question answering
Min, S.; Chen, D.; Zettlemoyer, L.; and Hajishirzi, H. 2019 · 1911
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M.; and Cohen, N. J. 1989 · 1989
Earlier work this paper cites.
Batchensemble: an alternative approach to efficient ensemble and lifelong learning
Wen, Y.; Tran, D.; and Ba, J. 2020 · 2002
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental Bayesian approach tested on 101 object categories
Fei-Fei, L.; Fergus, R.; and Perona, P. 2004 · 2004
Earlier work this paper cites.
Zhang, A.; Tay, Y.; Zhang, S.; Chan, A.; Luu, A. T.; Hui, S. C.; and Fu, J. 2021 · 2004
Earlier work this paper cites.
AdapterFusion: Non-Destructive Task Composition for Transfer Learning
Pfeiffer, J.; Kamath, A.; Rücklé, A.; Cho, K.; and Gurevych, I. 2021 · 2005
Earlier work this paper cites.
Funnel-transformer: Filtering out sequential redundancy for efficient language processing
Dai, Z.; Lai, G.; Yang, Y.; and Le, Q. V. 2020 · 2006
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E.; and Zisserman, A. 2008 · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A.; Hinton, G.; et al. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
The pascal visual object classes (VOC) challenge
Everingham, M.; Van Gool, L.; Williams, C. K.; Winn, J.; and Zisserman, A. 2010 · 2010
Earlier work this paper cites.
AdapterDrop: On the Efficiency of Adapters in Transformers
Rücklé, A.; Geigle, G.; Glockner, M.; Beck, T.; Pfeiffer, J.; Reimers, N.; and Gurevych, I. 2021 · 2010
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Xiao, J.; Hays, J.; Ehinger, K. A.; Oliva, A.; and Torralba, A. 2010 · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Coates, A.; Ng, A.; and Lee, H. 2011 · 2011
Earlier work this paper cites.
The German traffic sign recognition benchmark: a multi-class classification competition
Stallkamp, J.; Schlipsing, M.; Salmen, J.; and Igel, C. 2011 · 2011
Earlier work this paper cites.
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning
Aghajanyan, A.; Zettlemoyer, L.; and Gupta, S. 2020 · 2012
Earlier work this paper cites.
The MNIST database of handwritten digit images for machine learning research
Deng, L. 2012 · 2012
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M.; Vedaldi, A.; Zisserman, A.; and Jawahar, C. 2012 · 2012
Earlier work this paper cites.
A new performance measure and evaluation benchmark for road detection algorithms
Fritsch, J.; Kuehnl, T.; and Geiger, A. 2013 · 2013
Cited alongside, same era.
3d object representations for fine-grained categorization
Krause, J.; Stark, M.; Deng, J.; and Fei-Fei, L. 2013 · 2013
Cited alongside, same era.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
Sainath, T. N.; Kingsbury, B.; Sindhwani, V.; Arisoy, E.; and Ramabhadran, B. 2013 · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D.; Cho, K.; and Bengio, Y. 2014 · 2014
Cited alongside, same era.
Food-101–mining discriminative components with random forests
Bossard, L.; Guillaumin, M.; and Gool, L. V. 2014 · 2014
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018 · 2018
Later among the works it cites.
Adaptive Input Representations for Neural Language Modeling
Baevski, A.; and Auli, M. 2019 · 2019
Later among the works it cites.
EuroSat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P.; Bischke, B.; Dengel, A.; and Borth, D. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Later among the works it cites.
Sentiment Classification Using Document Embeddings Trained with Cosine Similarity
Thongtan, T.; and Phienthrakul, T. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kingma, D.; and Ba, J. 2014 · 2014
Cited alongside, same era.
Fastfood: Approximate kernel expansions in loglinear time
Le, Q. V.; Sarlós, T.; and Smola, A. J. 2014 · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2014 · 2014
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Ruder, S. 2016 · 2016
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J.; Salakhutdinov, R. R.; and Le, Q. V. 2019 · 2019
Later among the works it cites.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Later among the works it cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Kiela, D.; Firooz, H.; Mohan, A.; Goswami, V.; Singh, A.; Ringshia, P.; and Testuggine, D. 2020 · 2020
Later among the works it cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X.; Yin, X.; Li, C.; Zhang, P.; Hu, X.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; et al. 2020 · 2020
Later among the works it cites.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
Dong, X.; Bao, J.; Chen, D.; Zhang, W.; Yu, N.; Yuan, L.; Chen, D.; and Guo, B. 2021 · 2021
Later among the works it cites.
Kronecker decomposition for gpt compression
Edalati, A.; Tahaei, M.; Rashid, A.; Nia, V. P.; Clark, J. J.; and Rezagholizadeh, M. 2021 · 2021
Later among the works it cites.
Towards a unified view of parameter-efficient transfer learning
He, J.; Zhou, C.; Ma, X.; Berg-Kirkpatrick, T.; and Neubig, G. 2021 · 2021
Later among the works it cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Later among the works it cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Later among the works it cites.
Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Later among the works it cites.
Compacter: Efficient Low-Rank Hypercomplex Adapter Layers
Mahabadi, R. K.; Henderson, J.; and Ruder, S. 2021 · 2021
Later among the works it cites.
Carbon emissions and large neural network training
Patterson, D.; Gonzalez, J.; Le, Q.; Liang, C.; Munguia, L.-M.; Rothchild, D.; So, D.; Texier, M.; and Dean, J. 2021 · 2021
Later among the works it cites.
Tahaei, M. S.; Charlaix, E.; Nia, V. P.; Ghodsi, A.; and Rezagholizadeh, M. 2021 · 2021
Later among the works it cites.
BitFit: Simple Parameter-Efficient Fine-Tuning for Transformer-Based Masked Language-Models
Zaken, E. B.; Ravfogel, S.; and Goldberg, Y. 2021 · 2021
Later among the works it cites.
Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks
Sung, Y.-L.; Cho, J.; and Bansal, M. 2022 · 2022
Closest in time.