Fetching the paper…
Reading the bibliography…
When transferring a pretrained model to a downstream task, two popular methods are full fine-tuning (updating all the model parameters) and linear probing (updating only the last linear layer -- the "head").
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song · 1907
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Smallest singular value of a random rectangular matrix
Mark Rudelson and Roman Vershynin · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Adam Coates, Andrew Ng, and Honlak Lee · 2011
Earlier work this paper cites.
Matrix Computations
Gene H. Golub and Charles F. Van Loan · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Joel A. Tropp · 2015
Earlier work this paper cites.
Combining satellite imagery and machine learning to predict poverty
Neal Jean, Marshall Burke, Michael Xie, W. Matthew Davis, David B. Lobell, and Stefano Ermon · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, Chelsea Finn, Trevor Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass · 2017
Earlier work this paper cites.
Borrowing treasures from the wealthy: Deep transfer learning through selective joint fine-tuning
Weifeng Ge and Yizhou Yu · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Deep learning for segmentation of brain tumors: Impact of cross-institutional training and testing
EA AlBadawy, A Saha, and MA Mazurowski · 2018
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Earlier work this paper cites.
Functional map of the world
Gordon Christie, Neil Fendley, James Wilson, and Ryan Mukherjee · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon Shaolei Du, Wei Hu, and Jason Lee · 2018
Earlier work this paper cites.
Self-ensembling for visual domain adaptation
Geoff French, Michal Mackiewicz, and Mark Fisher · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Deep linear neural networks with arbitrary loss: All local minima are global
Thomas Laurent and James H. von Brecht · 2018
Earlier work this paper cites.
Explicit inductive bias for transfer learning with convolutional networks
Xuhong Li, Yves Grandvalet, and Franck Davoine · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Do CIFAR-10 classifiers generalize to CIFAR-10?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2018
Cited alongside, same era.
Benign overfitting in linear regression
Peter L. Bartlett, Philip M. Long, G´abor Lugosi, and Alexander Tsigler · 2019
Cited alongside, same era.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2019
Cited alongside, same era.
A new look at an old problem: A universal learning approach to linear regression
Koby Bibas, Yaniv Fogel, and Meir Feder · 2019
Cited alongside, same era.
Implicit regularization of discrete gradient dynamics in deep linear neural networks
Gauthier Gidel, Francis R. Bach, and Simon Lacoste-Julien · 2019
Cited alongside, same era.
Spottune: Transfer learning through adaptive fine-tuning
Yunhui Guo, Honghui Shi, Abhishek Kumar, Kristen Grauman, Tajana Rosing, and Rogerio Feris · 2019
On the theory of transfer learning: The importance of task diversity
Nilesh Tripuraneni, Michael I. Jordan, and Chi Jin · 2020
Later among the works it cites.
Understanding and improving information transfer in multi-task learning
Sen Wu, Hongyang R. Zhang, and Christopher Ré · 2020
Later among the works it cites.
Bdd100k: A diverse driving dataset for heterogeneous multitask learning
Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell · 2020
Later among the works it cites.
A large-scale study of representation learning with the visual task adaptation benchmark
Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, Andre Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, Lucas Beyer, Olivier Bachem, Michael Tschannen, Marcin Michalski, Olivier Bousquet, Sylvain Gelly, and Neil Houlsby · 2020
Later among the works it cites.
Side-tuning: A baseline for network adaptation via additive side networks
Jeffrey O Zhang, Alexander Sax, Amir Zamir, Leonidas Guibas, and Jitendra Malik · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Cited alongside, same era.
Do better imagenet models transfer better?
Simon Kornblith, Jonathon Shlens, and Quoc V. Le · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Cited alongside, same era.
Moment matching for multi-source domain adaptation
Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang · 2019
Cited alongside, same era.
FreeLB: Enhanced adversarial training for natural language understanding
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu · 2020
Later among the works it cites.
Better fine-tuning by reducing representational collapse
Armen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal, Luke Zettlemoyer, and Sonal Gupta · 2021
Later among the works it cites.
The evolution of out-of-distribution robustness throughout fine-tuning
Anders Andreassen, Yasaman Bahri, Behnam Neyshabur, and Rebecca Roelofs · 2021
Later among the works it cites.
A theory of label propagation for subpopulation shift
Tianle Cai, Ruiqi Gao, J. Lee, and Qi Lei · 2021
Later among the works it cites.
How fine-tuning allows for effective meta-learning
Kurtland Chua, Qi Lei, and Jason D Lee · 2021
Later among the works it cites.
Smart: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization
Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao · 2021
Later among the works it cites.
Does invariant risk minimization capture invariance?
Pritish Kamath, Akilesh Tangella, Danica J. Sutherland, and Nathan Srebro · 2021
Later among the works it cites.
Partial transfusion: on the expressive influence of trainable batch norm parameters for transfer learning
Fahdi Kanavati and Masayuki Tsuneki · 2021
Later among the works it cites.
Removing spurious features can hurt accuracy and affect groups disproportionately
Fereshte Khani and Percy Liang · 2021
Later among the works it cites.
WILDS: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton A. Earnshaw, Imran S. Haque, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Later among the works it cites.
Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization
John Miller, Rohan Taori, Aditi Raghunathan, Shiori Sagawa, Pang Wei Koh, Vaishaal Shankar, Percy Liang, Yair Carmon, and Ludwig Schmidt · 2021
Later among the works it cites.
Selective entropy optimization via committee consistency for unsupervised domain adaptation
Viraj Prabhu, Shivam Khare, Deeksha Karthik, and Judy Hoffman · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
The risks of invariant risk minimization
Elan Rosenfeld, Pradeep Ravikumar, and Andrej Risteski · 2021
Later among the works it cites.
Avoiding inference heuristics in few-shot prompt-based finetuning
Prasetya Ajie Utama, Nafise Sadat Moosavi, Victor Sanh, and Iryna Gurevych · 2021
Later among the works it cites.
Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Mike Li, Jong Wook Kim, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt · 2021
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2021
Later among the works it cites.