Mixout: Effective regularization to finetune large-scale pretrained language models
Original
Cheolhyoung Lee, Kyunghyun Cho, and Wanmo Kang · 1909
Earlier work this paper cites.
What would elsa do? freezing layers during transformer fine-tuning
Original
Jaejun Lee, Raphael Tang, and Jimmy Lin · 1911
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
On causal and anticausal learning
Bernhard Schölkopf, Dominik Janzing, Jonas Peters, Eleni Sgouritsa, Kun Zhang, and Joris Mooij · 2012
Earlier work this paper cites.
Learning and transferring mid-level image representations using convolutional neural networks
Maxime Oquab, Leon Bottou, Ivan Laptev, and Josef Sivic · 2014
Earlier work this paper cites.
Cnn features off-the-shelf: an astounding baseline for recognition
Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, and Stefan Carlsson · 2014
Earlier work this paper cites.
Deep domain confusion: Maximizing for domain invariance
Original
Eric Tzeng, Judy Hoffman, Ning Zhang, Kate Saenko, and Trevor Darrell · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Original
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Unsupervised domain adaptation with residual transfer networks
Mingsheng Long, Han Zhu, Jianmin Wang, and Michael I Jordan · 2016
Earlier work this paper cites.
Causal inference by using invariant prediction: identification and confidence intervals
Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen · 2016
Earlier work this paper cites.
Learning transferrable representations for unsupervised domain adaptation
Ozan Sener, Hyun Oh Song, Ashutosh Saxena, and Silvio Savarese · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Early stopping without a validation set
Original
Maren Mahsereci, Lukas Balles, Christoph Lassner, and Philipp Hennig · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2017
Earlier work this paper cites.
From detection of individual metastases to classification of lymph node status at the patient level: the camelyon17 challenge
Peter Bandi, Oscar Geessink, Quirine Manson, Marcory Van Dijk, Maschenka Balkenhol, Meyke Hermsen, Babak Ehteshami Bejnordi, Byungjae Lee, Kyunghyun Paeng, Aoxiao Zhong, et al · 2018
Earlier work this paper cites.
Functional map of the world
Gordon Christie, Neil Fendley, James Wilson, and Ryan Mukherjee · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Original
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Explicit inductive bias for transfer learning with convolutional networks
LI Xuhong, Yves Grandvalet, and Franck Davoine · 2018
Earlier work this paper cites.
Invariant risk minimization
Original
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Earlier work this paper cites.
What is the effect of importance weighting in deep learning?
Jonathon Byrd and Zachary Lipton · 2019
Earlier work this paper cites.
The intriguing role of module criticality in the generalization of deep networks
Original
Niladri S Chatterji, Behnam Neyshabur, and Hanie Sedghi · 2019
Earlier work this paper cites.
Spottune: transfer learning through adaptive fine-tuning
Yunhui Guo, Honghui Shi, Abhishek Kumar, Kristen Grauman, Tajana Rosing, and Rogerio Feris · 2019
Earlier work this paper cites.