Parameter-efficient transfer learning for NLP
Original
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 1902
Earlier work this paper cites.
On the mathematical foundations of theoretical statistics
Ronald A Fisher · 1922
Earlier work this paper cites.
Neural learning in structured parameter spaces-natural riemannian gradient
SI Amari · 1997
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Original
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton · 2002
Earlier work this paper cites.
Masking as an efficient alternative to finetuning for pretrained language models
Original
Mengjie Zhao, Tao Lin, Martin Jaggi, and Hinrich Schütze · 2004
Earlier work this paper cites.
Language models are few-shot learners
Original
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Big self-supervised models are strong semi-supervised learners
Original
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E. Hinton · 2006
Earlier work this paper cites.
Adapterhub: A framework for adapting transformers
Original
Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulic, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2009
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, Rajat Monga, Kai Chen, M. Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, A. Senior, P. Tucker, K. Yang, and A. Ng · 2012
Earlier work this paper cites.
Parameter-efficient transfer learning with diff pruning
Original
Demi Guo, Alexander M. Rush, and Yoon Kim · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Original
Razvan Pascanu and Yoshua Bengio · 2013
Earlier work this paper cites.