Fetching the paper…
Reading the bibliography…
While task-specific finetuning of pretrained networks has led to significant empirical advances in NLP, the large size of networks makes finetuning difficult to deploy in multi-task, memory-constrained settings.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Structured Pruning of Large Language Models
Ziheng Wang, Jeremy Wohlwend, and Tao Lei · 1910
Earlier work this paper cites.
Multitask Learning
Rich Caruana · 1997
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert French · 1999
Earlier work this paper cites.
AdapterFusion: Non-Destructive Task Composition for Transfer Learning
Jonas Pfeiffer, Aishwarya Kamath, Andreas Ruckle, and Kyunghyun Cho amd Iryna Gurevych · 2005
Earlier work this paper cites.
MAD-X: An Adapter-based Framework for Multi-task Cross-lingual Transfer
Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder · 2005
Earlier work this paper cites.
AdapterHub: A Framework for Adapting Transformers
Jonas Pfeiffer, Andreas Ruckle, Clifton Poth, Aishwarya Kamath, Ivan Vulic, Sebastian Ruder, and Iryna Gurevych Kyunghyun Cho · 2007
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Distributed Representations of Sentences and Documents
Quoc V. Le and Tomas Mikolov · 2014
Earlier work this paper cites.
Semi-Supervised Sequence Learning
Andrew Dai and Quoc V. Le · 2015
Earlier work this paper cites.
Skip-Thought Vectors
Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2015
Earlier work this paper cites.
Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Song Han, Huizi Mao, and William J. Dally · 2016
Earlier work this paper cites.
Learning distributed representations of sentences from unlabelled data
Felix Hill, Kyunghyun Cho, and Anna Korhonen · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Andrei A. Rusu, Neil C. Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Earlier work this paper cites.
Towards Universal Paraphrastic Sentence Embeddings
John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu · 2016
Earlier work this paper cites.
A Simple but Tough-to-Beat Baseline for Sentence Embeddings
Sanjeev Arora, Yingyu Liang, and Tengyu Ma · 2017
Earlier work this paper cites.
Supervised Learning of Universal Sentence Representations from Natural Language Inference Data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loic Barrault, and Antoine Bordes · 2017
Earlier work this paper cites.
Categorical Reparameterization with Gumbel-Softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting by incremental moment matching
Sang-Woo Lee, Jin-Hwa Kim, Jaehyun Jun, Jung-Woo Ha, and Byoung-Tak Zhang · 2017
Earlier work this paper cites.
Gradient Episodic Memory for Continual Learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Earlier work this paper cites.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh · 2017
Earlier work this paper cites.
Learned in translation: Contextualized word vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher · 2017
Cited alongside, same era.
Regularization techniques for fine-tuning in neural machine translation
Antonio Valerio Miceli Barone, Barry Haddow, Ulrich Germann, and Rico Sennrich · 2017
Cited alongside, same era.
Continual Learning with Deep Generative Replay
Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim · 2017
Cited alongside, same era.
Neural domain adaptation for biomedical question answering
Georg Wiese, Dirk Weissenborn, and Mariana Neves · 2017
Cited alongside, same era.
Universal sentence encoder for English
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil · 2018
Cited alongside, same era.
Universal Language Model Fine-tuning for Text Classification
Jeremy Howard and Sebastian Ruder · 2018
BERT and PALs: Projected attention layers for efficient adaptation in multi-task learning
Asa Cooper Stickland and Iain Murray · 2019
Later among the works it cites.
Patient Knowledge Distillation for BERT Model Compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 2019
Later among the works it cites.
BERT Rediscovers the Classical NLP Pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Later among the works it cites.
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning Sparse Neural Networks through L 0 L_{0} Regularization
Christos Louizos, Max Welling, Diederik P, and Kingma · 2018
Cited alongside, same era.
Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights
Arun Mallya, Dillon Davis, and Svetlana Lazebnik · 2018
Cited alongside, same era.
Deep Contextualized Word Representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
Efficient Parametrization of Multi-domain Deep Neural Networks
S. Rebuffi, A. Vedaldi, and H. Bilen · 2018
Cited alongside, same era.
Progress & Compress: A scalable framework for continual learning
Jonathan Schwarz, Jelena Luketina, Wojciech M. Czarnecki, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell · 2018
Cited alongside, same era.
Later among the works it cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Later among the works it cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Closest in time.
The Lottery Ticket Hypothesis for Pre-trained BERT Networks
Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin · 2020
Closest in time.
Compressing Large-Scale Transformer-Based Models: A Case Study on BERT
Prakhar Ganesh, Yao Chen, Xin Lou, Mohammad Ali Khan, Yin Yang, Deming Chen, Marianne Winslett, Hassan Sajjad, and Preslav Nakov · 2020
Closest in time.
Compressing BERT: Studying the Effects of Weight Pruning on Transfer Learning
Mitchell A. Gordon, Kevin Duh, and Nicholas Andrews · 2020
Closest in time.
TinyBERT: Distilling BERT for Natural Language Understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2020
Closest in time.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Closest in time.
Mixout: Effective Regularization to Finetune Large-scale Pretrained Language Models
Cheolhyoung Lee, Kyunghyun Cho, and Wanmo Kang · 2020
Closest in time.
How fine can fine-tuning be? Learning efficient language models
Evani Radiya-Dixit and Xin Wang · 2020
Closest in time.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Katherine Lee Adam Roberts, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Closest in time.
Poor Man’s BERT: Smaller and Faster Transformer Models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov · 2020
Closest in time.
Movement Pruning: Adaptive Sparsity by Fine-Tuning
Victor Sanh, Thomas Wolf, and Alexander M. Rush · 2020
Closest in time.
It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
Timo Schick and Hinrich Schutze · 2020
Closest in time.
Similarity Analysis of Contextual Word Representation Models
John M. Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass · 2020
Closest in time.
Masking as an Efficient Alternative to Finetuning for Pretrained Language Models
Mengjie Zhao, Tao Lin, Martin Jaggi, and Hinrich Schutze · 2020
Closest in time.
Low-Complexity Probing via Finding Subnetworks
Steven Cao, Victor Sanh, and Alexander M. Rush · 2021
Closest in time.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Xiang Lisa Li and Percy Liang · 2021
Closest in time.
Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
Guanghui Qin and Jason Eisner · 2021
Closest in time.