Fetching the paper…
Reading the bibliography…
Deep pretrained language models have achieved great success in the way of pretraining first and then fine-tuning.
Online continual learning with no task boundaries
Rahaf Aljundi, Min Lin, Baptiste Goujaud, and Yoshua Bengio. 2019 · 1903
Earlier work this paper cites.
Class-incremental learning via deep model consolidation
Junting Zhang, Jie Zhang, Shalini Ghosh, Dawei Li, Serafettin Tasci, Larry P. Heck, Heming Zhang, and C.-C. Jay Kuo. 2019b · 1903
Earlier work this paper cites.
Jiabin Xue, Jiqing Han, Tieran Zheng, Xiang Gao, and Jiaxing Guo. 2019 · 1904
Earlier work this paper cites.
Few-shot sequence labeling with label dependency transfer
Yutai Hou, Zhihan Zhou, Yijia Liu, Ning Wang, Wanxiang Che, Han Liu, and Ting Liu. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 1909
Earlier work this paper cites.
Continual learning: A comparative study on how to defy forgetting in classification tasks
Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Ales Leonardis, Gregory G. Slabaugh, and Tinne Tuytelaars. 2019 · 1909
Earlier work this paper cites.
Mixout: Effective regularization to finetune large-scale pretrained language models
Cheolhyoung Lee, Kyunghyun Cho, and Wanmo Kang. 2019 · 1909
Earlier work this paper cites.
Multi-task learning and catastrophic forgetting in continual reinforcement learning
Joao Ribeiro, Francisco S Melo, and Joao Dias. 2019 · 1909
Earlier work this paper cites.
Forget me not: Reducing catastrophic forgetting for domain adaptation in reading comprehension
Ying Xu, Xu Zhong, Antonio Jose Jimeno Yepes, and Jey Han Lau. 2019 · 1911
Earlier work this paper cites.
Side-tuning: Network adaptation via additive side networks
Jeffrey O. Zhang, Alexander Sax, Amir Roshan Zamir, Leonidas J. Guibas, and Jitendra Malik. 2019a · 1912
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen. 1989 · 1989
Earlier work this paper cites.
Multitask learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M. French. 1999 · 1999
Earlier work this paper cites.
Information theory, inference, and learning algorithms
David J. C. MacKay. 2003 · 2003
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Bill Dolan Lisa Ferro Danilo Giampiccolo Bernardo Magnini Roy Bar-Haim, Ido Dagan and Idan Szpektor. 2006 · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007 · 2007
Earlier work this paper cites.
The fifth PASCAL recognizing textual entailment challenge
Luisa Bentivogli, Bernardo Magnini, Ido Dagan, Hoa Trang Dang, and Danilo Giampiccolo. 2009 · 2009
Earlier work this paper cites.
The winograd schema challenge
Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
New perspectives on the natural gradient method
James Martens. 2014 · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Generating sentences from a continuous space
Samuel R. Bowman, Luke Vilnis, Oriol Vinyals, Andrew M. Dai, Rafal Józefowicz, and Samy Bengio. 2016 · 2016
Cited alongside, same era.
Less-forgetting learning in deep neural networks
Heechul Jung, Jeongwoo Ju, Minju Jung, and Junmo Kim. 2016 · 2016
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Later among the works it cites.
Explicit inductive bias for transfer learning with convolutional networks
Xuhong Li, Yves Grandvalet, and Franck Davoine. 2018 · 2018
Later among the works it cites.
Learning without forgetting
Zhizhong Li and Derek Hoiem. 2018 · 2018
Later among the works it cites.
Rotate your networks: Better weight consolidation and less catastrophic forgetting
Xialei Liu, Marc Masana, Luis Herranz, Joost van de Weijer, Antonio M. López, and Andrew D. Bagdanov. 2018 · 2018
Later among the works it cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Arun Mallya and Svetlana Lazebnik. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
How transferable are neural networks in NLP applications?
Lili Mou, Zhao Meng, Rui Yan, Ge Li, Yan Xu, Lu Zhang, and Zhi Jin. 2016 · 2016
Cited alongside, same era.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Andrei A. Rusu, Neil C. Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. 2016 · 2016
Cited alongside, same era.
Regularization techniques for fine-tuning in neural machine translation
Antonio Valerio Miceli Barone, Barry Haddow, Ulrich Germann, and Rico Sennrich. 2017 · 2017
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel M. Cer, Mona T. Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm
Bjarke Felbo, Alan Mislove, Anders Søgaard, Iyad Rahwan, and Sune Lehmann. 2017 · 2017
Cited alongside, same era.
On quadratic penalties in elastic weight consolidation
Ferenc Huszár. 2017 · 2017
Cited alongside, same era.
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
Progress & compress: A scalable framework for continual learning
Jonathan Schwarz, Wojciech Czarnecki, Jelena Luketina, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell. 2018 · 2018
Later among the works it cites.
Overcoming catastrophic forgetting with hard attention to the task
Joan Serrà, Didac Suris, Marius Miron, and Alexandros Karatzoglou. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
Reinforced continual learning
Ju Xu and Zhanxing Zhu. 2018 · 2018
Later among the works it cites.
Does an LSTM forget more than a cnn? an empirical study of catastrophic forgetting in NLP
Gaurav Arora, Afshin Rahimi, and Timothy Baldwin. 2019 · 2019
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Later among the works it cites.
To tune or not to tune? adapting pretrained representations to diverse tasks
Matthew E. Peters, Sebastian Ruder, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Neural transfer learning for natural language processing
Sebastian Ruder. 2019 · 2019
Later among the works it cites.
How to fine-tune bert for text classification?
Chi Sun, Xipeng Qiu, Yige Xu, and Xuanjing Huang. 2019 · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
Incremental learning through deep adaptation
Amir Rosenfeld and John K. Tsotsos. 2020 · 2020
Closest in time.
Overcoming catastrophic forgetting during domain adaptation of neural machine translation
Brian Thompson, Jeremy Gwinnup, Huda Khayrallah, Kevin Duh, and Philipp Koehn. 2019 · 2068
Closest in time.
An embarrassingly simple approach for transfer learning from pretrained language models
Alexandra Chronopoulou, Christos Baziotis, and Alexandros Potamianos. 2019 · 2095
Closest in time.