Fetching the paper…
Reading the bibliography…
Sequential fine-tuning and multi-task learning are methods aiming to incorporate knowledge from multiple tasks; however, they suffer from catastrophic forgetting and difficulties in dataset balancing.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen. 1989 · 1989
Earlier work this paper cites.
Multitask learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M French. 1999 · 1999
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
A unified architecture for natural language processing: deep neural networks with multitask learning
Ronan Collobert and Jason Weston. 2008 · 2008
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang. 2010 · 2010
Earlier work this paper cites.
Transfer learning
Lisa Torrey and Jude Shavlik. 2010 · 2010
Earlier work this paper cites.
The winograd schema challenge
Hector J. Levesque. 2011 · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014 · 2014
Earlier work this paper cites.
Large-scale multi-label text classification - revisiting neural networks
Jinseok Nam, Jungi Kim, Eneldo Loza Menc’ia, Iryna Gurevych, and Johannes Fürnkranz. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Recurrent neural network for text classification with multi-task learning
Pengfei Liu, Xipeng Qiu, and Xuanjing Huang. 2016 · 2016
Earlier work this paper cites.
Fully character-level neural machine translation without explicit segmentation
Jason Lee, Kyunghyun Cho, and Thomas Hofmann. 2017 · 2017
Earlier work this paper cites.
Adversarial multi-task learning for text classification
Pengfei Liu, Xipeng Qiu, and Xuanjing Huang. 2017 · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017 · 2017
Earlier work this paper cites.
An overview of multi-task learning in deep neural networks
Sebastian Ruder. 2017 · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
A survey on multi-task learning
Yu Zhang and Qiang Yang. 2017 · 2017
Cited alongside, same era.
Senteval: An evaluation toolkit for universal sentence representations
Alexis Conneau and Douwe Kiela. 2018 · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
Scitail: A textual entailment dataset from science question answering
Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018 · 2018
Cited alongside, same era.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Neural Transfer Learning for Natural Language Processing
Sebastian Ruder. 2019 · 2019
Later among the works it cites.
Latent multi-task architecture learning
Sebastian Ruder, Joachim Bingel, Isabelle Augenstein, and Anders Søgaard. 2019 · 2019
Later among the works it cites.
A hierarchical multi-task approach for learning embeddings from semantic tasks
Victor Sanh, Thomas Wolf, and Sebastian Ruder. 2019 · 2019
Later among the works it cites.
Social iqa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019 · 2019
Later among the works it cites.
BERT and pals: Projected attention layers for efficient adaptation in multi-task learning
Asa Cooper Stickland and Iain Murray. 2019 · 2019
Later among the works it cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jason Phang, Thibault Févry, and Samuel R. Bowman. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
Cross-topic argument mining from heterogeneous sources
Christian Stab, Tristan Miller, Benjamin Schiller, Pranav Rai, and Iryna Gurevych. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018 · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Cited alongside, same era.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Cited alongside, same era.
Simple, scalable adaptation for neural machine translation
Ankur Bapna and Orhan Firat. 2019 · 2019
Cited alongside, same era.
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Later among the works it cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
Parameter-efficient transfer learning with diff pruning
Demi Guo, Alexander M. Rush, and Yoon Kim. 2020 · 2020
Closest in time.
Anne Lauscher, Olga Majewska, Leonardo F. R. Ribeiro, Iryna Gurevych, Nikolai Rozanov, and Goran Glavaš. 2020 · 2020
Closest in time.
AdapterHub: A framework for adapting transformers
Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020a · 2020
Closest in time.
MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020c · 2020
Closest in time.
Monolingual adapters for zero-shot neural machine translation
Jerin Philip, Alexandre Berard, Matthias Gallé, and Laurent Besacier. 2020 · 2020
Closest in time.
Intermediate-task transfer learning with pretrained language models: When and why does it work?
Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, and Samuel Bowman. 2020 · 2020
Closest in time.
MultiCQA: Zero-shot transfer of self-supervised text matching models on a massive scale
Andreas Rücklé, Jonas Pfeiffer, and Iryna Gurevych. 2020b · 2020
Closest in time.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Closest in time.
UDapter: Language adaptation for truly Universal Dependency parsing
Ahmet Üstün, Arianna Bisazza, Gosse Bouma, and Gertjan van Noord. 2020 · 2020
Closest in time.
Orthogonal language and task adapters in zero-shot cross-lingual transfer
M. Vidoni, Ivan Vulić, and Goran Glavaš. 2020 · 2020
Closest in time.
K-adapter: Infusing knowledge into pre-trained models with adapters
Ruize Wang, Duyu Tang, Nan Duan, Zhongyu Wei, Xuanjing Huang, Jianshu Ji, Guihong Cao, Daxin Jiang, and Ming Zhou. 2020 · 2020
Closest in time.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Closest in time.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben-Zaken1 Shauli Ravfogel and Yoav Goldberg. 2021 · 2021
Closest in time.