Fetching the paper…
Reading the bibliography…
Multi-task learning shares information between related tasks, sometimes reducing the number of parameters required.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L · 2001
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2006
Earlier work this paper cites.
Natural language processing (almost) from scratch
Collobert, R., Weston, J., Bottou, L., Karlen, M., Kavukcuoglu, K., and Kuksa, P · 2011
Earlier work this paper cites.
New types of deep neural network learning for speech recognition and related applications: an overview
Deng, L., Hinton, G., and Kingsbury, B · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Learning hidden unit contributions for unsupervised speaker adaptation of neural network acoustic models
Swietojanski, P. and Renals, S · 2014
Earlier work this paper cites.
Semi-supervised sequence learning
Dai, A. M. and Le, Q. V · 2015
Earlier work this paper cites.
Low resource dependency parsing: Cross-lingual parameter sharing in a neural network parser
Duong, L., Cohn, T., Bird, S., and Cook, P · 2015
Earlier work this paper cites.
Ba, J., Kiros, R., and Hinton, G. E · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Cited alongside, same era.
A joint many-task model: Growing a neural network for multiple nlp tasks
Hashimoto, K., Xiong, C., Tsuruoka, Y., and Socher, R · 2017
Cited alongside, same era.
Gradient episodic memory for continual learning
Lopez-Paz, D. and Ranzato, M · 2017
Cited alongside, same era.
An overview of multi-task learning in deep neural networks
Ruder, S · 2017
Cited alongside, same era.
Distral: Robust multitask reinforcement learning
Teh, Y., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Improving language understanding by generative pre-training
Radford, A · 2018
Later among the works it cites.
Efficient parametrization of multi-domain deep neural networks
Rebuffi, S.-A., Bilen, H., and Vedaldi, A · 2018
Later among the works it cites.
A hierarchical multi-task approach for learning embeddings from semantic tasks
Sanh, V., Wolf, T., and Ruder, S · 2018
Later among the works it cites.
Learning general purpose distributed sentence representations via large scale multi-task learning
Subramanian, S., Trischler, A., Bengio, Y., and Pal, C. J · 2018
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trace norm regularised deep multi-task learning
Yang, Y. and Hospedales, T. M · 2017
Cited alongside, same era.
Looking for ELMo’s friends: Sentence-level pretraining beyond language modeling
Bowman, S. R., Pavlick, E., Grave, E., Durme, B. V., Wang, A., Hula, J., Xia, P., Pappagari, R., McCoy, R. T., Patel, R., Kim, N., Tenney, I., Huang, Y., Yu, K., Jin, S., and Chen, B · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S · 2018
Cited alongside, same era.
The natural language decathlon: Multitask learning as question answering
McCann, B., Keskar, N. S., Xiong, C., and Socher, R · 2018
Cited alongside, same era.
Sentence encoders on STILTSs: Supplementary training on intermediate labeled-data tasks
Phang, J., Févry, T., and Bowman, S. R · 2018
Cited alongside, same era.
Warstadt, A., Singh, A., and Bowman, S. R · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Later among the works it cites.
Swag: A large-scale adversarial dataset for grounded commonsense inference
Zellers, R., Bisk, Y., Schwartz, R., and Choi, Y · 2018
Later among the works it cites.
Self-attention generative adversarial networks
Zhang, H., Goodfellow, I. J., Metaxas, D. N., and Odena, A · 2018
Later among the works it cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J. G., Le, Q. V., and Salakhutdinov, R · 2019
Closest in time.
Parameter-Efficient Transfer Learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Closest in time.
Multi-task deep neural networks for natural language understanding
Liu, X., He, P., Chen, W., and Gao, J · 2019
Closest in time.