Fetching the paper…
Reading the bibliography…
Fine-tuning large pre-trained models is an effective transfer mechanism in NLP.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M. and Cohen, N. J · 1989
Earlier work this paper cites.
Class-based n-gram models of natural language
Brown, P. F., deSouza, P. V., Mercer, R. L., Pietra, V. J. D., and Lai, J. C · 1992
Earlier work this paper cites.
Newsweeder: Learning to filter netnews
Lang, K · 1995
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Learning to learn
Thrun, S · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, R. M · 1999
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., and Janvin, C · 2003
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, R. and Weston, J · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., jia Li, L., Li, K., and Fei-fei, L · 2009
Earlier work this paper cites.
Word representations: A simple and general method for semi-supervised learning
Turian, J., Ratinov, L., and Bengio, Y · 2010
Earlier work this paper cites.
Contributions to the Study of SMS Spam Filtering: New Collection and Results
Almeida, T. A., Hidalgo, J. M. G., and Yamakami, A · 2011
Earlier work this paper cites.
Cross-language knowledge transfer using multilingual deep neural network with shared hidden layers
Huang, J.-T., Li, J., Yu, D., Deng, L., and Gong, Y · 2013
Earlier work this paper cites.
UCI machine learning repository, 2013
Lichman, M · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., and Malik, J · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Earlier work this paper cites.
Distributed representations of sentences and documents
Le, Q. and Mikolov, T · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H · 2014
Cited alongside, same era.
Semi-supervised sequence learning
Dai, A. M. and Le, Q. V · 2015
Cited alongside, same era.
Skip-thought vectors
Kiros, R., Zhu, Y., Salakhutdinov, R. R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
Long, J., Shelhamer, E., and Darrell, T · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
What makes imagenet good for transfer learning?
Huh, M., Agrawal, P., and Efros, A. A · 2016
Cited alongside, same era.
Continual learning through synaptic intelligence
Zenke, F., Poole, B., and Ganguli, S · 2017
Later among the works it cites.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V · 2017
Later among the works it cites.
Cer, D., Yang, Y., Kong, S.-y., Hua, N., Limtiaco, N., John, R. S., Constant, N., Guajardo-Cespedes, M., Yuan, S., Tar, C., et al · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Later among the works it cites.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Cited alongside, same era.
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R · 2016
Cited alongside, same era.
Enriching word vectors with subword information
Bojanowski, P., Grave, E., Joulin, A., and Mikolov, T · 2017
Cited alongside, same era.
Coarse-to-fine question answering for long documents
Choi, E., Hewlett, D., Uszkoreit, J., Polosukhin, I., Lacoste, A., and Berant, J · 2017
Cited alongside, same era.
Supervised learning of universal sentence representations from natural language inference data
Conneau, A., Kiela, D., Schwenk, H., Barrault, L., and Bordes, A · 2017
Cited alongside, same era.
Modulating early visual processing by language
De Vries, H., Strub, F., Mary, J., Larochelle, H., Pietquin, O., and Courville, A. C · 2017
Cited alongside, same era.
Kornblith, S., Shlens, J., and Le, Q. V · 2018
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., de Vries, H., Dumoulin, V., and Courville, A. C · 2018
Later among the works it cites.
Deep contextualized word representations
Peters, M., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Later among the works it cites.
Know what you don’t know: Unanswerable questions for squad
Rajpurkar, P., Jia, R., and Liang, P · 2018
Later among the works it cites.
Efficient parametrization of multi-domain deep neural networks
Rebuffi, S., Vedaldi, A., and Bilen, H · 2018
Later among the works it cites.
Incremental learning through deep adaptation
Rosenfeld, A. and Tsotsos, J. K · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Later among the works it cites.
Transfer learning with neural automl
Wong, C., Houlsby, N., Lu, Y., and Gesmundo, A · 2018
Later among the works it cites.
Universal sentence encoder for english
Cer, D., Yang, Y., Kong, S.-y., Hua, N., Limtiaco, N., St. John, R., Constant, N., Guajardo-Cespedes, M., Yuan, S., Tar, C., Strope, B., and Kurzweil, R · 2019
Closest in time.
On self modulation for generative adversarial networks
Chen, T., Lucic, M., Houlsby, N., and Gelly, S · 2019
Closest in time.
BERT-A: Fine-tuning BERT with Adapters and Data Augmentation
Sina Semnani, Kaushik Sadagopan, F. T · 2019
Closest in time.
BERT and PALs: Projected Attention Layers for Efficient Adaptation in Multi-Task Learning
Stickland, A. C. and Murray, I · 2019
Closest in time.