Fetching the paper…
Reading the bibliography…
While most previous work has focused on different pretraining objectives and architectures for transfer learning, we ask how to best adapt the pretrained model to a given target task.
Comparison of the predicted and observed secondary structure of t4 phage lysozyme
Brian W Matthews. 1975 · 1975
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando Pereira. 2001 · 2001
Earlier work this paper cites.
Statistical phrase-based translation
Philipp Koehn, Franz Josef Och, and Daniel Marcu. 2003 · 2003
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang. 2010 · 2010
Earlier work this paper cites.
Word representations: A simple and general method for semi-supervised learning
Joseph P. Turian, Lev-Arie Ratinov, and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Y Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Convolutional Neural Networks for Sentence Classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
A sick cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, Roberto Zamparelli, et al. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. 2014 · 2014
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M. Dai and Quoc V. Le. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Neural Architectures for Named Entity Recognition
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016 · 2016
Cited alongside, same era.
How Transferable are Neural Networks in NLP Applications?
Lili Mou, Zhao Meng, Rui Yan, Ge Li, Yan Xu, Lu Zhang, and Zhi Jin. 2016 · 2016
Cited alongside, same era.
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Rethinking ImageNet Pre-training
Kaiming He, Ross Girshick, and Piotr Dollár. 2018 · 2018
Later among the works it cites.
Universal Language Model Fine-tuning for Text Classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Later among the works it cites.
Do Better ImageNet Models Transfer Better?
Simon Kornblith, Jonathon Shlens, Quoc V Le, and Google Brain. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel M. Cer, Mona T. Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
Enhanced LSTM for Natural Language Inference
Qian Chen, Xiaodan Zhu, Zhenhua Ling, Si Wei, Hui Jiang, and Diana Inkpen. 2017 · 2017
Cited alongside, same era.
Supervised Learning of Universal Sentence Representations from Natural Language Inference Data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes. 2017 · 2017
Cited alongside, same era.
Bjarke Felbo, Alan Mislove, Anders Søgaard, Iyad Rahwan, and Sune Lehmann. 2017 · 2017
Cited alongside, same era.
Allennlp: A deep semantic natural language processing platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke S. Zettlemoyer. 2017 · 2017
Cited alongside, same era.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Cited alongside, same era.
Learned in Translation: Contextualized Word Vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
An efficient framework for learning sentence representations
Lajanugen Logeswaran and Honglak Lee. 2018 · 2018
Later among the works it cites.
Scalable Mutual Information Estimation using Dependence Graphs
Morteza Noshad, Yu Zeng, and Alfred O. Hero III. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving Language Understanding by Generative Pre-Training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
On the Information Bottleneck Theory of Deep Learning
Andrew M Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D Tracey, and David D Cox. 2018 · 2018
Later among the works it cites.
Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning
Sandeep Subramanian, Adam Trischler, Yoshua Bengio, and Christopher J Pal. 2018 · 2018
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Later among the works it cites.