Fetching the paper…
Reading the bibliography…
Large pretrained language models have changed the way researchers approach discriminative natural language understanding tasks, leading to the dominance of approaches that adapt a pretrained model for arbitrary downstream tasks.
Pretraining-Based Natural Language Generation for Text Summarization
Haoyu Zhang, Yeyun Gong, Yu Yan, Nan Duan, Jianjun Xu, Ji Wang, Ming Gong, and Ming Zhou · 1902
Earlier work this paper cites.
A Maximum Likelihood Approach to Continuous Speech Recognition
Lalit R. Bahl, Frederick Jelinek, and Robert L. Mercer · 1983
Earlier work this paper cites.
Statistical Phrase-Based Translation
Philipp Koehn, Franz Josef Och, and Daniel Marcu · 2003
Earlier work this paper cites.
Overview of DUC 2006
Hoa Trang Dang · 2006
Earlier work this paper cites.
Hierarchical Neural Story Generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2011
Earlier work this paper cites.
Distributed Representations of Words and Phrases and their Compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
On Using Monolingual Corpora in Neural Machine Translation
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2015
Earlier work this paper cites.
Teaching Machines to Read and Comprehend
Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Earlier work this paper cites.
Bag of Tricks for Efficient Text Classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov · 2016
Earlier work this paper cites.
A hierarchical approach for generating descriptive image paragraphs
Jonathan Krause, Justin Johnson, Ranjay Krishna, and Li Fei-Fei · 2017
Earlier work this paper cites.
Learned in Translation: Contextualized Word Vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Unsupervised Pretraining for Sequence to Sequence Learning
Prajit Ramachandran, Peter J. Liu, and Quoc V. Le · 2017
Earlier work this paper cites.
Style transfer from non-parallel text by cross-alignment
Tianxiao Shen, Tao Lei, Regina Barzilay, and Tommi Jaakkola · 2017
Cited alongside, same era.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Lei Yu, Phil Blunsom, Chris Dyer, Edward Grefenstette, and Tomas Kocisky · 2017
Cited alongside, same era.
Diverse and coherent paragraph generation from images
Moitreya Chatterjee and Alexander G Schwing · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Simple Fusion: Return of the Language Model
Felix Stahlberg, James Cross, and Veselin Stoyanov · 2018
Later among the works it cites.
Adversarially regularized autoencoders
Junbo Zhao, Yoon Kim, Kelly Zhang, Alexander Rush, and Yann LeCun · 2018
Later among the works it cites.
An embarrassingly simple approach for transfer learning from pretrained language models
Alexandra Chronopoulou, Christos Baziotis, and Alexandros Potamianos · 2019
Closest in time.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon · 2019
Closest in time.
Pre-trained Language Model Representations for Language Generation
Sergey Edunov, Alexei Baevski, and Michael Auli · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sebastian Gehrmann, Yuntian Deng, and Alexander M. Rush · 2018
Cited alongside, same era.
Universal Language Model Fine-tuning for Text Classification
Jeremy Howard and Sebastian Ruder · 2018
Cited alongside, same era.
Training for diversity in image paragraph captioning
Luke Melas-Kyriazi, Alexander Rush, and George Han · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
GPT: Improving Language Understanding by Generative Pre-Training
Alec Radford and Tim Salimans · 2018
Cited alongside, same era.
Cold fusion: Training Seq2seq models together with language models
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates · 2018
Cited alongside, same era.
Learning Word Vectors for Sentiment Analysis
Andrew L Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts
Cited in the paper.
Closest in time.
Large-scale transfer learning for natural language generation
Sergey Golovanov, Rauf Kurbanov, Sergey Nikolenko, Kyryl Truskovskyi, Alexander Tselousov, and Thomas Wolf · 2019
Closest in time.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Closest in time.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau · 2019
Closest in time.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Closest in time.
Mass: Masked sequence to sequence pre-training for language generation
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and T. M. Liu · 2019
Closest in time.
Language models with transformers
Chenguang Wang, Mu Li, and Alexander J. Smola · 2019
Closest in time.