Fetching the paper…
Reading the bibliography…
Autoregressive language models, pretrained using large text corpora to do well on next word prediction, have been successful at solving many downstream tasks, even with zero-shot usage.
Distributional structure
Zellig Harris · 1954
Earlier work this paper cites.
A synopsis of linguistic theory, 1930-1955
John R Firth · 1957
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F Chen and Joshua Goodman · 1999
Earlier work this paper cites.
Can artificial neural networks learn language models?
Wei Xu and Alex Rudnicky · 2000
Earlier work this paper cites.
Learning question classifiers
Xin Li and Dan Roth · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin · 2003
Earlier work this paper cites.
Contrastive estimation reveals topic posterior information to linear models
Christopher Tosh, Akshay Krishnamurthy, and Daniel Hsu · 2003
Earlier work this paper cites.
Annotating opinions in the world press
Theresa Wilson and Janyce Wiebe · 2003
Earlier work this paper cites.
Mining and summarizing customer reviews
Minqing Hu and Bing Liu · 2004
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives · 2007
Earlier work this paper cites.
Contrastive learning, multi-view redundancy, and linear models
Christopher Tosh, Akshay Krishnamurthy, and Daniel Hsu · 2008
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Notes on noise contrastive estimation and negative sampling
Chris Dyer · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Neural word embedding as implicit matrix factorization
Omer Levy and Yoav Goldberg · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Russ R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Inferring networks of substitutable and complementary products
Julian J. McAuley, Rahul Pandey, and Jure Leskovec · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Cited alongside, same era.
A latent variable model approach to pmi-based word embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
A simple but tough-to-beat baseline for sentence embeddings
Sanjeev Arora, Yingyu Liang, and Tengyu Ma · 2017
Cited alongside, same era.
Learned in translation: Contextualized word vectors
Unsupervised learning of sentence embeddings using compositional n-gram features
Matteo Pagliardini, Prakhar Gupta, and Martin Jaggi · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2018
Later among the works it cites.
A theoretical analysis of contrastive unsupervised representation learning
Sanjeev Arora, Hrishikesh Khandeparkar, Mikhail Khodak, Orestis Plevrakis, and Nikunj Saunshi · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher · 2017
Cited alongside, same era.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever · 2017
Cited alongside, same era.
Unsupervised pretraining for sequence to sequence learning
Prajit Ramachandran, Peter Liu, and Quoc Le · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
A compressed sensing view of unsupervised text embeddings, bag-of-n-grams, and LSTMs
Sanjeev Arora, Mikhail Khodak, Nikunj Saunshi, and Kiran Vodrahalli · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Cited alongside, same era.
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Later among the works it cites.
Zero-shot text classification with generative language models
Raul Puri and Bryan Catanzaro · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Later among the works it cites.
olmpics–on what language model pre-training captures
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Closest in time.
Predicting what you already know helps: provable self-supervised learning
Jason D. Lee, Qi Lei, Nikunj Saunshi, and Jiacheng Zhuo · 2020
Closest in time.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze · 2020
Closest in time.
Tongzhou Wang and Phillip Isola · 2020
Closest in time.