Fetching the paper…
Reading the bibliography…
Recent progress in pretraining language models on large corpora has resulted in large performance gains on many NLP tasks.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
WordNet: a lexical database for english
George A Miller. 1995 · 1995
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Distributional semantics in linguistic and cognitive research
Alessandro Lenci. 2008 · 2008
Earlier work this paper cites.
The westbury lab wikipedia corpus
C. & Westbury C. Shaoul. 2010 · 2010
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, G.s Corrado, Kai Chen, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
SimLex-999: Evaluating semantic models with (genuine) similarity estimation
Felix Hill, Roi Reichart, and Anna Korhonen. 2015 · 2015
Cited alongside, same era.
Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn’t
Anna Gladkova, Aleksandr Drozd, and Satoshi Matsuoka. 2016 · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
Advances in pre-training distributed word representations
Tomas Mikolov, Edouard Grave, Piotr Bojanowski, Christian Puhrsch, and Armand Joulin. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
A. Radford, Jeffrey Wu, R. Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Rare words: A major problem for contextualized embeddings and how to fix it by attentive mimicking
Timo Schick and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
Pretrained encyclopedia: Weakly supervised knowledge-pretrained language model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018a · 2018
Cited alongside, same era.
Dissecting contextual word embeddings: Architecture and representation
Matthew Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih. 2018b · 2018
Cited alongside, same era.
Wenhan Xiong, Jingfei Du, William Yang Wang, and Veselin Stoyanov. 2020 · 2020
Later among the works it cites.