Fetching the paper…
Reading the bibliography…
Pretraining deep language models has led to large performance gains in NLP.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 1906
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Bad form: Comparing context-based and form-based few-shot learning in distributional semantic models
Jeroen Van Hautte, Guy Emerson, and Marek Rei. 2019 · 1910
Earlier work this paper cites.
What would Elsa do? Freezing layers during transformer fine-tuning
Jaejun Lee, Raphael Tang, and Jimmy Lin. 2019 · 1911
Earlier work this paper cites.
Wordnet: A lexical database for english
George A. Miller. 1995 · 1995
Earlier work this paper cites.
The westbury lab wikipedia corpus
Cyrus Shaoul and Chris Westbury. 2010 · 2010
Earlier work this paper cites.
Better word representations with recursive neural networks for morphology
Thang Luong, Richard Socher, and Christopher Manning. 2013 · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
DBpedia - a large-scale, multilingual knowledge base extracted from wikipedia
Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N. Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick van Kleef, Sören Auer, and Christian Bizer. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
context2vec: Learning generic context embedding with bidirectional LSTM
Oren Melamud, Jacob Goldberger, and Ido Dagan. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Cited alongside, same era.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
High-risk learning: acquiring new word vectors from tiny data
Aurélie Herbelot and Marco Baroni. 2017 · 2017
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Incorporating subword information into matrix factorization word embeddings
Alexandre Salle and Aline Villavicencio. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
Cloze-driven pretraining of self-attention networks
Alexei Baevski, Sergey Edunov, Yinhan Liu, Luke Zettlemoyer, and Michael Auli. 2019 · 2019
Closest in time.
Google T5 explores the limits of transfer learning
Sam Bowman. 2019 · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multimodal word meaning induction from minimal exposure to natural text
Angeliki Lazaridou, Marco Marelli, and Marco Baroni. 2017 · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017 · 2017
Cited alongside, same era.
Mimicking word embeddings using subword RNNs
Yuval Pinter, Robert Guthrie, and Jacob Eisenstein. 2017 · 2017
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2017
Cited alongside, same era.
HotFlip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Cited alongside, same era.
A la carte embedding: Cheap but effective induction of semantic feature vectors
Mikhail Khodak, Nikunj Saunshi, Yingyu Liang, Tengyu Ma, Brandon Stewart, and Sanjeev Arora. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Closest in time.
Second-order contexts from lexical substitutes for few-shot learning of word representations
Qianchu Liu, Diana McCarthy, and Anna Korhonen. 2019a · 2019
Closest in time.
Misspelling oblivious word embeddings
Aleksandra Piktus, Necati Bora Edizel, Piotr Bojanowski, Edouard Grave, Rui Ferreira, and Fabrizio Silvestri. 2019 · 2019
Closest in time.
Attentive mimicking: Better word embeddings by attending to informative contexts
Timo Schick and Hinrich Schütze. 2019a · 2019
Closest in time.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 2019
Closest in time.
Rare words: A major problem for contextualized embeddings and how to fix it by attentive mimicking
Timo Schick and Hinrich Schütze. 2020 · 2020
Closest in time.
Pattern for python
Tom De Smedt and Walter Daelemans. 2012 · 2067
Closest in time.