Fetching the paper…
Reading the bibliography…
To measure how well pretrained representations encode some linguistic property, it is common to use accuracy of a probe, i.e.
Learning and evaluating general linguistic intelligence
Dani Yogatama, Cyprien de Masson d’Autume, Jerome Connor, Tomas Kocisky, Mike Chrzanowski, Lingpeng Kong, Angeliki Lazaridou, Wang Ling, Lei Yu, Chris Dyer, and Phil Blunsom. 2019 · 1901
Earlier work this paper cites.
How can we know what language models know?
Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2019 · 1911
Earlier work this paper cites.
Bert is not a knowledge base (yet): Factual knowledge vs. name-based reasoning in unsupervised qa
Nina Poerner, Ulli Waltinger, and Hinrich Schütze. 2019 · 1911
Earlier work this paper cites.
olmpics – on what language model pre-training captures
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2019 · 1912
Earlier work this paper cites.
Universal coding, information, prediction, and estimation
Jorma Rissanen. 1984 · 1984
Earlier work this paper cites.
Keeping neural networks simple by minimising the description length of weights
GE Hinton and D von Cramp. 1993 · 1993
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Information theory, inference and learning algorithms
David JC MacKay. 2003 · 2003
Earlier work this paper cites.
A tutorial introduction to the minimum description length principle
Peter Grunwald. 2004 · 2004
Earlier work this paper cites.
Variational learning and bits-back coding: an information-theoretic view to bayesian learning
Antti Honkela and Harri Valpola. 2004 · 2004
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2014 · 2014
Earlier work this paper cites.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Assessing the ability of lstms to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
Convolutional neural network language models
Ngoc-Quan Pham, German Kruszewski, and Gemma Boleda. 2016 · 2016
Cited alongside, same era.
End-to-end neural coreference resolution
Kenton Lee, Luheng He, Mike Lewis, and Luke Zettlemoyer. 2017 · 2017
Cited alongside, same era.
Bayesian compression for deep learning
Christos Louizos, Karen Ullrich, and Max Welling. 2017 · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov. 2017 · 2017
Cited alongside, same era.
Arc-swift: A novel transition system for dependency parsing
Peng Qi and Christopher D. Manning. 2017 · 2017
Cited alongside, same era.
Language modeling teaches you more than translation does: Lessons learned through auxiliary syntactic task analysis
Kelly Zhang and Samuel Bowman. 2018 · 2018
Later among the works it cites.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2019 · 2019
Later among the works it cites.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass. 2019 · 2019
Later among the works it cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The description length of deep learning models
Léonard Blier and Yann Ollivier. 2018 · 2018
Cited alongside, same era.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato and Jörg Tiedemann. 2018 · 2018
Cited alongside, same era.
The importance of being recurrent for modeling hierarchical structure
Ke Tran, Arianna Bisazza, and Christof Monz. 2018 · 2018
Cited alongside, same era.
John Hewitt and Percy Liang. 2019 · 2019
Later among the works it cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Later among the works it cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Understanding learning dynamics of language models with SVCCA
Naomi Saphra and Adam Lopez. 2019 · 2019
Later among the works it cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al. 2019 · 2019
Later among the works it cites.
The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives
Elena Voita, Rico Sennrich, and Ivan Titov. 2019a · 2019
Later among the works it cites.
No training required: Exploring random encoders for sentence classification
John Wieting and Douwe Kiela. 2019 · 2019
Later among the works it cites.