Fetching the paper…
Reading the bibliography…
Contextualized word representations, such as ELMo and BERT, were shown to perform well on various semantic and syntactic tasks.
Assessing BERT’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F Liu, Matt Gardner, Yonatan Belinkov, Matthew Peters, and Noah A Smith. 2019a · 1903
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
BERT for coreference resolution: Baselines and analysis
Mandar Joshi, Omer Levy, Daniel S Weld, and Luke Zettlemoyer. 2019 · 1908
Earlier work this paper cites.
Probing natural language inference models through semantic fragments
Kyle Richardson, Hai Hu, Lawrence S Moss, and Ashish Sabharwal. 2019 · 1909
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek. 1999 · 1999
Earlier work this paper cites.
A tale of a probe and a parser
Rowan Hall Maudslay, Josef Valvoda, Tiago Pimentel, Adina Williams, and Ryan Cotterell. 2020 · 2005
Earlier work this paper cites.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Abhilasha Ravichander, Yonatan Belinkov, and Eduard H. Hovy. 2020 · 2005
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron C. Courville, and Pascal Vincent. 2013 · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013b · 2013
Earlier work this paper cites.
Linguistic regularities in sparse and explicit word representations
Omer Levy and Yoav Goldberg. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Deep metric learning using triplet network
Elad Hoffer and Nir Ailon. 2015 · 2015
Earlier work this paper cites.
An improved non-monotonic transition system for dependency parsing
Matthew Honnibal and Mark Johnson. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
FaceNet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015 · 2015
Cited alongside, same era.
Learning structured output representation using deep conditional generative models
Kihyuk Sohn, Honglak Lee, and Xinchen Yan. 2015 · 2015
Cited alongside, same era.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
Deep variational information bottleneck
Alexander Alemi, Ian Fischer, Joshua V. Dillon, and Murphy Murphy. 2016 · 2016
Cited alongside, same era.
Deep biaffine attention for neural dependency parsing
Timothy Dozat and Christopher D Manning. 2016 · 2016
Cited alongside, same era.
A two-step disentanglement method
Naama Hadad, Lior Wolf, and Moni Shahar. 2018 · 2018
Later among the works it cites.
Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem H. Zuidema. 2018 · 2018
Later among the works it cites.
Constituency parsing with a self-attentive encoder
Nikita Kitaev and Dan Klein. 2018 · 2018
Later among the works it cites.
Multiple-attribute text rewriting
Guillaume Lample, Sandeep Subramanian, Eric Smith, Ludovic Denoyer, Marc’Aurelio Ranzato, and Y-Lan Boureau. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Modeling garden path effects without explicit hierarchical syntax
Marten van Schijndel and Tal Linzen. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Disentangling factors of variation in deep representation using adversarial training
Michaël Mathieu, Junbo Jake Zhao, Pablo Sprechmann, Aditya Ramesh, and Yann LeCun. 2016 · 2016
Cited alongside, same era.
Controlling linguistic style aspects in neural language generation
Jessica Ficler and Yoav Goldberg. 2017 · 2017
Cited alongside, same era.
Texture synthesis and style transfer using perceptual image representations from convolutional neural networks
Leon A. Gatys. 2017 · 2017
Cited alongside, same era.
spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing
Matthew Honnibal and Ines Montani. 2017 · 2017
Cited alongside, same era.
Toward controlled generation of text
Zhiting Hu, Zichao Yang, Xiaodan Liang, Ruslan Salakhutdinov, and Eric P. Xing. 2017 · 2017
Cited alongside, same era.
Learning disentangled representations with semi-supervised deep generative models
Siddharth Narayanaswamy, Brooks Paige, Jan-Willem van de Meent, Alban Desmaison, Noah D. Goodman, Pushmeet Kohli, Frank D. Wood, and Philip H. S. Torr. 2017 · 2017
Cited alongside, same era.
Reconstruction-based disentanglement for pose-invariant face recognition
Xi Peng, Xiang Yu, Kihyuk Sohn, Dimitris N. Metaxas, and Manmohan Chandraker. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
What do RNN language models learn about filler–gap dependencies?
Ethan Wilcox, Roger Levy, Takashi Morita, and Richard Futrell. 2018 · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Later among the works it cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
Specializing word embeddings (for parsing) by information bottleneck
Xiang Lisa Li and Jason Eisner. 2019 · 2019
Later among the works it cites.
Open sesame: Getting inside bert’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 2019
Later among the works it cites.
Visualizing and measuring the geometry of bert
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019 · 2019
Later among the works it cites.
End-to-end open-domain question answering with BERTserini
Wei Yang, Yuqing Xie, Aileen Lin, Xingyu Li, Luchen Tan, Kun Xiong, Ming Li, and Jimmy Lin. 2019 · 2019
Later among the works it cites.
When bert forgets how to pos: Amnesic probing of linguistic properties and mlm predictions
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2020 · 2020
Closest in time.