Fetching the paper…
Reading the bibliography…
In this work, we study the representation space of contextualized embeddings and gain insight into the hidden topology of large language models.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
An introduction to conditional random fields
Charles Sutton, Andrew McCallum, et al · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Max Welling Diederik P. Kingma · 2013
Earlier work this paper cites.
Conditional random field autoencoders for unsupervised structured prediction
Waleed Ammar, Chris Dyer, and Noah A. Smith · 2014
Earlier work this paper cites.
From word embeddings to document distances
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger · 2015
Earlier work this paper cites.
Differentiable dynamic programming for structured prediction and attention
Arthur Mensch and Mathieu Blondel · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning · 2019
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith · 2019
Earlier work this paper cites.
Visualizing and measuring the geometry of bert
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim · 2019
Cited alongside, same era.
Finding universal grammatical relations in multilingual BERT
Ethan A. Chi, John Hewitt, and Christopher D. Manning · 2020
Cited alongside, same era.
Latent template induction with gumbel-crf
Yao Fu, Chuanqi Tan, Bin Bi, Mosha Chen, Yansong Feng, and Alexander M. Rush · 2020
Cited alongside, same era.
A tale of a probe and a parser
Rowan Hall Maudslay, Josef Valvoda, Tiago Pimentel, Adina Williams, and Ryan Cotterell · 2020
Cited alongside, same era.
Are pre-trained language models aware of phrases? simple but strong baselines for grammar induction
Taeuk Kim, Jihun Choi, Daniel Edmiston, and Sang goo Lee · 2020
Cited alongside, same era.
Posterior control of blackbox generation
Xiang Lisa Li and Alexander Rush · 2020
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2020
Later among the works it cites.
Perturbed masking: Parameter-free probing for analyzing and interpreting BERT
Zhiyong Wu, Yun Chen, Ben Kao, and Qun Liu · 2020
Later among the works it cites.
Isotropy in the contextual embedding space: Clusters and manifolds
Xingyu Cai, Jiaji Huang, Yuchen Bian, and Kenneth Church · 2021
Later among the works it cites.
Probing {bert} in hyperbolic spaces
Boli Chen, Yao Fu, Guangwei Xu, Pengjun Xie, Chuanqi Tan, Mosha Chen, and Liping Jing · 2021
Later among the works it cites.
Obtaining better static word embeddings using contextual embedding models
Prakhar Gupta and Martin Jaggi · 2021
Later among the works it cites.
Conditional probing: measuring usable information beyond a baseline
John Hewitt, Kawin Ethayarajh, Percy Liang, and Christopher Manning · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
What do neural networks learn when trained with random labels?
Hartmut Maennel, Ibrahim M Alabdulmohsin, Ilya O Tolstikhin, Robert Baldock, Olivier Bousquet, Sylvain Gelly, and Daniel Keysers · 2020
Cited alongside, same era.
Emergence of separable manifolds in deep language representations
Jonathan Mamou, Hang Le, Miguel A Del Rio, Cory Stephenson, Hanlin Tang, Yoon Kim, and SueYeon Chung · 2020
Cited alongside, same era.
Asking without telling: Exploring latent ontologies in contextual representations
Julian Michael, Jan A. Botha, and Ian Tenney · 2020
Cited alongside, same era.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick
Cited in the paper.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick
Cited in the paper.
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
Discovering latent concepts learned in BERT
Fahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani, Jia Xu, and Hassan Sajjad · 2022
Closest in time.
Scaling structured inference with randomization
Yao Fu, John P. Cunningham, and Mirella Lapata · 2022
Closest in time.