Fetching the paper…
Reading the bibliography…
The recent success of distributed word representations has led to an increased interest in analyzing the properties of their spatial distribution.
Improving neural language modeling via adversarial training
Dilin Wang, Chengyue Gong, and Qiang Liu. 2019 · 1906
Earlier work this paper cites.
Kawin Ethayarajh. 2019 · 1909
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
What do you mean, bert? assessing BERT as a distributional semantics model
Timothee Mickus, Denis Paperno, Mathieu Constant, and Kees van Deemter. 2019 · 1911
Earlier work this paper cites.
Maximum likelihood estimation of intrinsic dimension
Elizaveta Levina and Peter J. Bickel. 2004 · 2004
Earlier work this paper cites.
Random walks on context spaces: Towards an explanation of the mysteries of semantic word embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2015 · 2015
Earlier work this paper cites.
Intrinsic dimension estimation: Relevant techniques and a benchmark framework
Paola Campadelli, Elena Casiraghi, Claudio Ceruti, and Alessandro Rozza. 2015 · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
Word re-embedding via manifold dimensionality retention
S. Hasan and E. Curry. 2017 · 2017
Cited alongside, same era.
All-but-the-top: Simple and effective postprocessing for word representations
Jiaqi Mu, Suma Bhat, and Pramod Viswanath. 2017 · 2017
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Frage: Frequency-agnostic word representation
Chengyue Gong, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan. 2018 · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, R. Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Refining word representations by manifold learning
Chu Yonghe, Hongfei Lin, Liang Yang, Yufeng Diao, Zhang Shaowu, and Fan Xiaochao. 2019 · 2019
Later among the works it cites.
Getting in shape: Word embedding subspaces
Tianyuan Zhou, João Sedoc, and J. Rodu. 2019 · 2019
Later among the works it cites.
Too much in common: Shifting of embeddings in transformer language models and its implications
Daniel Biś, Maksim Podkorytov, and Xiuwen Liu. 2021 · 2021
Closest in time.
Isotropy in the contextual embedding space: Clusters and manifolds
Xingyu Cai, Jiaji Huang, Yuchen Bian, and Kenneth Church. 2021 · 2021
Closest in time.
Learning to remove: Towards isotropic pre-trained BERT embedding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Representation degeneration problem in training natural language generation models
Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2019 · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Visualizing and measuring the geometry of bert
Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Pearce, F. Viégas, and M. Wattenberg. 2019a
Cited in the paper.
Visualizing and measuring the geometry of bert
Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Pearce, Fernanda Viégas, and Martin Wattenberg. 2019b
Cited in the paper.
Yuxin Liang, Rui Cao, Jie Zheng, Jie Ren, and Ling Gao. 2021 · 2021
Closest in time.
Isobn: Fine-tuning bert with isotropic batch normalization
Wenxuan Zhou, Bill Yuchen Lin, and Xiang Ren. 2021 · 2021
Closest in time.